Decision model of the ufakzeka family
ufakzeka-karar
It reads a Turkish text and answers the typed questions you ask about it. Depending on the question, it picks one option from a set, places the text at a level on an ordered scale, or says yes or no. It gives every option a probability, temperature-scaled on validation data, and alongside it works out an expected error you can use as a not-sure signal (abstain). It generates no text, and it reads each question in one forward pass on an ordinary CPU.
- 182,494,466parameters
- 102.5 msmedian per question, one thread on an Apple M1 CPU
- 7 / 16rank on the HakemBench board
- Apache-2.0licence
Text in, probabilities out
A request is a text (the state) and a set of named questions. The model reads each question together with the text and returns a probability distribution over the options. Options do not see each other and share positions, so their order leaves the probabilities bit for bit the same up to ten options, and the same up to floating-point noise above ten. One pass reads at most 10 options; larger sets are split and the probabilities joined into one distribution.
The text and the question are read together within 448 tokens. The end of a longer text is cut; the question is always read. Each option is read up to 48 tokens.
Choice
The model picks one option from a set of 2 to 255. The answer holds the most likely option, every option’s probability and a confidence.
Score
The model places the text at one level of an ordered scale of 2 to 10 levels. The answer holds the expected level, the distribution over levels and a confidence.
Yes/no (noul)
The model answers a single yes or no question, and the answer is the probability of yes.
The not-sure signal is the model’s estimate of how likely it is to be wrong on this question, read from a map learned from its own errors. Pick a confidence threshold and send the answers below it to a person. On HakemBench, answers with confidence 0.90 or more cover 29 percent of the questions, and 6 percent of those answers are wrong. The not-sure signal does miss guardrail errors, though. Of its answers to guardrail questions at 0.99 confidence or more, 12 percent are wrong, and on customer support almost no answer reaches 0.90. These guardrail and support numbers carry the flag “shaped by reading the test results”, explained under the HakemBench results below.
Where it helps
Customer support
It decides whether a request needs a person, whether a reply answers the question and whether the topic is sensitive. The model is weak here; see the limits below.
Moderation
It tells whether a message contains profanity, insults or abusive language.
Guardrails
It judges whether a message to an assistant tries to override its instructions or make it do something it may not (prompt injection). Use it as one layer among others, never as the only one.
Spam and phishing
It tells what kind of message it is (scam, unsolicited ad, marketing the reader signed up for, transaction notice, personal message) and whether it tries to trick the reader.
Fact-check triage
It judges whether a sentence holds a checkable claim and how much of a priority checking it is. The model does not decide whether the claim is true.
Grading in education
It grades a student’s answer and tells which subject a question belongs to.
Legal routing
It tells which court or office a dispute or request goes to, and which fundamental right an individual application concerns.
Request and response
Requests and responses use the field names of Jev’s typed-decision API (state, questions, type, instructions, criteria). Our one addition is the abstain field. The example texts are in Turkish, as the model reads them.
{
"model": "ufakzeka-karar",
"state": "Faturam iki kez kesildi, iade istiyorum.",
"questions": {
"konu": {
"type": "choice",
"instructions": "Mesajın konusu nedir?",
"criteria": {
"fatura": null,
"kargo": null,
"iade": null
}
},
"temsilci": {
"type": "noul",
"instructions": "Bu mesaj bir müşteri temsilcisine yönlendirilmeli mi?"
},
"ofke": {
"type": "score",
"instructions": "Müşteri ne kadar öfkeli?",
"criteria": [
"Sakin",
"Rahatsız",
"Öfkeli"
]
}
}
}The release model’s answer to this request, unchanged; numbers are rounded to three decimals. The text mentions both a bill and a refund, so on the topic question the model is split between two options; its confidence is under 50 percent and it says “not sure”.
{
"model": "ufakzeka-karar",
"answers": {
"konu": {
"type": "choice",
"choice": "iade",
"probabilities": {
"fatura": 0.459,
"kargo": 0.07,
"iade": 0.471
},
"confidence": 0.42,
"abstain": 0.58
},
"temsilci": {
"type": "noul",
"noul": 0.126,
"abstain": 0.145
},
"ofke": {
"type": "score",
"score": 0.915,
"legend": {
"0": "Sakin",
"1": "Rahatsız",
"2": "Öfkeli"
},
"probabilities": {
"0": 0.404,
"1": 0.278,
"2": 0.319
},
"confidence": 0.369,
"abstain": 0.631
}
}
}# hf download ufakai/ufakzeka-karar karar.py requirements.txt --local-dir .
# pip install -r requirements.txt
from karar import Karar
karar = Karar.from_pretrained("ufakai/ufakzeka-karar")
karar.decide(
"Faturam iki kez kesildi, iade istiyorum.",
{"konu": {"type": "choice", "instructions": "Mesajın konusu nedir?",
"criteria": {"fatura": None, "kargo": None, "iade": None}}},
)probabilities- temperature-scaled probabilities, summing to 1
confidence- 1 minus abstain; on choice and score answers
abstain- estimated probability that the model is wrong here
score- expected level, counted from 0
noul- probability of yes
Results on HakemBench
- 0.660composite score, [0.642, 0.677]
- 7 / 16rank on the board
- 0.705decision quality (macro F1), [0.688, 0.718]
- 0.482calibration axis (higher is better), [0.456, 0.507]
Values in square brackets are 95 percent bootstrap intervals.
On the 4,275 questions of HakemBench’s open set, ufakzeka-karar is 7th of the board’s 16 rows, with a composite of 0.660 [0.642, 0.677]. The models above it are Gemini 3.8 Flash, GPT-5.6 Sol, GLM 5.3, Jev 1.13, Kev 4B and Kev 9B. Three of them are large language models reached through an API, and Jev 1.13 is a decision model served through an API. Kev 4B and Kev 9B are open-weight decision models of billions of parameters. Some named hosted models share a family with models that wrote or labelled parts of the benchmark data (HakemBench page).
Its strengths lie elsewhere. Its weights are open (Apache-2.0) and at 182,494,466 parameters it runs on an ordinary CPU, taking a median of 102.5 ms a question on one thread. Option order cannot change its answer (order sensitivity 0.000), so half of its robustness value is 1.0 by construction; that value is 0.941, against 0.943 for Jev 1.13. On the measured half, paraphrase agreement, it is 0.883 [0.833, 0.932] against 0.926 [0.889, 0.963] for Jev 1.13, and the intervals overlap.
On a small set of 84 support questions answered by hand, its accuracy is 0.476 [0.369, 0.583], below the 0.619 [0.512, 0.714] of the surface-cue baseline, which looks only at cues such as length, a digit or a question mark, not at meaning. The human answers on this page come from one person; one annotator makes mistakes too, so the figures based on them are indicative.
| Track | Questions | Accuracy | Macro F1 | Brier score | Calibration error |
|---|---|---|---|---|---|
| Fact-check triage | 602 | 0.640 [0.602, 0.681] | 0.437 [0.390, 0.481] | 0.499 [0.460, 0.538] | 0.037 [0.033, 0.066] |
| Education | 640 | 0.787 [0.756, 0.817] | 0.802 [0.773, 0.826] | 0.278 [0.249, 0.308] | 0.025 [0.025, 0.047] |
| Guardrails* | 418 | 0.770 [0.730, 0.811] | 0.764 [0.722, 0.805] | 0.367 [0.305, 0.433] | 0.158 [0.119, 0.200] |
| Legal routing | 308 | 0.782 [0.737, 0.825] | 0.754 [0.690, 0.797] | 0.291 [0.238, 0.350] | 0.056 [0.036, 0.097] |
| Moderation* | 353 | 0.952 [0.929, 0.975] | 0.951 [0.927, 0.973] | 0.099 [0.077, 0.123] | 0.084 [0.063, 0.105] |
| Spam and phishing | 614 | 0.806 [0.774, 0.839] | 0.746 [0.708, 0.779] | 0.271 [0.237, 0.307] | 0.035 [0.029, 0.060] |
| Customer support* | 1,340 | 0.569 [0.534, 0.605] | 0.479 [0.446, 0.511] | 0.526 [0.497, 0.554] | 0.069 [0.048, 0.095] |
The values are computed on the open set from the probabilities the model returns. Brackets are 95 percent bootstrap intervals. Calibration error is smooth expected calibration error (smooth ECE), the gap between the confidence the model states and how often it is right; lower is better. The three starred rows (guardrails, moderation and customer support) carry the note below.
The 181 court items and the 28 source guardrail items are the test splits of datasets whose train splits are in ufakzeka-karar’s training data (3,000 court rows, and 1,791 rows from three prompt-injection sets that include the guardrail items' two source sets; near-duplicates removed), so on those questions the lab’s model is supervised in-distribution while the other models answer zero-shot; its legal routing score and the guardrail counts should be read with that.
On guardrails it caught 194 of 217 prompt-injection attacks and let 128 of 201 harmless messages through. On moderation it flagged 150 of 157 offensive messages and let 186 of 196 clean ones through.*
* Its numbers are not blind. Earlier runs’ test results shaped its training data, so its guardrail, moderation and customer support numbers carry the flag “shaped by reading the test results”; they are starred in the table and in the paragraph above, and marked where the page quotes them elsewhere. How this happened is told below, under “The road to the release”. With every model scored on the other four tracks only, its composite is 0.678, 6th of 16.
Where it is weak
- It is weak on customer support. In that track its accuracy is 0.490 on choice questions and 0.481 on score questions. These numbers carry the starred note above. On a small hand-answered set it stays below the surface-cue baseline (0.476 against 0.619), and the baseline is also ahead on the support choice (0.516) and score (0.507) questions of the open set (partly in-sample). Do not leave support routing to this model alone.
- It is weak at scoring how much of a priority a fact check is; on those questions its accuracy is 0.500 and its macro F1 0.268.
- On guardrails it raises many false alarms: it flags 73 of 201 harmless messages and lets 128 through. These numbers carry the starred note above.
- Its moderation result may be inflated. The moderation comments written for its training resemble in style the moderation texts written for the benchmark, which make up all of the open set’s moderation items.
- Its world knowledge is weak. On the MMLU-Pro-TR knowledge exam its accuracy is 0.098 [0.092, 0.103], slightly below the chance level of 0.111. Because the model does not know facts the text does not state, questions whose answer lies in general knowledge rather than in the text can go wrong.
- The shipped temperature makes calibration worse on held-out support questions (0.036 to 0.064 for the released model, which has since trained on those questions). On HakemBench the model is somewhat overconfident (post-hoc temperature 1.1733). On guardrails, 12 percent of the answers at 0.99 confidence or more are wrong (the starred note above applies). On a kind of question it never saw in training, trust the confidence less.
- All human answers come from one person; there is no inter-annotator agreement. Most of the gold comes from the passes of one AI model family.
- It does not explain itself; it gives only probabilities and does not say why it chose an answer.
The road to the release
The released model is the last of three runs scored on HakemBench, and its numbers are not blind. The second run’s new training data was aimed at the first run’s errors on the full test set in guardrails, moderation and customer support. The second run’s guardrail results on the full test set then showed that it caught only 106 of 217 attacks, most likely through a shortcut on the writing style, since its only guardrail training texts written like a real user’s message were harmless. The released run was trained after those results were read, under a protocol fixed in writing before any of its data, code or runs. To correct that shortcut it added an openly licensed prompt-injection set written as conversation, whose texts a model wrote, and the model was chosen from six candidates under a rule fixed in writing beforehand. The selection rule and its result ship with the code. Every number of the model on this page comes after these readings. The research post tells the whole road.
On other Turkish test sets
| Set | Questions | Chance | Accuracy | Macro F1 |
|---|---|---|---|---|
| MASSIVE 1.1, tr-TR intents | 2,937 | 0.017 | 0.745 | 0.691 |
| MiDe22 Turkish tweets | 1,012 | 0.333 | 0.741 | 0.727 |
| MMLU-Pro-TR | 11,838 | 0.111 | 0.098 | |
| OffensEval-TR 2020, subtask A | 3,528 | 0.500 | 0.844 | 0.755 |
| XCOPA, Turkish | 500 | 0.500 | 0.584 | |
| X-FACT, Turkish test claims | 169 | 0.250 | 0.343 | 0.207 |
Each test item was asked as a typed question with no examples in the request, and each run was scored with its own calibration; the table shows the released model. The train splits of OffensEval-TR, MASSIVE and MiDe22 are in the training data, so on those three sets the numbers are supervised, not zero-shot transfer. Against Run 1, the first run scored on HakemBench, there are small drops on MMLU-Pro-TR, OffensEval-TR, MASSIVE and MiDe22 and small gains on XCOPA and X-FACT. MMLU-Pro-TR is a knowledge exam; its result is given above, among the weak spots.
Runs on a CPU
- On one machine (Apple M1, one torch thread), measured on 13 sample questions, the median was 102.5 ms per question and the peak resident memory 986,038,272 bytes. This is a one-machine measurement, not a benchmark.
- In the HakemBench run, on a Mac’s CPU, the median time per question was 103.6 [102.4, 106.8] ms. Speed does not enter the composite, and the hardware differs from model to model; open-jev and Laya at its shipped length also ran on a Mac’s CPU, Kev, Qwen3.5-4B and decider-2b on an NVIDIA L4 GPU, Laya at full length on a GPU, and the API models were measured over the network.
- The model is released in fp32 (float32). No int8 file is published, because the int8 files tried on an earlier checkpoint did not stay within 0.5 points (0.005) of fp32 on macro F1, although their calibration error stayed within 0.11 points (0.0011). No int8 version of the release model was measured.
What calibration changed
| Temperature | Held-out support questions, before temperature | Held-out support questions, after temperature | Development set, before temperature | Development set, after temperature |
|---|---|---|---|---|
| 1.2614 | 0.036 | 0.064 | 0.050 | 0.038 |
Apart from the temperature, the values are calibration error. Temperature scaling raises the calibration error on the held-out support questions (the four questions HakemBench asks of every support item) and lowers it on the development set, a set of HakemBench-style support questions kept apart from the test set. The released model has since trained on those four questions, so its held-out result is not an unseen-question test; for the first model scored on HakemBench, which never trained on them, its own temperature raised the error from 0.027 to 0.045. The criterion on the selection split was fixed before the held-out result was seen, so T = 1.2614 ships as chosen; T = 1 would have given 0.036 on the held-out set.
Files and code
- Demokarar.ufakzeka.com/en
- Hugging Faceufakai/ufakzeka-karar
- GitHubufakai/ufakzeka-karar
- Research postMethod, measurements and limits
- BenchmarkHakemBench