Turkish language model from ufak AI

ufakzeka-1

A 151M parameter Turkish language model trained from scratch. Open weights, open data recipe, Apache-2.0.

At a glance
  • 151Mparameters
  • 13.5Bpretraining tokens
  • < $300total cost
  • Apache-2.0license
Samples5real exchanges, as they came

Example conversations

ufakzeka-1
merhaba ben Elif
Memnun oldum, Elif!
Kerem adında bir çocuk hakkında üç cümlelik bir masal anlat
Kerem, köy evinin bahçesinde oynarken çiçeklerin arasına saklandı. Biraz sonra gözleri görünmüyordu. Yoktu! Kerem çok üzüldü. Yapraklara baktı, toprağa baktı, ama bir türlü bulamadı. Sonra küçük bir karınca gördü. Karınca bir şeyler arıyordu. Kerem merakla karıncayı takip etti. Karınca büyük bir yaprağın altında kalmış! Uçamıyordu çünkü ağırlaşmış. Kerem üzüldü. Karınca, 'Yardım edin,' dedi. Kerem önce korktu, ama sonra cesaretini topladı. Yavaşça karıncayı avucuna aldı. Yaprağı dikkatlice karıncanın üstüne koydu. Karınca havalanıp uçmayı denedi. Çok mutlu oldu! 'Teşekkürler Kerem!' dedi. Kerem ve karınca bahçede birlikte oynadılar. Çiçeklerin yanından koştular, kelebekleri takip ettiler. Güneş batarken Kerem annesine koştu. 'Bahçedeki maceramı anlatmıştım,' dedi. Annesi onu kucakladı. Kerem o gün öğrendi ki, korktuğunda bile yardım istemekten çekinmemek gerekir. Ve bir minik canlının, büyük bir dostluk gösterebileceğini gördü. Artık her akşam bahçede bugünü birlikte kutluyorlar.
benim adım ne
Elif. Sohbetin başında söylemiştin.

Turkish model outputs shown as produced; temperature 0, no edits.

Evaluation7tasks, one harness, 5 models

Measured in one harness

Our model and the baselines were measured with the same harness. The five models sit well above random guessing on four of the seven tasks and close to it on the other three. On ARC-e and XCOPA, two of those four, ufakzeka-1 is less than three points behind the Turkish models five times its size; on TurBLiMP, which measures grammar, it is 5 to 8 points behind.

ufakzeka-1151MQwen2.5-0.5B494MKanarya750Mturkish-gpt2-large774M

Stones are proportional to parameter count.

ModelParamsHellaSwagARC-cARC-eXCOPABelebeleTurBLiMPTurkishMMLU
ufakzeka-1151M33.227.638.859.827.490.123.3
ufakzeka-1-base151M35.327.639.160.027.692.419.2
Kanarya-750M750M37.826.541.461.223.495.316.2
turkish-gpt2-large774M35.325.440.460.622.498.219.0
Qwen2.5-0.5B494M29.223.628.654.629.970.018.2
Random guess25252550255020

Every model measured with the same harness and no worked examples; the score is the share of questions where the model finds the correct option most likely. HellaSwag and ARC are length-corrected accuracy, the rest raw. TurBLiMP averages 16 subsets, TurkishMMLU 9 subjects. Details in the model card.

Cost$286total spend

Where the money went

Pretraining (GPU)13.5B token · H100
$66
Fine-tuning and evaluation (GPU, Colab)SFT · sweeps · gates · L4 / L40S / H100 · Colab
$183
Data generation and judge model (API)LLM API
$37
Total
$286

The total and the API line are from the bills; pretraining is its run hours at list price, and fine-tuning and evaluation are the rest. Pretraining ran in August, fine-tuning and every measurement in September. The full ledger is in the repository.

Limits5measured limits, also in the card

What it cannot do

  • It invents facts it does not have. Ask the distance from Ankara to İstanbul and it answers with a number rather than saying it does not know. It only declines the classes it was taught: time, date, weather, news, prices, personal details, the future. Quantities in recipes are unreliable too.
  • No code. Ask for Python and it writes prose.
  • Arithmetic one question at a time, working shown. Ask it to check the same sum and it can repeat the working correctly and write a wrong number on the result line. One calculation per message.
  • After a long story it can confuse who is who: in a 200-conversation run it mixed the user’s name up with a character’s, or kept an earlier name, in 38 conversations.
  • It knows nothing after summer 2026.
Usage196 MBq8_0, runs on an ordinary laptop

Run it

The architecture is plain Qwen3; transformers needs no custom code. The GGUF files need a small tokenizer patch applied to llama.cpp; details in the docs.

llama-cli -m ufakzeka-1-q8_0.gguf --temp 0.3 --top-k 40 --top-p 0.9 --repeat-penalty 1.0 --dry-multiplier 0 -n 520