cd /news/large-language-models/ufakzeka-1-building-and-evaluating-a… · home topics large-language-models article
[ARTICLE · art-137805] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

ufakzeka-1: Building and Evaluating a 151M-Parameter Turkish Language Model from Scratch

Researchers released ufakzeka-1, a 151M-parameter (182M with embeddings) decoder-only Turkish language model pretrained from scratch on 13.5B tokens of openly licensed text and instruction-tuned for chat at a total cost of about $286 in cloud GPU, API and notebook time, according to the arXiv paper 2609.25081v1. The paper reports three transferable findings: a safety gate "fixed" with training data written from its own questions read 64/64 while the honest figure was 34/64, training-seed variance matched the spread across every recipe tried, and data rounds repaired only what was absent from the data while identity tracking over long context and multi-turn arithmetic did not move. Weights, the data recipe, evaluation code and spend ledger are released under Apache-2.0.

by read1 min views1 publishedSep 23, 2026

arXiv:2609.25081v1 Announce Type: new Abstract: We describe ufakzeka-1, a 151M-parameter (182M with embeddings) decoder-only Turkish language model pretrained from scratch on 13.5B tokens of openly licensed text and instruction-tuned for chat, at a total cost of about $286 in cloud GPU, API and notebook time. The contribution is not the model's capability, which is what a model this size can be expected to have, but the record of building and measuring it: a Turkish byte-level tokenizer at 1.77 tokens per word, a three-stage pretraining schedule, a post-training mixture of openly licensed and generated data, and an evaluation battery of release gates, a rule-checked sweep of 5,508 conversations, judged conversations and hand tests, all with prompts held out from the training data, enforced by decontamination inside the data build and by a checked-in invariant script we run before each build. We report three findings that we believe transfer to other small-model efforts: a safety gate that had been "fixed" with training data written from its own questions read 64/64 while the honest figure was 34/64; training-seed variance was as large as the spread across every recipe we tried, so single-seed comparisons at this scale are uninformative; and data rounds repaired only what was absent from the data, while identity tracking over long context and multi-turn arithmetic did not move across any data change we tried, which we read as limits of the model size rather than gaps in the data, a reading the next, larger model will test. Weights, the data recipe, the evaluation code and the spend ledger are released under Apache-2.0.

── more in #large-language-models 4 stories · sorted by recency
── more on @ufakzeka-1 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ufakzeka-1-building-…] indexed:0 read:1min 2026-09-23 ·