cd /news/artificial-intelligence/moganbert-tr-a-turkish-encoder-found… · home topics artificial-intelligence article
[ARTICLE · art-112710] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM-to-MLM Curriculum

MoganAI released MoganBert-TR, a 149M-parameter Turkish encoder foundation model trained from scratch on 237.3B tokens with a two-stage CLM-to-MLM curriculum, achieving 78.41 on TrGLUE, the best among compared Turkish ModernBERT models, and 77.73 on TabiBench. The derived MoganBert-Embed ranks first among student models on MTEB(Turkish) with 68.30, reaching 99.5% of its 7.57B-parameter teacher's score with a 51x smaller backbone.

read1 min views1 publishedAug 27, 2026

arXiv:2608.25768v1 Announce Type: new Abstract: Turkish encoder models have adopted modern architectures while leaving the pretraining objective fixed at masked language modelling. This paper introduces MoganBert-TR, a 149M-parameter Turkish encoder foundation model trained from scratch on a language-specifically filtered corpus, together with an embedding model derived from it (MoganBert-Embed). MoganBert-TR is trained over 237.3B tokens with a two-stage CLM-to-MLM curriculum: causal language modelling first, masked language modelling for the remainder, with the transition made inside the stable phase of a WSD schedule. In a controlled ablation under an equal step budget, this design outperforms pure MLM by 2.7-3.7x on Turkish MS MARCO retrieval; the measured mechanism is embedding geometry, where a single direction absorbs 28.1% of the variance under pure MLM against 11.9% under the curriculum. Long-context extension and learning-rate decay are then split into two branches after a shared prefix: running the final portion of decay at 1024 context improves the TrGLUE average by 0.49 +/- 0.26 points across five paired seeds (p = 0.013) and beats a model-soup alternative by 0.75 points at ~4.3% additional cost. MoganBert-TR attains 78.41 on TrGLUE, the best among the Turkish ModernBERT models compared, and 77.73 on TabiBench, where it leads two of the eight categories with the largest margin on code retrieval (+3.62 points over TabiBERT). MoganBert-Embed, produced through teacher distillation and multi-signal contrastive fine-tuning, ranks first among student models on the MTEB(Turkish) overall average with 68.30 and reaches 99.5% of its 7.57B-parameter teacher's score with a 51x smaller backbone. The accompanying 50,048-token tokenizer outperforms all compared Turkish tokenizers on compression and fertility across two independent test sets. Weights, tokenizer, embedding model and evaluation code: https://huggingface.co/moganai

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @moganai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/moganbert-tr-a-turki…] indexed:0 read:1min 2026-08-27 ·