cd /news/large-language-models/looped-gpt-bert-trading-parameters-f… · home topics large-language-models article
[ARTICLE · art-125382] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Looped GPT-BERT: Trading Parameters for Computation in Small Language Modeling

Researchers studying Looped GPT-BERT in the BabyLM 2026 Strict-small setting trained a 12.18M-parameter model on a preprocessed 7.48M-word English corpus, using four physical layers for twelve recurrent traversals, and the BabyLM 2026 leaderboard reports an Overall Average of 35.42 and an NLP Average of 48.48. The looped model achieved comparable performance to public BabyLM 10M Strict-small GPT-2 and GPT-BERT baselines on selected linguistic and downstream metrics including BLiMP and GLUE with fewer parameters. Loop ablations showed additional recurrent computation can improve training and preserve strong performance on selected linguistic tasks, while poorer results on other tasks may reveal a limitation of the looped design, since using only a few physical layers restricts the model's representational space.

by read1 min views2 publishedSep 10, 2026

arXiv:2609.09691v1 Announce Type: new Abstract: When training data are limited, increasing parameter count is not the only way to improve language-model performance. A small parameter set, when repeatedly applied, can also deliver comparable performance. We study Looped GPT-BERT in the BabyLM 2026 Strict-small setting, combining GPT-BERT's masked next-token and causal language-modeling objectives with depth-wise parameter sharing. We train on a preprocessed 7.48M-word English corpus and compare objective ratios, non-looped and looped architectures, and loop counts. Our final $4\times12$ model uses four physical layers for twelve recurrent traversals and contains 12.18M parameters. The BabyLM 2026 leaderboard reports an Overall Average of 35.42 and an NLP Average of 48.48. Compared with public BabyLM 10M Strict-small GPT-2 and GPT-BERT baselines, it achieves comparable performance on selected linguistic and downstream metrics, including BLiMP and GLUE, with fewer parameters. The loop ablations show that additional recurrent computation can improve training and preserve strong performance on selected linguistic tasks, whereas poorer performance on other tasks may reveal an inherent limitation of the looped design: using only a few physical layers restricts the model's representational space.

── more in #large-language-models 4 stories · sorted by recency
── more on @looped gpt-bert 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/looped-gpt-bert-trad…] indexed:0 read:1min 2026-09-10 ·