cd /news/natural-language-processing/feeding-babylms-macaroni-code-switch… · home › topics › natural-language-processing › article
[ARTICLE · art-140743] src=arxiv.org ↗ pub= topic=natural-language-processing verified=true sentiment=↑ positive

Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence

Pretraining small decoder-only transformers on code-switched text aligns representations of parallel text across English, Dutch, and Chinese and improves scores on the BabyLM evaluation suite, according to an arXiv paper (2609.30535v1) from researchers who released code, data, and models at github.com/drooryck/multilingual-macaroni. The study built two 100M-word corpora — a base mix of the English, Dutch, and Chinese BabyBabelLM datasets and an LLM-generated version with word- and sentence-level code-switching — and found the alignment persisted through subsequent monolingual training. Under a curriculum progressing from word-level code-switching to sentence-level code-switching to monolingual documents, models trained on code-switched data outperformed baselines trained without it.

by read1 min views1 publishedSep 28, 2026

arXiv:2609.30535v1 Announce Type: new Abstract: Children in multilingual communities often code-switch, using multiple languages in a single utterance. Can we induce cross-lingual alignment in language models by training on code-switched text? We pretrain small decoder-only transformers on two 100M-word multilingual corpora: a base corpus formed by mixing the English, Dutch, and Chinese BabyBabelLM datasets, and a corpus generated from it by inserting word- and sentence-level code-switching using an LLM. We find that training on code-switched data aligns the representations of parallel text, particularly across different scripts, and that this alignment persists through training on monolingual documents. Under a learning curriculum that progresses from word-level code-switching, to sentence-level code-switching, to monolingual documents, models trained on code-switched data outperform baselines trained without it on the BabyLM evaluation suite. Our work characterizes code-switching curriculum learning as an effective data augmentation method for multilingual pretraining. We release our code, data, and models at https://github.com/drooryck/multilingual-macaroni.

── more in #natural-language-processing 4 stories · sorted by recency
── more on @babybabellm 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/feeding-babylms-maca…] indexed:0 read:1min 2026-09-28 · —