{"slug": "feeding-babylms-macaroni-code-switching-curricula-cause-cross-lingual", "title": "Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence", "summary": "Pretraining small decoder-only transformers on code-switched text aligns representations of parallel text across English, Dutch, and Chinese and improves scores on the BabyLM evaluation suite, according to an arXiv paper (2609.30535v1) from researchers who released code, data, and models at github.com/drooryck/multilingual-macaroni. The study built two 100M-word corpora — a base mix of the English, Dutch, and Chinese BabyBabelLM datasets and an LLM-generated version with word- and sentence-level code-switching — and found the alignment persisted through subsequent monolingual training. Under a curriculum progressing from word-level code-switching to sentence-level code-switching to monolingual documents, models trained on code-switched data outperformed baselines trained without it.", "body_md": "arXiv:2609.30535v1 Announce Type: new \nAbstract: Children in multilingual communities often code-switch, using multiple languages in a single utterance. Can we induce cross-lingual alignment in language models by training on code-switched text? We pretrain small decoder-only transformers on two 100M-word multilingual corpora: a base corpus formed by mixing the English, Dutch, and Chinese BabyBabelLM datasets, and a corpus generated from it by inserting word- and sentence-level code-switching using an LLM. We find that training on code-switched data aligns the representations of parallel text, particularly across different scripts, and that this alignment persists through training on monolingual documents. Under a learning curriculum that progresses from word-level code-switching, to sentence-level code-switching, to monolingual documents, models trained on code-switched data outperform baselines trained without it on the BabyLM evaluation suite. Our work characterizes code-switching curriculum learning as an effective data augmentation method for multilingual pretraining. We release our code, data, and models at https://github.com/drooryck/multilingual-macaroni.", "url": "https://wpnews.pro/news/feeding-babylms-macaroni-code-switching-curricula-cause-cross-lingual", "canonical_source": "https://arxiv.org/abs/2609.30535", "published_at": "2026-09-28 04:00:00+00:00", "updated_at": "2026-09-28 04:18:49.595316+00:00", "lang": "en", "topics": ["natural-language-processing", "large-language-models", "machine-learning", "ai-research"], "entities": ["BabyBabelLM", "BabyLM", "arXiv", "github.com/drooryck/multilingual-macaroni"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/feeding-babylms-macaroni-code-switching-curricula-cause-cross-lingual", "markdown": "https://wpnews.pro/news/feeding-babylms-macaroni-code-switching-curricula-cause-cross-lingual.md", "text": "https://wpnews.pro/news/feeding-babylms-macaroni-code-switching-curricula-cause-cross-lingual.txt", "jsonld": "https://wpnews.pro/news/feeding-babylms-macaroni-code-switching-curricula-cause-cross-lingual.jsonld"}}