{"slug": "wiring-beats-blending-what-transfers-between-transformer-sizes-and-what-doesn-t", "title": "Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't", "summary": "A new arXiv preprint (2608.02829) finds that converting a pretrained 1.4B-parameter Pythia model into a 410M-parameter sibling works best through initialization transfer, not weight projection. The authors show that dense weight projection is functionally destructive, while least-squares compensation and variance-preserving rescale provide token-efficient gains at low budgets, beating subcloning variants (e.g., 84.0 vs. 89.7 perplexity on a width-reduced pair) but converging to parity at larger budgets. The study also identifies over-correction at larger donor scales (6.9B->1.4B) due to ill-conditioning, suggesting dimension-aware regularization.", "body_md": "arXiv:2608.02829v1 Announce Type: new\nAbstract: Model families train every size from scratch. Can a pretrained large model be converted into a smaller sibling? We characterize the 1.4B->410M conversion in the Pythia family end-to-end: (i) representations align strongly across sizes (ridge R^2=0.84) while parameters align weakly; (ii) dense weight projection is functionally destructive -- provably not an assembly artifact -- because basis mixing breaks rotary, per-head, GELU, and LayerNorm structure; (iii) after the best-fit linear operator, weight residuals are statistically indistinguishable from noise under shuffle controls; (iv) conversion value therefore lives in initialization. In matched-budget continued pre-training we decompose conversion into two independent levers -- least-squares compensation (function: best zero-shot) and variance-preserving rescale (dynamics: best endpoints). Compensation is a token-efficient, low-budget win rather than a universal one: at 30M tokens it beats the strongest subcloning variant on both a width-reduced pair (84.0 +/- 1.8 vs. 89.7 +/- 3.7, 3/3 seeds) and a held-out depth-reduced pair (109.3 vs. 117.9, 3/3 seeds), reaching a given quality with fewer tokens; at a 33x larger budget the two converge to parity (40.0 vs. 40.0), both far ahead of from-scratch, which transfer initialization always beats -- by up to 18x at low budget, the margin narrowing at convergence and at the largest scale. We further map the method's boundary: at ~5x the donor scale (6.9B->1.4B) stacking both levers over-corrects, which we trace to ill-conditioning of the compensation solve at large width, pointing to dimension-aware regularization as the fix. Code, checkpoints, and the frozen evaluation corpus are released.", "url": "https://wpnews.pro/news/wiring-beats-blending-what-transfers-between-transformer-sizes-and-what-doesn-t", "canonical_source": "https://arxiv.org/abs/2608.02829", "published_at": "2026-08-05 04:00:00+00:00", "updated_at": "2026-08-05 04:02:41.459113+00:00", "lang": "en", "topics": ["machine-learning", "artificial-intelligence"], "entities": ["arXiv", "Pythia"], "alternates": {"html": "https://wpnews.pro/news/wiring-beats-blending-what-transfers-between-transformer-sizes-and-what-doesn-t", "markdown": "https://wpnews.pro/news/wiring-beats-blending-what-transfers-between-transformer-sizes-and-what-doesn-t.md", "text": "https://wpnews.pro/news/wiring-beats-blending-what-transfers-between-transformer-sizes-and-what-doesn-t.txt", "jsonld": "https://wpnews.pro/news/wiring-beats-blending-what-transfers-between-transformer-sizes-and-what-doesn-t.jsonld"}}