{"slug": "bayesian-wind-tunnels-for-model-selection", "title": "Bayesian Wind Tunnels for Model Selection", "summary": "A 2.8M-parameter transformer achieves 0.01-bit entropy agreement with the Bayesian optimum for model selection in Bayesian wind tunnels using fixed-point-free involutions, but fails completely with opaque symbols when arithmetic is required, a boundary that persists under 112x scaling to 316M parameters. The study, published on arXiv (2607.19379v1), introduces controlled environments where ground-truth posteriors over hypothesis classes are available in closed form, and finds that frontier LLMs show qualitative Bayesian behavior but a large calibration gap of ~55x.", "body_md": "arXiv:2607.19379v1 Announce Type: new\nAbstract: Prior work has shown that transformers can perform exact Bayesian filtering within a fixed\nhypothesis class. Can they also perform Bayesian model selection -- identifying the correct\nhypothesis class from data? We introduce model-selection Bayesian wind tunnels: controlled\nenvironments where ground-truth posteriors over hypothesis classes are available in closed\nform. Using fixed-point-free involutions -- whose defining property f(f(x))=x is purely\nrelational -- a 2.8M-parameter transformer achieves 0.01-bit entropy agreement with the\nBayesian optimum (3 seeds), with both integer tokens and opaque symbols whose meanings\nchange every episode. This extends to non-nested comparisons: involutions vs. 3-cycles\n(where neither class is a subset of the other) achieve class-posterior MAE under 0.001,\ndemonstrating genuine model selection beyond simplicity/subset bias. We then identify a\nsharp perceptual access condition: when the discriminative statistic requires arithmetic --\nmodular addition (rotations) or multiplication (f(x)=cx mod p) -- model selection succeeds\nwith integer tokens but fails completely with opaque symbols, and this boundary persists\nunder 112x scaling (2.8M to 316M parameters). A stationarity control confirms the operative\nfactor: opaque tokens with a fixed relabeling succeed (0.009-bit MAE), showing that stable\nsemantics, not integer identity, enable circuit compilation. Header subtask diagnostics\nlocalize the failure to the composition of header inversion with arithmetic rather than\nheader parsing itself. Probing frontier LLMs on the same tasks shows qualitative Bayesian\nbehavior but a large calibration gap (~55x), measured through lossy probes and therefore\ndirectional rather than exact.", "url": "https://wpnews.pro/news/bayesian-wind-tunnels-for-model-selection", "canonical_source": "https://arxiv.org/abs/2607.19379", "published_at": "2026-07-23 04:00:00+00:00", "updated_at": "2026-07-23 04:12:20.718274+00:00", "lang": "en", "topics": ["machine-learning", "artificial-intelligence", "large-language-models"], "entities": ["arXiv", "Bayesian wind tunnels", "fixed-point-free involutions", "LLMs"], "alternates": {"html": "https://wpnews.pro/news/bayesian-wind-tunnels-for-model-selection", "markdown": "https://wpnews.pro/news/bayesian-wind-tunnels-for-model-selection.md", "text": "https://wpnews.pro/news/bayesian-wind-tunnels-for-model-selection.txt", "jsonld": "https://wpnews.pro/news/bayesian-wind-tunnels-for-model-selection.jsonld"}}