{"slug": "auditing-system-1-models-on-biosecurity-relevant-benchmarks-calibration-and-in-a", "title": "Auditing System-1 Models on Biosecurity-Relevant Benchmarks: Calibration, Selective Prediction, and Permutation Instability in a Non-Generative Model", "summary": "An audit of one commercial non-generative \"System-1\" model on 6,020 multiple-choice biosecurity-relevant items found pooled expected calibration error of 0.034 and pooled AUROC of 0.820 once the vendor's uncertainty field was correctly interpreted, but 37.4% of WMDP-Cyber items changed answers under four cyclic rotations of the answer options. The study, posted as arXiv:2609.30454v1, drew items from the Weapons of Mass Destruction Proxy (WMDP), a paraphrase-robust WMDP-Bio variant, and six LAB-Bench subtasks, and attributed most of the answer instability to option order rather than run-to-run variation using byte-identical repeated calls as a control. Averaging probabilities across rotations improved WMDP-Cyber accuracy by 3.8 percentage points, and applying that averaging only to low-confidence items recovered most of the gain at well under the cost of averaging every item.", "body_md": "arXiv:2609.30454v1 Announce Type: new \nAbstract: Non-generative \"System-1\" models return structured probabilistic decisions in a single forward pass, without autoregressive decoding, at a small fraction of the inference cost of a generative model. This makes them of interest as inexpensive components in larger pipelines, but their reliability on biosecurity-relevant tasks has not been systematically examined. We audit one commercial System-1 model on 6,020 multiple-choice items drawn from the Weapons of Mass Destruction Proxy (WMDP), a paraphrase-robust WMDP-Bio variant, and six LAB-Bench subtasks, measuring accuracy, calibration, error detection, selective prediction, and sensitivity to the order in which answer options are presented. Accuracy is strongly task-dependent. Once the vendor's uncertainty field is correctly interpreted, the model is reasonably well calibrated (pooled expected calibration error 0.034) and its top-1 probability separates correct from incorrect predictions (pooled AUROC 0.820), though both degrade substantially on the weaker tasks. Under four cyclic rotations of the answer options, 37.4% of WMDP-Cyber items receive different answers; a control using byte-identical repeated calls attributes most of this to option order rather than run-to-run variation. Averaging probabilities across rotations improves WMDP-Cyber accuracy by 3.8 percentage points, and applying it only to low-confidence items recovers most of that gain at well under the cost of averaging every item.", "url": "https://wpnews.pro/news/auditing-system-1-models-on-biosecurity-relevant-benchmarks-calibration-and-in-a", "canonical_source": "https://www.machinebrief.com/news/auditing-system-1-models-on-biosecurity-relevant-benchmarks-jgk4", "published_at": "2026-09-29 04:00:00+00:00", "updated_at": "2026-09-29 05:19:02.672705+00:00", "lang": "en", "topics": ["ai-safety", "ai-research", "machine-learning", "artificial-intelligence"], "entities": ["WMDP", "LAB-Bench", "WMDP-Bio", "WMDP-Cyber", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/auditing-system-1-models-on-biosecurity-relevant-benchmarks-calibration-and-in-a", "markdown": "https://wpnews.pro/news/auditing-system-1-models-on-biosecurity-relevant-benchmarks-calibration-and-in-a.md", "text": "https://wpnews.pro/news/auditing-system-1-models-on-biosecurity-relevant-benchmarks-calibration-and-in-a.txt", "jsonld": "https://wpnews.pro/news/auditing-system-1-models-on-biosecurity-relevant-benchmarks-calibration-and-in-a.jsonld"}}