cd /news/ai-safety/auditing-system-1-models-on-biosecur… · home › topics › ai-safety › article
[ARTICLE · art-141478] src=machinebrief.com ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Auditing System-1 Models on Biosecurity-Relevant Benchmarks: Calibration, Selective Prediction, and Permutation Instability in a Non-Generative Model

An audit of one commercial non-generative "System-1" model on 6,020 multiple-choice biosecurity-relevant items found pooled expected calibration error of 0.034 and pooled AUROC of 0.820 once the vendor's uncertainty field was correctly interpreted, but 37.4% of WMDP-Cyber items changed answers under four cyclic rotations of the answer options. The study, posted as arXiv:2609.30454v1, drew items from the Weapons of Mass Destruction Proxy (WMDP), a paraphrase-robust WMDP-Bio variant, and six LAB-Bench subtasks, and attributed most of the answer instability to option order rather than run-to-run variation using byte-identical repeated calls as a control. Averaging probabilities across rotations improved WMDP-Cyber accuracy by 3.8 percentage points, and applying that averaging only to low-confidence items recovered most of the gain at well under the cost of averaging every item.

by read1 min views1 publishedSep 29, 2026

arXiv:2609.30454v1 Announce Type: new Abstract: Non-generative "System-1" models return structured probabilistic decisions in a single forward pass, without autoregressive decoding, at a small fraction of the inference cost of a generative model. This makes them of interest as inexpensive components in larger pipelines, but their reliability on biosecurity-relevant tasks has not been systematically examined. We audit one commercial System-1 model on 6,020 multiple-choice items drawn from the Weapons of Mass Destruction Proxy (WMDP), a paraphrase-robust WMDP-Bio variant, and six LAB-Bench subtasks, measuring accuracy, calibration, error detection, selective prediction, and sensitivity to the order in which answer options are presented. Accuracy is strongly task-dependent. Once the vendor's uncertainty field is correctly interpreted, the model is reasonably well calibrated (pooled expected calibration error 0.034) and its top-1 probability separates correct from incorrect predictions (pooled AUROC 0.820), though both degrade substantially on the weaker tasks. Under four cyclic rotations of the answer options, 37.4% of WMDP-Cyber items receive different answers; a control using byte-identical repeated calls attributes most of this to option order rather than run-to-run variation. Averaging probabilities across rotations improves WMDP-Cyber accuracy by 3.8 percentage points, and applying it only to low-confidence items recovers most of that gain at well under the cost of averaging every item.

── more in #ai-safety 4 stories · sorted by recency
── more on @wmdp 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/auditing-system-1-mo…] indexed:0 read:1min 2026-09-29 · —