04:00
2026-08-05
machinebrief.com
artificial-intelligence
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity
A new arXiv preprint (2608.02665v1) finds that evaluating large language models on a single canonical prompt surface underestimates unsafe compliance by 3.3 to 12.9 percentage points across five modelβ¦