04:00
2026-10-08
arxiv.org
ai-safety
Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs
A new arXiv paper (2610.09033v1) introduces the Adversarial Surface-Form Robustness Dataset (ASRD), 2,100 prompts across seven surface-form families, and evaluates five open-weight language models to …