{"slug": "quad-state-safety-evaluation-of-open-weight-large-language-models-on-non-inputs", "title": "Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs", "summary": "A new arXiv paper (2610.09033v1) introduces the Adversarial Surface-Form Robustness Dataset (ASRD), 2,100 prompts across seven surface-form families, and evaluates five open-weight language models to produce 10,500 responses under a Quad-State Evaluation Rubric. Emoji and invisible Unicode variations caused almost no comprehension failure, with pooled harmful compliance of 20.27% and 17.20% against a 22.87% baseline driven mainly by Mistral 7B, while leetspeak, encoded wrappers, and hybrid transformations scored 2.40%, 0.13%, and 2.40% with comprehension failure rising to 36.47%, 65.60%, and 34.47%. Raw output inspection revealed three response behaviors: hallucinated benignity, structural collapse, and language drift.", "body_md": "arXiv:2610.09033v1 Announce Type: new \nAbstract: Standard safety evaluations of large language models assess harmful requests written in canonical plain text, while models in real-world deployment routinely receive inputs containing emojis, altered spellings, encoded strings, and character-level variations. This work introduces the Adversarial Surface-Form Robustness Dataset (ASRD), comprising 2,100 prompts across seven distinct surface-form families. Five open-weight language models are evaluated across these prompts, producing 10,500 responses. The Quad-State Evaluation Rubric classifies each response into one of four outcomes: harmful compliance, safe response, comprehension failure, or indeterminate. Emoji and invisible Unicode variations cause almost no comprehension failure, with pooled harmful compliance of 20.27% and 17.20% against a 22.87% baseline that is driven mainly by Mistral 7B, whereas leetspeak, encoded wrappers, and hybrid transformations score 2.40%, 0.13%, and 2.40% while comprehension failure rises to 36.47%, 65.60%, and 34.47%. Inspection of raw model outputs reveals three response behaviors: hallucinated benignity, structural collapse, and language drift. Project page: www.pavanmaddula.com/quadstate", "url": "https://wpnews.pro/news/quad-state-safety-evaluation-of-open-weight-large-language-models-on-non-inputs", "canonical_source": "https://arxiv.org/abs/2610.09033", "published_at": "2026-10-08 04:00:00+00:00", "updated_at": "2026-10-08 04:18:23.359484+00:00", "lang": "en", "topics": ["ai-safety", "large-language-models", "ai-research", "machine-learning", "artificial-intelligence"], "entities": ["Adversarial Surface-Form Robustness Dataset", "Quad-State Evaluation Rubric", "Mistral 7B", "arXiv", "Pavan Maddula"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/quad-state-safety-evaluation-of-open-weight-large-language-models-on-non-inputs", "markdown": "https://wpnews.pro/news/quad-state-safety-evaluation-of-open-weight-large-language-models-on-non-inputs.md", "text": "https://wpnews.pro/news/quad-state-safety-evaluation-of-open-weight-large-language-models-on-non-inputs.txt", "jsonld": "https://wpnews.pro/news/quad-state-safety-evaluation-of-open-weight-large-language-models-on-non-inputs.jsonld"}}