{"slug": "beyond-routine-compliance-cunning-data-cultivates-safety-vigilance-in-large", "title": "Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models", "summary": "A new arXiv paper (2609.18515v1) introduces \"cunning questions\" — prompts with misleading premises, atypical reasoning, or subtle inconsistencies — as training data that improved large language model robustness to out-of-distribution jailbreak attacks and strengthened subsequent safety fine-tuning. Augmenting an existing state-of-the-art safety alignment pipeline with Cunning training reduced mean attack success rate across nine backbone-benchmark combinations from 17.40% to 15.05%, establishing a new state of the art in the evaluated settings. Trace analysis after matched safety fine-tuning indicated safety judgments more often govern responses before harmful planning begins, and a conditional theoretical analysis characterizes when invariance learned from cunning data transfers to safety-related inputs.", "body_md": "arXiv:2609.18515v1 Announce Type: new \nAbstract: Safety alignment teaches large language models (LLMs) to recognize harmful requests and reject risky instructions. Yet aligned models can fail when harmful intent is concealed within seemingly benign contexts. Robust safety therefore requires both knowledge of safety boundaries and \\textbf{vigilance}: the ability to detect unusual premises, misleading reasoning, and latent risks beneath surface-level semantics. Vigilance requires models to scrutinize a request's underlying intent and assumptions before acting. To cultivate this capability, we introduce \\textbf{cunning questions}, which are not necessarily safety-related but contain misleading premises, atypical reasoning, or subtle inconsistencies. We hypothesize that learning to look beyond such reasoning traps can transfer to safety-critical scenarios. Experiments show that Cunning training improves robustness to out-of-distribution jailbreak attacks and strengthens subsequent safety fine-tuning. Furthermore, augmenting an existing state-of-the-art safety alignment pipeline with Cunning establishes a new state of the art across our evaluated settings, reducing mean ASR across nine backbone--benchmark combinations from 17.40\\% to 15.05\\%. Trace analysis after matched safety fine-tuning suggests that safety judgments are more likely to govern responses before harmful planning begins. A conditional theoretical analysis further characterizes when invariance learned from cunning data can transfer to safety-related inputs. These findings suggest that cunning data can strengthen model vigilance and complement conventional safety alignment.", "url": "https://wpnews.pro/news/beyond-routine-compliance-cunning-data-cultivates-safety-vigilance-in-large", "canonical_source": "https://www.machinebrief.com/news/beyond-routine-compliance-cunning-data-cultivates-safety-vig-7f1h", "published_at": "2026-09-17 04:00:00+00:00", "updated_at": "2026-09-17 06:25:00.493602+00:00", "lang": "en", "topics": ["ai-safety", "large-language-models", "ai-research", "machine-learning"], "entities": ["arXiv", "Cunning questions"], "alternates": {"html": "https://wpnews.pro/news/beyond-routine-compliance-cunning-data-cultivates-safety-vigilance-in-large", "markdown": "https://wpnews.pro/news/beyond-routine-compliance-cunning-data-cultivates-safety-vigilance-in-large.md", "text": "https://wpnews.pro/news/beyond-routine-compliance-cunning-data-cultivates-safety-vigilance-in-large.txt", "jsonld": "https://wpnews.pro/news/beyond-routine-compliance-cunning-data-cultivates-safety-vigilance-in-large.jsonld"}}