cd /news/ai-safety/how-narrative-wrapping-affects-llm-r… · home › topics › ai-safety › article
[ARTICLE · art-148043] src=arxiv.org ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense

A new arXiv paper (2610.11005v1) reports that narrative or role-play wrappers let harmful requests bypass safety refusals in Qwen3-1.7B at 89.4% attack success in English, 93.0% in modern Chinese, and 95.7% in Classical Chinese. The authors built GUISE, a cross-language benchmark with parallel English, modern Chinese, and Classical Chinese requests, matched harmful and benign pairs, held-out wrapper types, and a stricter criterion counting warn-then-answer responses as successes, and found that language and register shift harmful-request representations only slightly from the refusal direction while narrative wrappers move them much farther. They propose AXIS, which combines preference optimisation with a rotation objective aligning harmful-request representations to the refusal direction and a commitment objective training full refusal, and report AXIS achieves the highest combined safety and usability score across Qwen3-1.7B, Qwen3-4B, and GLM-4-9B.

by read1 min views1 publishedOct 9, 2026

arXiv:2610.11005v1 Announce Type: new Abstract: Safety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper. We measure this vulnerability across languages and registers: attack success on Qwen3-1.7B is already 89.4% in English and 93.0% in modern Chinese, and reaches 95.7% in Classical Chinese. We build GUISE, a benchmark for systematically studying this vulnerability. It includes parallel requests in English, modern Chinese, and Classical Chinese, matched harmful and benign pairs, wrapper types held out for evaluation, and a stricter criterion that counts warn-then-answer responses as attack successes. Representation analysis shows that language and register move harmful-request representations only slightly away from the model's refusal direction, whereas narrative wrappers move them much farther away. We propose AXIS, which combines preference optimisation with a rotation objective that aligns harmful-request representations with the refusal direction and a commitment objective that trains the model to refuse completely rather than produce a warn-then-answer response. Across Qwen3-1.7B, Qwen3-4B and GLM-4-9B, AXIS achieves the highest combined safety and usability score among the compared methods.

── more in #ai-safety 4 stories · sorted by recency
── more on @qwen3-1.7b 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-narrative-wrappi…] indexed:0 read:1min 2026-10-09 · —