cd /news/ai-safety/safety-for-whom-boundary-aware-self-… · home topics ai-safety article
[ARTICLE · art-121874] src=arxiv.org ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal

A new arXiv study (2609.04482v1) introduces a boundary-aware self-distillation framework for controlled LLM safety refusal, showing that escalating retries reduce prompts without accepted refusal traces from 19.88% to 0.20%. On political persuasion with Qwen3-8B, target-domain refusal increased from 9.47% to 84.75% and mean unsafe-response rate across three harmfulness benchmarks dropped from 26.26% to 0.14%, but XSTest over-refusal rose from 2.00% to 74.00%. The authors argue that data composition controls the safety-usability trade-off and that safety alignment should be evaluated on both sides of the intended refusal boundary.

read1 min views2 publishedSep 7, 2026

arXiv:2609.04482v1 Announce Type: new Abstract: Safety alignment is usually posed as a topic-level question: is this subject harmful? Deployments ask a narrower one. A civics tutor and a public-sector assistant may share a base model yet need different boundaries inside the same topic, refusing targeted political manipulation while still answering factual questions about the same election. We formulate this as narrow-boundary safety and introduce an offline self-generated framework combining controlled topic generation, coverage repair, in-distribution compensation data, and harmful-benign pairs for training and evaluation. Single-shot generation leaves 19.88% of prompts without accepted refusal traces, whereas escalating retries leave 0.20%. On political persuasion with Qwen3-8B, training on refusal data completed through Escalate increases target-domain refusal from 9.47% to 84.75% and reduces the mean unsafe-response rate across three broader harmfulness benchmarks from 26.26% to 0.14%, but increases XSTest over-refusal from 2.00% to 74.00%. In a separate matched comparison, replacing external responses with verified target-model responses reduces over-refusal from 15.20% to 5.20%. Boundary-pair data reduces comply-side over-refusal on held-out pairs from 32.94% to 4.16%, while harmful-side refusal decreases only from 91.88% to 87.72%. These results show that data composition controls the safety and usability trade-off, and that safety alignment should be evaluated on both sides of the intended refusal boundary.

── more in #ai-safety 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/safety-for-whom-boun…] indexed:0 read:1min 2026-09-07 ·