{"slug": "safety-for-whom-boundary-aware-self-distillation-for-controlled-llm-safety", "title": "Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal", "summary": "A new arXiv study (2609.04482v1) introduces a boundary-aware self-distillation framework for controlled LLM safety refusal, showing that escalating retries reduce prompts without accepted refusal traces from 19.88% to 0.20%. On political persuasion with Qwen3-8B, target-domain refusal increased from 9.47% to 84.75% and mean unsafe-response rate across three harmfulness benchmarks dropped from 26.26% to 0.14%, but XSTest over-refusal rose from 2.00% to 74.00%. The authors argue that data composition controls the safety-usability trade-off and that safety alignment should be evaluated on both sides of the intended refusal boundary.", "body_md": "arXiv:2609.04482v1 Announce Type: new \nAbstract: Safety alignment is usually posed as a topic-level question: is this subject harmful? Deployments ask a narrower one. A civics tutor and a public-sector assistant may share a base model yet need different boundaries inside the same topic, refusing targeted political manipulation while still answering factual questions about the same election. We formulate this as narrow-boundary safety and introduce an offline self-generated framework combining controlled topic generation, coverage repair, in-distribution compensation data, and harmful-benign pairs for training and evaluation. Single-shot generation leaves 19.88% of prompts without accepted refusal traces, whereas escalating retries leave 0.20%. On political persuasion with Qwen3-8B, training on refusal data completed through Escalate increases target-domain refusal from 9.47% to 84.75% and reduces the mean unsafe-response rate across three broader harmfulness benchmarks from 26.26% to 0.14%, but increases XSTest over-refusal from 2.00% to 74.00%. In a separate matched comparison, replacing external responses with verified target-model responses reduces over-refusal from 15.20% to 5.20%. Boundary-pair data reduces comply-side over-refusal on held-out pairs from 32.94% to 4.16%, while harmful-side refusal decreases only from 91.88% to 87.72%. These results show that data composition controls the safety and usability trade-off, and that safety alignment should be evaluated on both sides of the intended refusal boundary.", "url": "https://wpnews.pro/news/safety-for-whom-boundary-aware-self-distillation-for-controlled-llm-safety", "canonical_source": "https://arxiv.org/abs/2609.04482", "published_at": "2026-09-07 04:00:00+00:00", "updated_at": "2026-09-07 04:25:41.669672+00:00", "lang": "en", "topics": ["ai-safety", "large-language-models", "artificial-intelligence"], "entities": ["arXiv", "Qwen3-8B", "XSTest"], "alternates": {"html": "https://wpnews.pro/news/safety-for-whom-boundary-aware-self-distillation-for-controlled-llm-safety", "markdown": "https://wpnews.pro/news/safety-for-whom-boundary-aware-self-distillation-for-controlled-llm-safety.md", "text": "https://wpnews.pro/news/safety-for-whom-boundary-aware-self-distillation-for-controlled-llm-safety.txt", "jsonld": "https://wpnews.pro/news/safety-for-whom-boundary-aware-self-distillation-for-controlled-llm-safety.jsonld"}}