Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic
A new paper, 'Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal,' argues that topic-level safety taxonomies like LlamaGuard-3 are too blunt for real deployments, which need to refuse only a harmful subset of a topic (e.g., political manipulation) while answering benign prompts (e.g., factual election questions). The authors propose boundary-aware self-distillation…