14:23
2026-09-08
huggingface.co
ai-safety
Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic
A new paper, 'Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal,' argues that topic-level safety taxonomies like LlamaGuard-3 are too blunt for real deployments, whicβ¦