{"slug": "the-role-of-fine-grained-harm-signals-in-llm-safety", "title": "The Role of Fine-grained Harm Signals in LLM Safety", "summary": "A new arXiv paper (2609.19366v1) reports that category-specific harm representations in large language models carry safety-relevant information beyond a shared general harmfulness direction. Using activation steering with category residuals across 11 risk categories in 3 instruction-tuned LLMs, the researchers found that whether category residuals encode harmfulness varies by category and follows a similar pattern across models, while whether they induce refusal is more model-dependent. The authors conclude that fine-grained category residuals, not just shared general harmfulness representation, must be considered to fully understand LLM safety.", "body_md": "arXiv:2609.19366v1 Announce Type: new \nAbstract: Prior work has shown that internal harmfulness representations in large language models vary across risk categories, while sharing a common general harm representation component. This raises a question about the role of the category-specific component beyond general harm representation in LLM safety. To answer this question, we isolate the category-specific component by removing shared general harmfulness representation from each categorical harmfulness representation, yielding a category residual that is orthogonal to general harmfulness at every layer. Using activation steering with category residuals across 11 risk categories in 3 instruction-tuned LLMs, we find that whether category residuals encode harmfulness varies across categories, and that this category-wise pattern is similar across models. Whether category residuals induce refusal also varies across categories, but this category-wise pattern is more model-dependent. We also find that category residuals increase LLMs' downstream internal alignment with shared general harmfulness representation. Together, these findings demonstrate that more fine-grained category residuals should also be considered beyond shared general harmfulness representation to fully understand LLM safety. More broadly, our findings show that even a direction orthogonal to a concept at one layer can contribute to the concept's downstream amplification.", "url": "https://wpnews.pro/news/the-role-of-fine-grained-harm-signals-in-llm-safety", "canonical_source": "https://arxiv.org/abs/2609.19366", "published_at": "2026-09-18 04:00:00+00:00", "updated_at": "2026-09-18 04:26:21.397777+00:00", "lang": "en", "topics": ["ai-safety", "large-language-models", "ai-research", "artificial-intelligence"], "entities": ["arXiv", "large language models", "activation steering"], "alternates": {"html": "https://wpnews.pro/news/the-role-of-fine-grained-harm-signals-in-llm-safety", "markdown": "https://wpnews.pro/news/the-role-of-fine-grained-harm-signals-in-llm-safety.md", "text": "https://wpnews.pro/news/the-role-of-fine-grained-harm-signals-in-llm-safety.txt", "jsonld": "https://wpnews.pro/news/the-role-of-fine-grained-harm-signals-in-llm-safety.jsonld"}}