04:00
2026-09-18
arxiv.org
ai-safety
The Role of Fine-grained Harm Signals in LLM Safety
A new arXiv paper (2609.19366v1) reports that category-specific harm representations in large language models carry safety-relevant information beyond a shared general harmfulness direction. Using act…