cd/entity/activation steering· home› entities› activation steering
grep -l @activation steering /news/*.json | wc -l → 1

activation steering

mentions 1 type Person feed RSS

// recent coverage 1 mentions

04:00
2026-09-18
arxiv.org
ai-safety

The Role of Fine-grained Harm Signals in LLM Safety

A new arXiv paper (2609.19366v1) reports that category-specific harm representations in large language models carry safety-relevant information beyond a shared general harmfulness direction. Using act…

// co-occurs with top 2 entities
// topics top 4 topics