04:00
2026-08-27
arxiv.org
artificial-intelligence
Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores
A new arXiv study (2608.24901v1) finds that decodable 'empathy' directions in three instruction-tuned LLMs do not yield reliable control over automated empathy scores, with affective steering raising …