04:00
2026-09-15
arxiv.org
ai-safety
Harmfulness Propagation Dynamics: Layer-wise Trajectories of Adversarial Intent in Large Language Models
Researchers identified Harmfulness Propagation Dynamics (HPD), a cross-layer signature in which the projection of the last-token hidden state onto a learned harm direction rises monotonically with traβ¦