04:00
2026-07-27
arxiv.org
large-language-models
Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study
A new study analyzing self-harm representations in large language models finds that self-harm information crystallizes in the final 3-7% of network layers across four models, with Gemma-3-4B representβ¦