03:07
2026-08-12
arxiv.org
ai-safety
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful
A new arXiv paper by researchers auditing internal safety scores finds that these scores anti-rank successful jailbreaks, with harmful generation rising from 0.05 to 0.27 on Llama while harmful intentβ¦