22:10
2026-07-20
lesswrong.com
ai-safety
The AI Safety Illusion: Why Current Safety Datasets Fool Us on Model Safety
A new paper from researchers systematically evaluating widely used AI safety datasets AdvBench and HarmBench finds that these benchmarks contain contrived data points with excessive triggering cues anβ¦