cd/entity/HarmBench· home› entities› HarmBench
grep -l @harmbench /news/*.json | wc -l → 9

HarmBench

mentions 9 type Organization feed RSS

// recent coverage 9 mentions

12:00
2026-10-01
gizmodo.com
ai-safety

AI’s ‘Abliteration’ Problem Is Bigger Than China

Anthropic researchers reported on Tuesday that attackers can bypass the safeguards of GLM-5.3, an open-weight model released last month by Chinese AI lab Z.ai, between 64% and 100% of the time using s…

09:08
2026-08-06
aiunderstanding.org
ai-safety

Study Finds AI Safety Benchmarks Can Use Far Fewer Tests

A research paper posted on August 5 by two independent researchers and two researchers affiliated with the UK AI Security Institute found that AI safety benchmarks can be compressed to as few as 10-25…

16:58
2026-07-16
lesswrong.com
ai-safety

Jailbreak Patching with SOO-Style Conceptual Fusion

A new jailbreak patching method using Self-Other Overlap (SOO) conceptual fusion, described by Carauleanu et al. (2025), successfully reduces jailbreak evasion in Qwen 2.5 1.5b by fusing the model's a…

// co-occurs with top 8 entities
// topics top 6 topics