cd/entity/HarmbenchΒ· homeβ€Ί entitiesβ€Ί Harmbench
grep -l @harmbench /news/*.json | wc -l β†’ 1

Harmbench

mentions 1 type Organization feed RSS

// recent coverage 1 mentions

04:00
2026-07-24
arxiv.org
artificial-intelligence

Robust Critics: Defending LLMs Against Multi-Turn Attacks

Researchers propose Dialogue Critic Guided Sampling (DCGS), a framework that infers user intent at every turn of dialogue to defend large language models against multi-turn attacks. DCGS models advers…

// co-occurs with top 4 entities
// topics top 4 topics