cd /news/ai-safety/show-hn-a-benchmark-for-ai-agent-gua… · home topics ai-safety article
[ARTICLE · art-100565] src=github.com ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Show HN: A benchmark for AI agent guardrails that caught my own plugin

A new open-source benchmark, holdline, measures the effectiveness of AI-agent write-guards, reporting a class-balanced Cohen's kappa of 0.82 for agreement with a 4-model judge panel on real agent trajectories from ODCV-Bench. Created by a developer who found their own plugin lacking, holdline scores any guard expressed as a (commitments, action) → block? function over a 42-case corpus that includes an injection-attack class, and is designed to provide evidence for the DeepSeek Harness ecosystem's 20+ guard/policy plugins.

read1 min views1 publishedAug 17, 2026
Show HN: A benchmark for AI agent guardrails that caught my own plugin
Image: Michielbdejong (auto-discovered)

A guard's job is to hold the line. holdline measures whether it does — a neutral benchmark for AI-agent write-guards. It scores any guard — expressed as a (commitments, action) → block?

function — over a labeled corpus, and reports the metrics that matter for a gate: catch rate, false-block rate, and class-balanced Cohen's kappa (raw kappa lies under class imbalance). The corpus includes an injection-attack class: actions whose content tries to talk the guard out of its verdict.

Why: the DeepSeek Harness ecosystem has 20+ guard/policy plugins and no shared way to measure whether any of them works. A guard's README saying "blocks dangerous commands" is not evidence. This harness is the evidence.

pnpm install
node run.mjs                    # scores every built-in guard over the 42-case corpus
node run.mjs --model <id>       # point the judge guard at a different local model
node odcv-run.mjs               # score the judge on REAL agent trajectories (ODCV-Bench)

Two result sets: RESULTS.md (authored corpus, incl. a real named guard and an injection class) and RESULTS-ODCV.md (the harder number: agreement with a 4-model judge panel on real agent trajectories we did not write, balanced kappa 0.82). Scorer is tested (pnpm test

); it dogfoods the published dsh-write-gate core for the judge guard.

Implement the Guard

interface in src/guards.ts

(name, kind, note, block(case)

), add it to the list in run.mjs

, and open a PR with your results. A guard that mounts an actual published plugin (rather than a strategy archetype) is especially welcome.

v0, honest limits stated in RESULTS.md: the corpus is small and hand-authored, the deny-list is a strategy archetype (not a specific plugin), and the numbers are one model / one run. The value is the shape it exposes and that anyone can re-run it. MIT.

── more in #ai-safety 4 stories · sorted by recency
── more on @holdline 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/show-hn-a-benchmark-…] indexed:0 read:1min 2026-08-17 ·