cd /news/ai-agents/claude-code-plugin-eval-gate-skills-… · home topics ai-agents article
[ARTICLE · art-129298] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Claude Code plugin eval: gate skills on delta before you merge

Anthropic shipped a first-party evaluation command for Claude Code, `claude plugin eval`, that runs a test suite with a plugin loaded and again without it to report the delta the plugin actually contributes. The workflow, available in Claude Code v2.1.269+, lets teams create cases under an `evals/` directory, grade outcomes, and fail CI builds when scores drop below a threshold. The documentation warns that a case scoring 1.00 with a plugin loaded can still be useless if the model already solves the prompt unaided, since a near-zero delta means the plugin did not move the outcome.

by read5 min views3 publishedSep 14, 2026

Claude Code now ships a first-party way to prove a plugin or skill actually helps. The command is claude plugin eval. It runs your suite with the plugin loaded and again without it, then reports WITH, W/OUT, and Δ — the lift the plugin contributed. That changes what production teams should do before they merge a skill into a kit repo or pin it in CI.

This is not another “write better prompts” essay. It is a builder workflow for Claude Code v2.1.269+: create cases under evals/, grade outcomes, and fail the build when the score drops. If you already ship agent docs in an owned monorepo, pair this with Claude Code on an OTF kit: the CLAUDE.md and prompts that ship with the repo. For feature-level evals outside plugins, keep using A practical LLM evaluation loop for AI features that need to ship — different layer, different artifact.

Anthropic’s plugin evals docs describe a dedicated CLI path for plugin and skill authors. The important claims for builders:

claude --version, then claude update if needed).evals/ directory next to your plugin manifest (plugin.json or .claude-plugin/plugin.json).Δ is with-arm score minus without-arm score.claude plugin eval init can propose cases and graders interactively; claude plugin eval init --bare <case> writes a blank template.--threshold, pin --model / --judge-model, keep reports local with --no-publish, and cap spend with --max-cost-usd. That is the DO: stop merging skills because a demo chat looked fine. Measure contribution, then gate.

A case that scores 1.00 with the plugin loaded can still be useless. Claude might already solve the prompt without your skill. The docs call this out: if WITH and W/OUT are both high, Δ near zero means the plugin did not move the outcome.

Production implications:

Δ ≈ 0 with a tool_used: Skill grader failing — Claude never chose your skill on natural phrasing. Fix the skill description, re-run, compare. tool_used graders on Skill are excluded from both arms’ scored totals so you do not invent lift by checking something impossible without the plugin.regex, tool_order, and file_exists are free transcript/file checks. llm and baseline call a judge model and add to the run’s list-price estimate.--runs 1 --ablation none; trust needs the default three. If you only ever run the with-arm, you are grading absolute behavior. That is fine while drafting graders. Before you call a skill “done,” turn the baseline back on and read Δ.

From the plugin root (the directory that contains the manifest):

claude --version   # need >= 2.1.269
claude plugin eval init --bare commit-message

You get something shaped like:

evals/commit-message/
├── prompt.md
└── graders/
    └── criteria.md

Edit prompt.md so the body is a request a user would type — do not name the skill in the prompt:

---
max_turns: 10
allowed_tools: [Read, Glob, Grep, Skill]
---

Write me a commit message for this change: I renamed getUser to fetchUser and updated the three call sites.

Add a result grader (llm or regex) and a skill-fired grader:

---
type: tool_used
tool: Skill
input_match: '"skill"\s*:\s*"(?:[\w-]+:)?your-skill-name"'
---

Then run:

claude plugin eval .

Expect six runs for one case at defaults (3× with, 3× without). Open the HTML report path printed at the end. Iterate on description and graders until Δ is positive for the prompts you care about.

For MCP-backed skills, put mocks under evals/mocks/<server>/<tool>.md so CI does not need the real service. Use --scaffold only for suites you trust — scaffold scripts run as you, outside the agent sandbox.

The docs’ recommended CI shape is explicit:

claude plugin eval . \
  --trust-plugin \
  --json results.json \
  --threshold 0.8 \
  --model claude-sonnet-5 \
  --judge-model claude-haiku-4-5 \
  --no-publish \
  --max-cost-usd 20

Exit codes that matter:

0 — every case met 1 — a case scored below threshold, files failed to load, or trust was missing without --trust-plugin 2 — partial run (partial: true Pin both models so a provider rollout is not mistaken for a plugin regression. Keep every-change suites on free graders when you can; reserve llm judges for short outputs with concrete PASS/FAIL rubrics. Leave partial: true results out of trend charts.

Credentials and install still belong on the runner (ANTHROPIC_API_KEY or your normal Claude Code auth). Without --trust-plugin, a non-interactive job against an untrusted checkout exits 1.

OTF’s wedge is not “another sandboxed chat.” It is an owned repo agents keep editing — CLAUDE.md, prompts, and skills that travel with the product. Plugin eval is the missing acceptance gate for those skills:

Layer Question Artifact
Repo conventions Can an agent find the seams? CLAUDE.md /.cursorrules (see the kit docs post above)
Change acceptance Did this PR do what we asked? AI coding agent acceptance checklist
Plugin contribution Did the skill raise the score vs no plugin? claude plugin eval +Δ
Product AI features Does the user-facing model path hold? LLM evaluation loop

Put the suite next to the plugin you ship. Fail CI when Δ collapses after a description tweak or a model pin change. That is how skills stay production assets instead of chat folklore.

Security note from the same docs: pointing claude plugin eval at a plugin is the same trust decision as claude --plugin-dir. Only evaluate plugins you trust. Hooks and real MCP servers you opt into with --allow-real-servers or --mocks off run outside the agent sandbox.

If you are still choosing between a sandboxed MVP host and a repo you own, start with the kit spine, then add this gate. Skills without Δ are demos. Skills with a CI threshold are release machinery.

── more in #ai-agents 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/claude-code-plugin-e…] indexed:0 read:5min 2026-09-14 ·