{"slug": "i-ran-claude-plugin-eval-on-89-skills-the-trigger-rate-fell-from-100-to-55-6", "title": "I ran claude plugin eval on 89 skills. The trigger rate fell from 100% to 55.6%.", "summary": "A developer measured Claude Code's new `claude plugin eval` tool against 89 skill directories in a coffee e-commerce project, finding that trigger rates for positive test cases fell from 100% with a single skill loaded to 55.6% when all 89 competed for the model's attention. The eval's default second arm, which runs each case without the plugin loaded, revealed that the negative case still scored 100% in both configurations. The developer then deleted the sandbox trace directories that would have shown which skills the model invoked instead, losing the data needed to distinguish correct delegation from description overlap.", "body_md": "My coffee e-commerce project has 89 skill directories under it. Each one exists because I got burned: remember the GRANT on that migration, keep those three layers in sync when you rename a column, query the database before you guess at the code.\n\nSkills work by having the model read their descriptions and decide, on its own, whether to invoke one. That's the whole problem. **When it doesn't invoke, nothing happens.** No error, no warning, no log line. You just watch the AI start working, and find out two days later that it skipped the rule.\n\nI measured this by hand once, back in July: four natural-language prompts, four sessions, counting how many times the skill name showed up in the reply. Sonnet hit 3 of 4, Haiku hit 1 of 4. I wrote the numbers into my project rules and never measured again, because doing it by hand is expensive and gives you a one-off number that goes stale.\n\nLast week Claude Code quietly shipped `claude plugin eval`.\n\nYou write test cases — realistic sentences a user might type — plus graders that score the result. Each case runs in a fresh isolated session, three times by default because the model is non-deterministic.\n\nStandard stuff. What surprised me: **it runs a second arm by default, with your plugin not loaded at all.**\n\nThe difference between the two scores is `Δ`. The docs put it bluntly:\n\nIf a case scores 1.0 both with and without the plugin, the plugin isn't what made it pass.\n\nVendors don't usually ship a default that proves their users' work is useless. This design concedes something: **your skill might be doing nothing, and you'd have no way to know.**\n\nI reused the four prompts from July rather than writing new ones. Writing your own test corpus has a trap — you unconsciously write sentences your skill happens to catch. The old corpus was written before I knew what I'd be measuring.\n\nThree positive cases (add a back-in-stock notification, add a ship-date column, add a daily-summary page for staff) and one negative (change the homepage headline and button colour). One grader: did that skill get invoked, yes or no.\n\nResult: **12 for 12.** All three positives fired, the negative correctly didn't. $2.49, 273 seconds.\n\nI knew immediately the number was useless, because the test plugin contained **exactly one skill**. When there's only one candidate and its name matches the task, picking it costs the model nothing. My real environment has 89 competing for the same attention.\n\nI wrote an assembly script that builds a throwaway plugin from all 89 skills at run time. (Deliberately no skill copies in version control — copies drift from the original, and this project has already paid that bill.)\n\nSame prompts, same model. The only variable is 1 skill versus 89:\n\n| Case | 1 skill | 89 skills | \n|---|---|---|\n| Back-in-stock notification | 100% | 67% | \n| Ship-date column | 100% | 33% | \n| Staff daily-summary page | 100% | 67% | \n| Homepage copy (should NOT fire) | 100% | 100% | \n\nPositives combined: 5 of 9, **55.6%**. The negative case stayed correct in both.\n\nThat lands in the same range as my July hand-measurement (Sonnet 75%, Haiku 25%), which is mildly reassuring — two completely independent methods, one manual and one automated, agreeing.\n\nI had a conclusion forming: too many skills dilutes attention, time to prune.\n\nBefore writing that conclusion I wanted one more thing: in the runs where the skill *didn't* fire, **what did the model invoke instead?**\n\nIf it picked a more appropriate skill, that's correct delegation, not a defect. If it picked one with an overlapping description, the fix is editing descriptions, not deleting skills. Both leave fingerprints in the trace.\n\nThen I realised I had deleted every trace ten minutes earlier.\n\nThe tool leaves a sandbox directory per run and prints \"remove it when you're done.\" I cleaned them all up in one sweep. Those directories held the only thing that could answer the question.\n\nThis was the expensive lesson of the day, and it's a dumb one: **I deleted something far more valuable than tidiness, in order to be tidy.** The tool tells you to remove them because they may contain agent-written content. That is not the same as telling you to remove them before you've looked.\n\nRe-run with `--keep-temp`: $0.82.\n\nAll three re-runs failed to invoke. Here's every tool call each one made:\n\n```\nrun 1: Glob, Glob, Glob, Glob, Glob, Grep\nrun 2: Glob, Glob, Glob, Glob, Glob\nrun 3: Glob, Glob, Glob, Glob, Glob\n```\n\n**Not one alternative skill invoked.** It went looking through files instead.\n\nSo \"correct delegation\" and \"overlapping descriptions\" are both ruled out. But the thing that actually changed my conclusion was the final reply. All three looked like this (excerpt, translated):\n\nBased on this project's skill descriptions, here's what I expect to touch (**this is expectation, not fact** — I'll confirm by reading files once you give me the path):\n\n**DB**: new migration adding `orders.estimated_ship_date`, with the tenant_id / RLS / GRANT boilerplate\n**Write path**: whitelist the field in the admin order edit payload (don't spread formData raw)\n**Middle layer**: if there's an RPC `jsonb_build_object` or a `normalizeOrder` whitelist, add it there too, otherwise you get \"value in DB, null in frontend\"\n**Presentation**: `orderPresentation.js` SSOT, plus the admin order list…\n\nThose four points **are what that skill would have produced.** GRANT boilerplate, no raw form spreading, the middle-layer whitelist that silently nulls your field, the five order surfaces that drift apart. It even labelled its own confidence level.\n\nIt never invoked the Skill tool. It read the description and reasoned from it.\n\nSo my 55.6% measures **invocation**, not **effect**. Those are separable in the trace, and I had assumed they were the same thing.\n\nAll three traces show the model stuck in the same place: the working directory is empty, so it burns five Glob calls hunting for project files and ends with \"please tell me where the project is.\"\n\nThat's my fault. Every eval run starts in an empty workspace, and the tool supports seeding one with a scaffold script. I skipped it. An empty directory pushes the model's attention toward \"find the files\" instead of \"is this an architectural change?\"\n\nSo the honest framing is: **these numbers measure whether the model formally invokes a skill in an empty workspace.** That's some distance from the question I care about.\n\nOne more trap, because it will fool anyone who doesn't read the surrounding columns.\n\nMy first attempt at proving the grader could actually fail produced this:\n\n```\nrun 1/1 [with]: score 0.00  $0.00  error: exit 1: Credit balance is too low\n  ✗ skill-fired: Skill called 0x (expected 1..∞)\n```\n\nI wanted a red light, and it handed me a red light.\n\nBut that 0.00 isn't \"the grader caught something.\" **The run never happened.** API credit ran out, the session never started, the grader searched an empty transcript, found no skill invocation, and rendered 0.\n\n(The credit thing turned out to be a separate wallet from my subscription — eval uses whatever credentials your normal sessions use, and `ANTHROPIC_API_KEY` in my environment was routing it to the API account. Unsetting it for the child process fell back to subscription auth.)\n\nThe real one, once it ran:\n\n```\nrun 1/1 [with]:    score 1.00  $0.30\nrun 1/1 [without]: score 0.00  $0.15   ← this 0.00 is real\nΔ +1.00\n```\n\nTwo identical-looking zeros. **The discriminator is the columns next to them**: the fake one cost $0.00 and carried an error string; the real one cost $0.15 with a null error.\n\nThis is the part I find most worth thinking about, and I can only offer speculation.\n\nThey could have published a \"how to write good skill descriptions\" guide for a fraction of the cost. Instead they built a measuring instrument — one whose default behaviour is to **prove your work contributed nothing.**\n\nThree signals I can read from that:\n\n**One: this isn't a documentation problem.** If unclear descriptions were the issue, a style guide would fix it. Building a measurement tool implies they think this needs per-case empirical testing, and that results shift with model versions. The docs say to pin your model in CI so \"a model rollout isn't mistaken for a plugin regression.\" That sentence concedes the same skill performs differently across models.\n\n**Two: the ablation arm is the default, not a flag.** That's the strongest signal in the design. It assumes your plugin might not be contributing and computes the answer whether you asked or not.\n\n**Three: this is a tool that only becomes necessary at a certain ecosystem size.** With five skills you don't need to measure. Past some threshold the question shifts from \"do I have a skill for this?\" to \"will it win the attention contest?\" I have 89. That number didn't exist two years ago.\n\nMy guess is they saw this curve before I did. That's a guess — I have no inside information.\n\n**Known:**\n\n**Not known:**\n\nIf you've got a pile of skills, prompts or rule files, my suggestion isn't to prune. It's to **measure once** — and when you do, remember that \"was the tool invoked\" is the easy grader to write, and probably not the question you care about.\n\nAs for whether I should cut my 89 down: I have no evidence either way yet. Next step is seeding the workspace and running two models. That's when I get to have an opinion.\n\nRelated, from the same toolchain: [Why does the AI always leave a few blocks out?](https://dev.to/content/completeness-baseline-en) — the other half of this problem, where the verifier checks whether your declarations are internally consistent and never whether they match the world. And [compiling AI reasoning into scripts](https://dev.to/content/crystallize-compile-ai-reasoning) (in Chinese), on paying for frontier reasoning once and replaying it for free.\n\n*Originally published on my blog: [I ran claude plugin eval on 89 skills. The trigger rate fell from 100% to 55.6%.](https://coffeeshooters.com/content/skill-fired-is-not-skill-worked-en?utm_source=devto&utm_medium=social&utm_campaign=blog-skill-fired-is-not-skill-worked-en)*\n\n*I keep a running index of every pothole I've hit building a real production system solo — symptom on the left, what to grep in your own repo on the right: [coffeeshooters.com/potholes](https://coffeeshooters.com/potholes?utm_source=devto&utm_medium=social&utm_campaign=potholes-index)*\n\n*And if your team is shipping AI-written code faster than anyone can read it, that's the thing I do for a living: [coffeeshooters.com/code-audit](https://coffeeshooters.com/code-audit?utm_source=devto&utm_medium=social&utm_campaign=code-audit-offer)*", "url": "https://wpnews.pro/news/i-ran-claude-plugin-eval-on-89-skills-the-trigger-rate-fell-from-100-to-55-6", "canonical_source": "https://dev.to/dexterlung/i-ran-claude-plugin-eval-on-89-skills-the-trigger-rate-fell-from-100-to-556-n29", "published_at": "2026-09-28 13:05:19+00:00", "updated_at": "2026-09-28 13:20:09.616292+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "large-language-models", "mlops"], "entities": ["Claude Code", "Anthropic", "Sonnet", "Haiku"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-ran-claude-plugin-eval-on-89-skills-the-trigger-rate-fell-from-100-to-55-6", "markdown": "https://wpnews.pro/news/i-ran-claude-plugin-eval-on-89-skills-the-trigger-rate-fell-from-100-to-55-6.md", "text": "https://wpnews.pro/news/i-ran-claude-plugin-eval-on-89-skills-the-trigger-rate-fell-from-100-to-55-6.txt", "jsonld": "https://wpnews.pro/news/i-ran-claude-plugin-eval-on-89-skills-the-trigger-rate-fell-from-100-to-55-6.jsonld"}}