{"slug": "i-a-b-tested-9-popular-ai-agent-skills-4-of-them-did-nothing", "title": "I A/B-tested 9 popular AI agent skills. 4 of them did nothing.", "summary": "A developer A/B-tested nine popular AI coding-agent skills against Sonnet 5.5 and Haiku 4.5 using Anthropic's claude plugin eval, running 19 test cases three times per arm, and found four of the nine added nothing. Workflow-changing skills like 'grill' (interview the user one question at a time) gained +75 points on Sonnet and +58 on Haiku, while habit-restating skills such as 'surgical' and 'stdlib-first' scored zero and never even loaded in 0 of 6 runs. The developer also reported that a faulty grader initially made 'handoff' appear to hurt Sonnet by 13 points before the corrected run showed no effect.", "body_md": "Agent skills are everywhere this year. A skill is a `SKILL.md` file that teaches a coding agent\n\n(Claude Code, Codex, Cursor, Gemini CLI…) how to do something: verify before saying \"done\",\n\nkeep diffs small, review code for real bugs. Some skill repos have hundreds of thousands of stars.\n\nI noticed nobody measures them. A skill is a prompt, and whether a prompt helps depends on the\n\nmodel reading it. A skill written for last year's model might do nothing on this year's, or\n\nmake it worse. So I tested them.\n\nI took 9 of the most popular skill ideas and rewrote them for current models: short, calm, no\n\nwalls of `MUST` and `NEVER`. Then I gave each one an eval suite using Anthropic's\n\n[`claude plugin eval`](https://code.claude.com/docs/en/plugin-evals), which runs every task\n\ntwice:\n\nThe difference between the two scores (Δ) is what the skill actually adds. That came to 19 test\n\ncases, 3 runs each per arm, on **Sonnet 5.5** and **Haiku 4.5**. The prompts read like what a\n\nreal user would type, and they never name the skill. The graders check outcomes (\"were the\n\nunrelated lines left untouched?\", \"did it admit the fix was untested?\"), not whether the reply\n\nfollowed the skill's own formatting.\n\n| Skill | Sonnet 5.5 Δ | Haiku 4.5 Δ | Verdict | \n|---|---|---|---|\n| `grill` : interview me before coding, one question at a time, each with a recommended answer | **+75** | **+58** | ✅ keep | \n| `bug-hunt-review` : report only real bugs, each with a concrete failing input | **+10** | **+17** | ✅ keep | \n| `handoff` : write a note a fresh session can resume from | 0 | **+13** | ✅ keep (small models) | \n| `prove-it` : don't say \"fixed\" without the command that shows it | 0 | **+11** | ✅ keep (small models) | \n| `root-cause` : fix the bug where it starts, not where it was reported | +13 | −11 | ⚠️ on probation | \n| `surgical` : smallest possible diff | 0 | 0 | ✂️ cut | \n| `stdlib-first` : built-ins before new packages | 0 | 0 | ✂️ cut | \n| `answer-first` : first sentence is the answer | 0 | +3 | ✂️ cut | \n| `secure-defaults` : parameterized SQL, no shell strings | 0 | −8 | ✂️ cut | \n\nFour of nine were cut. They're still in the repo under `retired/`, with their evals, so anyone\n\ncan re-test them on a future model.\n\n**1. Skills that add a workflow help. Skills that restate good habits don't.**\n\n`grill` makes the model do something it wouldn't choose on its own: ask one question at a time\n\nand recommend an answer for each. Without it, Sonnet asked five or more questions at once and didn't recommend\n\nan answer for any of them. That's a +75 point difference. But \"keep your diff small\" and \"parameterize\n\nyour SQL\"? Sonnet 5.5 already does that. The skill adds nothing.\n\n**2. Most \"be careful\" skills never even loaded.**\n\nOn natural prompts, `surgical`, `stdlib-first`, `answer-first` and `secure-defaults` were loaded\n\nin **0 of 6** runs. The model decided they weren't relevant, and it was right: it already behaved\n\nthat way. A skill that never fires still costs context on every turn, because its description\n\nis always loaded.\n\n**3. Smaller models benefit more.**\n\n`handoff` and `prove-it` did nothing for Sonnet, which already writes accurate handoffs and\n\nadmits when it couldn't run the tests. Haiku gained 11–13 points from them.\n\n**4. A skill's description alone can change behavior.**\n\n`root-cause` never loaded on either model, yet scored +13 on Sonnet and −11 on Haiku. The only\n\npart of it the model saw was its one-line description in the skill list. That's a weak, noisy\n\neffect, so it stays on probation instead of claiming a win.\n\n**5. Check your graders before you trust your numbers.**\n\nMy first run showed `handoff` *hurting* Sonnet by 13 points. The cause was my grader: it\n\nrequired the first \"next step\" to name a function to change, and it failed the correct answer,\n\n\"re-run the tests first\". After I fixed the grader, the effect was 0. The fix is noted in the\n\nchangelog, and the README table only uses the corrected run.\n\nThree runs per arm is noisy, and some of my cases are probably too easy: when the baseline\n\nalready scores 100%, a skill can't show a benefit. Harder eval cases are the most useful thing\n\nanyone could contribute.\n\nWhile doing this I built **`skill-vet`**, a zero-dependency scanner you can point at any skill\n\nrepo *before* installing it:\n\n```\nnpx @menadirali/skill-vet vet owner/repo\n```\n\nIt checks for download-and-execute (`curl … | sh`), hidden Unicode, prompt-injection phrasing,\n\ncredential access, the skill spec, and how many tokens a skill pack adds to every session. I ran\n\nit on 115 skills from 9 of the most popular skill repos. It found no download-and-execute,\n\nprompt-injection, or hidden-Unicode problems, a couple of spec errors, and lots of all-caps\n\n\"shouting\" that current models don't need.\n\n`npx skills add nadirali1350/vetted`, or in Claude Code,\n`/plugin marketplace add nadirali1350/vetted`\nIf you've seen your coding agent repeatedly get something wrong on a current model, I'd love to\n\nhear it. That's how the next skill gets written, with its eval first.", "url": "https://wpnews.pro/news/i-a-b-tested-9-popular-ai-agent-skills-4-of-them-did-nothing", "canonical_source": "https://dev.to/menadirali/i-ab-tested-9-popular-ai-agent-skills-4-of-them-did-nothing-5ba5", "published_at": "2026-10-04 12:29:23+00:00", "updated_at": "2026-10-04 12:42:41.025571+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "large-language-models", "developer-tools", "ai-research"], "entities": ["Anthropic", "Claude Code", "Codex", "Cursor", "Gemini CLI", "Sonnet 5.5", "Haiku 4.5"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-a-b-tested-9-popular-ai-agent-skills-4-of-them-did-nothing", "markdown": "https://wpnews.pro/news/i-a-b-tested-9-popular-ai-agent-skills-4-of-them-did-nothing.md", "text": "https://wpnews.pro/news/i-a-b-tested-9-popular-ai-agent-skills-4-of-them-did-nothing.txt", "jsonld": "https://wpnews.pro/news/i-a-b-tested-9-popular-ai-agent-skills-4-of-them-did-nothing.jsonld"}}