cd /news/ai-agents/i-a-b-tested-9-popular-ai-agent-skil… · home › topics › ai-agents › article
[ARTICLE · art-144827] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

I A/B-tested 9 popular AI agent skills. 4 of them did nothing.

A developer A/B-tested nine popular AI coding-agent skills against Sonnet 5.5 and Haiku 4.5 using Anthropic's claude plugin eval, running 19 test cases three times per arm, and found four of the nine added nothing. Workflow-changing skills like 'grill' (interview the user one question at a time) gained +75 points on Sonnet and +58 on Haiku, while habit-restating skills such as 'surgical' and 'stdlib-first' scored zero and never even loaded in 0 of 6 runs. The developer also reported that a faulty grader initially made 'handoff' appear to hurt Sonnet by 13 points before the corrected run showed no effect.

by read4 min views2 publishedOct 4, 2026

Agent skills are everywhere this year. A skill is a SKILL.md file that teaches a coding agent

(Claude Code, Codex, Cursor, Gemini CLI…) how to do something: verify before saying "done",

keep diffs small, review code for real bugs. Some skill repos have hundreds of thousands of stars.

I noticed nobody measures them. A skill is a prompt, and whether a prompt helps depends on the

model reading it. A skill written for last year's model might do nothing on this year's, or

make it worse. So I tested them.

I took 9 of the most popular skill ideas and rewrote them for current models: short, calm, no

walls of MUST and NEVER. Then I gave each one an eval suite using Anthropic's

claude plugin eval, which runs every task

twice:

The difference between the two scores (Δ) is what the skill actually adds. That came to 19 test

cases, 3 runs each per arm, on Sonnet 5.5 and Haiku 4.5. The prompts read like what a

real user would type, and they never name the skill. The graders check outcomes ("were the

unrelated lines left untouched?", "did it admit the fix was untested?"), not whether the reply

followed the skill's own formatting.

Skill Sonnet 5.5 Δ Haiku 4.5 Δ Verdict
grill : interview me before coding, one question at a time, each with a recommended answer +75 +58 ✅ keep
bug-hunt-review : report only real bugs, each with a concrete failing input +10 +17 ✅ keep
handoff : write a note a fresh session can resume from 0 +13 ✅ keep (small models)
prove-it : don't say "fixed" without the command that shows it 0 +11 ✅ keep (small models)
root-cause : fix the bug where it starts, not where it was reported +13 −11 ⚠️ on probation
surgical : smallest possible diff 0 0 ✂️ cut
stdlib-first : built-ins before new packages 0 0 ✂️ cut
answer-first : first sentence is the answer 0 +3 ✂️ cut
secure-defaults : parameterized SQL, no shell strings 0 −8 ✂️ cut

Four of nine were cut. They're still in the repo under retired/, with their evals, so anyone

can re-test them on a future model.

1. Skills that add a workflow help. Skills that restate good habits don't.

grill makes the model do something it wouldn't choose on its own: ask one question at a time

and recommend an answer for each. Without it, Sonnet asked five or more questions at once and didn't recommend

an answer for any of them. That's a +75 point difference. But "keep your diff small" and "parameterize

your SQL"? Sonnet 5.5 already does that. The skill adds nothing.

2. Most "be careful" skills never even loaded.

On natural prompts, surgical, stdlib-first, answer-first and secure-defaults were loaded

in 0 of 6 runs. The model decided they weren't relevant, and it was right: it already behaved

that way. A skill that never fires still costs context on every turn, because its description

is always loaded.

3. Smaller models benefit more.

handoff and prove-it did nothing for Sonnet, which already writes accurate handoffs and

admits when it couldn't run the tests. Haiku gained 11–13 points from them.

4. A skill's description alone can change behavior.

root-cause never loaded on either model, yet scored +13 on Sonnet and −11 on Haiku. The only

part of it the model saw was its one-line description in the skill list. That's a weak, noisy

effect, so it stays on probation instead of claiming a win.

5. Check your graders before you trust your numbers.

My first run showed handoff hurting Sonnet by 13 points. The cause was my grader: it

required the first "next step" to name a function to change, and it failed the correct answer,

"re-run the tests first". After I fixed the grader, the effect was 0. The fix is noted in the

changelog, and the README table only uses the corrected run.

Three runs per arm is noisy, and some of my cases are probably too easy: when the baseline

already scores 100%, a skill can't show a benefit. Harder eval cases are the most useful thing

anyone could contribute.

While doing this I built skill-vet, a zero-dependency scanner you can point at any skill

repo before installing it:

npx @menadirali/skill-vet vet owner/repo

It checks for download-and-execute (curl … | sh), hidden Unicode, prompt-injection phrasing,

credential access, the skill spec, and how many tokens a skill pack adds to every session. I ran

it on 115 skills from 9 of the most popular skill repos. It found no download-and-execute,

prompt-injection, or hidden-Unicode problems, a couple of spec errors, and lots of all-caps

"shouting" that current models don't need.

npx skills add nadirali1350/vetted, or in Claude Code, /plugin marketplace add nadirali1350/vetted If you've seen your coding agent repeatedly get something wrong on a current model, I'd love to

hear it. That's how the next skill gets written, with its eval first.

── more in #ai-agents 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/i-a-b-tested-9-popul…] indexed:0 read:4min 2026-10-04 · —