# I A/B-tested 9 popular AI agent skills. 4 of them did nothing.

> Source: <https://dev.to/menadirali/i-ab-tested-9-popular-ai-agent-skills-4-of-them-did-nothing-5ba5>
> Published: 2026-10-04 12:29:23+00:00

Agent skills are everywhere this year. A skill is a `SKILL.md` file that teaches a coding agent

(Claude Code, Codex, Cursor, Gemini CLI…) how to do something: verify before saying "done",

keep diffs small, review code for real bugs. Some skill repos have hundreds of thousands of stars.

I noticed nobody measures them. A skill is a prompt, and whether a prompt helps depends on the

model reading it. A skill written for last year's model might do nothing on this year's, or

make it worse. So I tested them.

I took 9 of the most popular skill ideas and rewrote them for current models: short, calm, no

walls of `MUST` and `NEVER`. Then I gave each one an eval suite using Anthropic's

[`claude plugin eval`](https://code.claude.com/docs/en/plugin-evals), which runs every task

twice:

The difference between the two scores (Δ) is what the skill actually adds. That came to 19 test

cases, 3 runs each per arm, on **Sonnet 5.5** and **Haiku 4.5**. The prompts read like what a

real user would type, and they never name the skill. The graders check outcomes ("were the

unrelated lines left untouched?", "did it admit the fix was untested?"), not whether the reply

followed the skill's own formatting.

| Skill | Sonnet 5.5 Δ | Haiku 4.5 Δ | Verdict | 
|---|---|---|---|
| `grill` : interview me before coding, one question at a time, each with a recommended answer | **+75** | **+58** | ✅ keep | 
| `bug-hunt-review` : report only real bugs, each with a concrete failing input | **+10** | **+17** | ✅ keep | 
| `handoff` : write a note a fresh session can resume from | 0 | **+13** | ✅ keep (small models) | 
| `prove-it` : don't say "fixed" without the command that shows it | 0 | **+11** | ✅ keep (small models) | 
| `root-cause` : fix the bug where it starts, not where it was reported | +13 | −11 | ⚠️ on probation | 
| `surgical` : smallest possible diff | 0 | 0 | ✂️ cut | 
| `stdlib-first` : built-ins before new packages | 0 | 0 | ✂️ cut | 
| `answer-first` : first sentence is the answer | 0 | +3 | ✂️ cut | 
| `secure-defaults` : parameterized SQL, no shell strings | 0 | −8 | ✂️ cut | 

Four of nine were cut. They're still in the repo under `retired/`, with their evals, so anyone

can re-test them on a future model.

**1. Skills that add a workflow help. Skills that restate good habits don't.**

`grill` makes the model do something it wouldn't choose on its own: ask one question at a time

and recommend an answer for each. Without it, Sonnet asked five or more questions at once and didn't recommend

an answer for any of them. That's a +75 point difference. But "keep your diff small" and "parameterize

your SQL"? Sonnet 5.5 already does that. The skill adds nothing.

**2. Most "be careful" skills never even loaded.**

On natural prompts, `surgical`, `stdlib-first`, `answer-first` and `secure-defaults` were loaded

in **0 of 6** runs. The model decided they weren't relevant, and it was right: it already behaved

that way. A skill that never fires still costs context on every turn, because its description

is always loaded.

**3. Smaller models benefit more.**

`handoff` and `prove-it` did nothing for Sonnet, which already writes accurate handoffs and

admits when it couldn't run the tests. Haiku gained 11–13 points from them.

**4. A skill's description alone can change behavior.**

`root-cause` never loaded on either model, yet scored +13 on Sonnet and −11 on Haiku. The only

part of it the model saw was its one-line description in the skill list. That's a weak, noisy

effect, so it stays on probation instead of claiming a win.

**5. Check your graders before you trust your numbers.**

My first run showed `handoff` *hurting* Sonnet by 13 points. The cause was my grader: it

required the first "next step" to name a function to change, and it failed the correct answer,

"re-run the tests first". After I fixed the grader, the effect was 0. The fix is noted in the

changelog, and the README table only uses the corrected run.

Three runs per arm is noisy, and some of my cases are probably too easy: when the baseline

already scores 100%, a skill can't show a benefit. Harder eval cases are the most useful thing

anyone could contribute.

While doing this I built **`skill-vet`**, a zero-dependency scanner you can point at any skill

repo *before* installing it:

```
npx @menadirali/skill-vet vet owner/repo
```

It checks for download-and-execute (`curl … | sh`), hidden Unicode, prompt-injection phrasing,

credential access, the skill spec, and how many tokens a skill pack adds to every session. I ran

it on 115 skills from 9 of the most popular skill repos. It found no download-and-execute,

prompt-injection, or hidden-Unicode problems, a couple of spec errors, and lots of all-caps

"shouting" that current models don't need.

`npx skills add nadirali1350/vetted`, or in Claude Code,
`/plugin marketplace add nadirali1350/vetted`
If you've seen your coding agent repeatedly get something wrong on a current model, I'd love to

hear it. That's how the next skill gets written, with its eval first.
