# I ran the same Claude Code skill on Haiku 5.5 twice. Run 1: +55 points. Run 2: nothing. Driftproofhq

> Source: <https://dev.to/driftproofhq/i-ran-the-same-claude-code-skill-on-haiku-55-twice-run-1-55-points-run-2-nothing-driftproofhq-50c4>
> Published: 2026-10-08 05:15:26+00:00

Claude Haiku 5.5 came out on 7 October. Within a day I ran three popular Claude Code skills on it, and on Haiku 4.5 next to it. Three runs each, same task, same grader, same Claude Code version.

Here's the result that made me laugh.

The git workflow skill on Haiku 5.5, scored 0 to 1:

| Run | Without the skill | With the skill | Gap | 
|---|---|---|---|
| 1 | 0.300 | 0.849 | **+0.549** | 
| 2 | 0.849 | 0.839 | -0.010 | 
| 3 | 0.613 | 0.848 | +0.235 | 

Run 1 is the screenshot that goes viral. "This one skill makes Haiku 5.5 55 points better."

Run 2 says the skill does nothing.

Same setup both times. And look at where the swing actually is: the skill arm sits at about 0.84 every time. It's Haiku 5.5 *without* the skill that jumped from 0.300 to 0.849 between runs. So the "+55 points" was mostly one bad baseline run.

Run it once and you would have published a fake win, with a real number attached.

That last line is the whole story. A lot of "I tried skill X and Claude got way better" posts are one run. Some popular skill-eval tools run a single attempt by default. Any one-run test would have reported run 1 above as a big win.

One task per skill, and only 3 to 10 answers per side in each run, so "too few answers to tell" means these runs couldn't see a 0.05 gap. It does not mean the skill is useless. The Haiku 5.5 calls also ran with Claude Code's per-turn effort setting and the Haiku 4.5 calls didn't, so the two models differ in more than the model. Every answer was scored three times by Claude Opus 5. This is not a model ranking.

The skills are from Addy Osmani's agent-skills pack, and I'm not knocking them. The point is that one run can't tell you whether any skill works, good or bad.

Full report, every number and every receipt: [driftproofhq.com/reports/014](https://driftproofhq.com/reports/014/)

So, honest question: how many times do you run a skill before you decide it works?

*I maintain Driftproof, the open-source Claude Code plugin that ran this.
