cd /news/ai-agents/i-ran-the-same-claude-code-skill-on-… · home › topics › ai-agents › article
[ARTICLE · art-147362] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

I ran the same Claude Code skill on Haiku 5.5 twice. Run 1: +55 points. Run 2: nothing. Driftproofhq

A developer at Driftproof ran the same Claude Code git-workflow skill on Claude Haiku 5.5 three times and found the measured benefit swung from +0.549 to -0.010 to +0.235, with the skill arm holding steady near 0.84 while the no-skill baseline jumped from 0.300 to 0.849 between runs. The report concludes the headline "+55 points" result was driven by a single bad baseline run, and warns that single-run skill evaluations can publish fake wins with real numbers attached.

by read2 min views1 publishedOct 8, 2026

Claude Haiku 5.5 came out on 7 October. Within a day I ran three popular Claude Code skills on it, and on Haiku 4.5 next to it. Three runs each, same task, same grader, same Claude Code version.

Here's the result that made me laugh.

The git workflow skill on Haiku 5.5, scored 0 to 1:

Run Without the skill With the skill Gap
1 0.300 0.849 +0.549
2 0.849 0.839 -0.010
3 0.613 0.848 +0.235

Run 1 is the screenshot that goes viral. "This one skill makes Haiku 5.5 55 points better."

Run 2 says the skill does nothing.

Same setup both times. And look at where the swing actually is: the skill arm sits at about 0.84 every time. It's Haiku 5.5 without the skill that jumped from 0.300 to 0.849 between runs. So the "+55 points" was mostly one bad baseline run.

Run it once and you would have published a fake win, with a real number attached.

That last line is the whole story. A lot of "I tried skill X and Claude got way better" posts are one run. Some popular skill-eval tools run a single attempt by default. Any one-run test would have reported run 1 above as a big win.

One task per skill, and only 3 to 10 answers per side in each run, so "too few answers to tell" means these runs couldn't see a 0.05 gap. It does not mean the skill is useless. The Haiku 5.5 calls also ran with Claude Code's per-turn effort setting and the Haiku 4.5 calls didn't, so the two models differ in more than the model. Every answer was scored three times by Claude Opus 5. This is not a model ranking.

The skills are from Addy Osmani's agent-skills pack, and I'm not knocking them. The point is that one run can't tell you whether any skill works, good or bad.

Full report, every number and every receipt: driftproofhq.com/reports/014 So, honest question: how many times do you run a skill before you decide it works?

*I maintain Driftproof, the open-source Claude Code plugin that ran this.

── more in #ai-agents 4 stories · sorted by recency
── more on @claude haiku 5.5 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/i-ran-the-same-claud…] indexed:0 read:2min 2026-10-08 · —