{"slug": "i-ran-the-same-claude-code-skill-on-haiku-5-5-twice-run-1-55-points-run-2", "title": "I ran the same Claude Code skill on Haiku 5.5 twice. Run 1: +55 points. Run 2: nothing. Driftproofhq", "summary": "A developer at Driftproof ran the same Claude Code git-workflow skill on Claude Haiku 5.5 three times and found the measured benefit swung from +0.549 to -0.010 to +0.235, with the skill arm holding steady near 0.84 while the no-skill baseline jumped from 0.300 to 0.849 between runs. The report concludes the headline \"+55 points\" result was driven by a single bad baseline run, and warns that single-run skill evaluations can publish fake wins with real numbers attached.", "body_md": "Claude Haiku 5.5 came out on 7 October. Within a day I ran three popular Claude Code skills on it, and on Haiku 4.5 next to it. Three runs each, same task, same grader, same Claude Code version.\n\nHere's the result that made me laugh.\n\nThe git workflow skill on Haiku 5.5, scored 0 to 1:\n\n| Run | Without the skill | With the skill | Gap | \n|---|---|---|---|\n| 1 | 0.300 | 0.849 | **+0.549** | \n| 2 | 0.849 | 0.839 | -0.010 | \n| 3 | 0.613 | 0.848 | +0.235 | \n\nRun 1 is the screenshot that goes viral. \"This one skill makes Haiku 5.5 55 points better.\"\n\nRun 2 says the skill does nothing.\n\nSame setup both times. And look at where the swing actually is: the skill arm sits at about 0.84 every time. It's Haiku 5.5 *without* the skill that jumped from 0.300 to 0.849 between runs. So the \"+55 points\" was mostly one bad baseline run.\n\nRun it once and you would have published a fake win, with a real number attached.\n\nThat last line is the whole story. A lot of \"I tried skill X and Claude got way better\" posts are one run. Some popular skill-eval tools run a single attempt by default. Any one-run test would have reported run 1 above as a big win.\n\nOne task per skill, and only 3 to 10 answers per side in each run, so \"too few answers to tell\" means these runs couldn't see a 0.05 gap. It does not mean the skill is useless. The Haiku 5.5 calls also ran with Claude Code's per-turn effort setting and the Haiku 4.5 calls didn't, so the two models differ in more than the model. Every answer was scored three times by Claude Opus 5. This is not a model ranking.\n\nThe skills are from Addy Osmani's agent-skills pack, and I'm not knocking them. The point is that one run can't tell you whether any skill works, good or bad.\n\nFull report, every number and every receipt: [driftproofhq.com/reports/014](https://driftproofhq.com/reports/014/)\n\nSo, honest question: how many times do you run a skill before you decide it works?\n\n*I maintain Driftproof, the open-source Claude Code plugin that ran this.", "url": "https://wpnews.pro/news/i-ran-the-same-claude-code-skill-on-haiku-5-5-twice-run-1-55-points-run-2", "canonical_source": "https://dev.to/driftproofhq/i-ran-the-same-claude-code-skill-on-haiku-55-twice-run-1-55-points-run-2-nothing-driftproofhq-50c4", "published_at": "2026-10-08 05:15:26+00:00", "updated_at": "2026-10-08 05:17:12.385505+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-tools", "developer-tools", "ai-research"], "entities": ["Claude Haiku 5.5", "Claude Haiku 4.5", "Claude Code", "Claude Opus 5", "Driftproof", "Addy Osmani"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-ran-the-same-claude-code-skill-on-haiku-5-5-twice-run-1-55-points-run-2", "markdown": "https://wpnews.pro/news/i-ran-the-same-claude-code-skill-on-haiku-5-5-twice-run-1-55-points-run-2.md", "text": "https://wpnews.pro/news/i-ran-the-same-claude-code-skill-on-haiku-5-5-twice-run-1-55-points-run-2.txt", "jsonld": "https://wpnews.pro/news/i-ran-the-same-claude-code-skill-on-haiku-5-5-twice-run-1-55-points-run-2.jsonld"}}