cd /news/ai-agents/we-re-tested-69-skills-on-harder-inp… · home › topics › ai-agents › article
[ARTICLE · art-144905] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

We re-tested 69 skills on harder inputs. 16 failed. Here is what broke.

A developer behind the ProSkillPacks agent-skills project re-tested all 69 of its published skills on harder inputs and found 16 defects — roughly one in four — none of which had surfaced in the skills' original passing runs. The most common failure, seen in 7 of the 16 cases, was restating uncertain input as fact in user-facing output, followed by invented claims, miscounts, unnamed conflicts, a script misreading a compressed response, and one skill writing a file it should not have. Each fix required only one or two sentences of instruction changes, except for a single script fix, and every repaired skill passed a re-run on the same input.

by read2 min views1 publishedOct 4, 2026

Disclosure: we make agent skills (see the end). The checklist below is ours and needs no purchase. Written with AI assistance.

Every skill we ship had already passed one run on a real public input. We read that output, it looked right, and we moved on. Then we gave all 69 skills a second run on an input made to be hard. Sixteen had a defect. About one in four.

None of the 16 failed the first run. That is the point: one passing run tells you the skill can do the job, not that it will.

For each skill we picked one of these, using public pages or text we wrote: Then we read each output against the input ourselves, with no grader and no second model.

What broke Defects
Unsure input restated as fact 7
Invented or unsupported claim 3
Miscount in a summary 2
Conflict in the brief not named 2
Script misread a compressed response 1
Wrote a file it should not have 1

The largest group is the quiet one. The skill noticed that the input was unsure, marked it once, and then stated it as fact in the text meant for someone else.

A newsletter skill was given notes where a postage change was unconfirmed. Its first subject line was "Postage prices are going up". After the fix it was "Is postage going up this April?".

A host wrote "Smoking outside I guess". The paste-ready rules said "Smoking is outside only." After the fix the setting read "Not settled. Please confirm."

A listing said "5 min from the beach (maybe 10 if the tide is in)". Title option 1 was "Studio 5 min from the beach, queen bed". After the fix the title said "near the beach" and the hedge stayed in the description.

A claim checker wrote "There are 9 claims, and 8 are high risk." The table under it had 7 High and 2 Medium. A CSV profiler that is meant to be read-only wrote "I left orders.csv untouched and wrote the cleaned version to orders_clean.csv", which contradicts itself in one sentence.

Every fix was one or two sentences in the skill's instructions, except one script fix. Every fixed skill passed a re-run on the same input.

You can run this on any prompt or skill in an afternoon.

All runs used Claude Sonnet, so another model may fail elsewhere. We judged the outputs ourselves. One hard input per skill found 16, and a third run might find more. Several inputs were written by us to be hard, so they show behaviour on that kind of input, not how often real use will hit it.

The full table, the quoted before and after lines and the data are on the study page: https://proskillpacks.github.io/study/retest/?utm_source=devto&utm_medium=post&utm_campaign=devto-retest

We make agent skills. The free ones are at https://github.com/proskillpacks/skills

── more in #ai-agents 4 stories · sorted by recency
── more on @proskillpacks 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/we-re-tested-69-skil…] indexed:0 read:2min 2026-10-04 · —