Caveman vs Ponytail vs Chisle: I benchmarked the Claude Code token-saving plugins on 20 tasks A developer built Chisle, a Claude Code plugin that trims token usage, and benchmarked it against the caveman and ponytail plugins plus a no-plugin baseline across 20 live tasks and 59+ model runs on Haiku and Sonnet. Chisle cut the total bill to 52% of baseline versus 80% for caveman and 68% for ponytail, with the smallest worst-case blowup (173% vs 424% and 227%) and only one backfire in 20 tasks, while all answers across every arm were graded correct in the July re-verification run. If you use Claude Code , you've probably seen the two popular plugins that promise to cut your token bill: caveman https://github.com/JuliusBrussee/caveman , which makes Claude talk like a caveman, and ponytail https://github.com/dietrichgebert/ponytail , which pushes it to write less code. I built a third one, Chisle https://chisle.jaypokale.me , and benchmarked all three against Claude with no plugin at all : 20 live tasks, 59+ model runs on Haiku and Sonnet, with the arms differing only in the injected ruleset. TL;DR: caveman cut the bill to 80% , ponytail to 68% , and Chisle to 52% , nearly half. Chisle also had the smallest worst day 173% vs 424% and 227% and backfired once in 20 tasks, where caveman backfired 6 times and ponytail 8. | 20 live tasks, billed output vs no plugin | total bill | average task | worst case | backfires | |---|---|---|---|---| | no plugin baseline | 100% | 100% | | | | caveman | 80% | 98% | 424% | 6 / 20 | | ponytail | 68% | 91% | 227% | 8 / 20 | | Chisle | 52% | 69% | 173% | 1 / 20 | In the July re-verification run, every answer from every arm was graded correct . None of these tools buys its savings with wrong answers. Every number comes from committed raw transcripts in the Chisle repo https://github.com/JayPokale/Chisle/tree/main/benchmarks/results . | | tasks | caveman | ponytail | Chisle | |---|---|---|---|---| | coding wants working code | 12 | 74% | 59% | 44% | | explanation wants prose | 8 | 103% | 104% | 87% | On coding, Chisle bills 44% of a bare model, a third less than ponytail, the closest thing to a dedicated "lazy code" tool. On explanation prompts both specialists go above 100% : tools built to write less made Claude write more than using nothing. Chisle is the only one that stays under. Split by answer length, the gap widens: on long answers Chisle bills 45% , caveman 79%, ponytail 59%. caveman is genuinely good at compressing prose , and on some short prose prompts it's a hair leaner than Chisle. But it has no judgment about what to build. Asked to "add caching" , it produced three implementations 330 tokens . Chisle gave one @cache decorator and a one-line upgrade path 151 tokens . Its worst day cost 4.2× a bare model. ponytail has the right instinct : smallest thing that works. But it pads prose so much that it backfires: on a "retry logic" prompt it ran 227% of baseline. A tool whose whole job is writing less wrote more than twice as much. Installing both to cover both axes doesn't fix it either: they fight over prose style and double per-session overhead. On one task the pair did worse 605 tokens than Chisle alone 595 . | | prose | code judgment | input / context | publishes failures | |---|---|---|---|---| | caveman | ✅ | ❌ | ❌ | ❌ | | ponytail | ❌ | ✅ | ❌ | ❌ | | Chisle | ✅ | ✅ | ✅ | ✅ | PostToolUse hook trims oversized tool output by Read / Edit / Write are never touched . "Add debounce to a search input that currently fires an API call on every keystroke." Verbatim committed output: useDebounce