cd /news/ai-tools/caveman-vs-ponytail-vs-chisle-i-benc… · home › topics › ai-tools › article
[ARTICLE · art-146330] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=↑ positive

Caveman vs Ponytail vs Chisle: I benchmarked the Claude Code token-saving plugins on 20 tasks

A developer built Chisle, a Claude Code plugin that trims token usage, and benchmarked it against the caveman and ponytail plugins plus a no-plugin baseline across 20 live tasks and 59+ model runs on Haiku and Sonnet. Chisle cut the total bill to 52% of baseline versus 80% for caveman and 68% for ponytail, with the smallest worst-case blowup (173% vs 424% and 227%) and only one backfire in 20 tasks, while all answers across every arm were graded correct in the July re-verification run.

by read5 min views1 publishedOct 6, 2026

If you use Claude Code, you've probably seen the two popular plugins that promise to cut your token bill: caveman, which makes Claude talk like a caveman, and ponytail, which pushes it to write less code. I built a third one, Chisle, and benchmarked all three against Claude with no plugin at all: 20 live tasks, 59+ model runs on Haiku and Sonnet, with the arms differing only in the injected ruleset.

TL;DR: caveman cut the bill to 80%, ponytail to 68%, and Chisle to 52%, nearly half. Chisle also had the smallest worst day (** 173%** vs 424% and 227%) and backfired once in 20 tasks, where caveman backfired 6 times and ponytail 8.

20 live tasks, billed output vs no plugin total bill average task worst case backfires
no plugin (baseline) 100% 100%
caveman 80% 98% 424% 6 / 20
ponytail 68% 91% 227% 8 / 20
Chisle 52% 69% 173% 1 / 20

In the July re-verification run, every answer from every arm was graded correct. None of these tools buys its savings with wrong answers. Every number comes from committed raw transcripts in the Chisle repo.

tasks caveman ponytail Chisle
coding (wants working code) 12 74% 59% 44%
explanation (wants prose) 8 103% 104% 87%

On coding, Chisle bills 44% of a bare model, a third less than ponytail, the closest thing to a dedicated "lazy code" tool. On explanation prompts both specialists go above 100%: tools built to write less made Claude write more than using nothing. Chisle is the only one that stays under.

Split by answer length, the gap widens: on long answers Chisle bills 45%, caveman 79%, ponytail 59%.

caveman is genuinely good at compressing prose, and on some short prose prompts it's a hair leaner than Chisle. But it has no judgment about what to build. Asked to "add caching", it produced three implementations (330 tokens). Chisle gave one @cache decorator and a one-line upgrade path (** 151 tokens**). Its worst day cost 4.2× a bare model.

ponytail has the right instinct: smallest thing that works. But it pads prose so much that it backfires: on a "retry logic" prompt it ran 227% of baseline. A tool whose whole job is writing less wrote more than twice as much.

Installing both to cover both axes doesn't fix it either: they fight over prose style and double per-session overhead. On one task the pair did worse (605 tokens) than Chisle alone (595).

prose code judgment input / context publishes failures
caveman ✅ ❌ ❌ ❌
ponytail ❌ ✅ ❌ ❌
Chisle ✅ ✅ ✅ ✅

PostToolUse hook trims oversized tool output by Read/ Edit/ Write are never touched). "Add debounce to a search input that currently fires an API call on every keystroke." Verbatim committed output:

useDebounce<T> hook in its own file, then Option 2, Option 3, a comparison table and caveats.setTimeout in the effect you already have, two lines on why, and lodash.debounce if it's already installed."

useEffect(() => {
  const timer = setTimeout(async () => {
    if (query.trim()) { /* fetch */ }
  }, 300);
  return () => clearTimeout(timer);
}, [query]);

Not golfed, just boring: one less file, one less abstraction, same behaviour.

cache task where the bare model wrote an unusually long answer (4,910 tokens, against 375–813 in later runs), and a few cells that an older ruleset example may have primed. Without those cells, Chisle's 20-task total is Chisle is the only tool in this class that publishes the runs where it lost: benchmarks/results.

npx chisle             # installs for Claude Code and any other agents it finds
npx chisle --dry-run   # preview first

Zero dependencies, MIT licensed. The same ruleset ships to Cursor, Codex, Gemini CLI, GitHub Copilot, Windsurf, Cline, OpenCode, Kiro, Antigravity, Hermes and Pi. Claude Code and Pi also get the input-side compressor and live modes (lite, full, ultra). /chisle-audit flags over-engineered code and bloated prose/docs in one ranked report; ponytail's audit is code-only and caveman has none.

Is Chisle better than caveman?

Across 20 tasks: 52% vs 80% of the bare-model bill, worst case 173% vs 424%, 1 vs 6 backfires. On coding prompts 44% vs 74%. caveman is a little leaner on some short prose prompts.

Is Chisle better than ponytail?

Across 20 tasks: 52% vs 68%, worst case 173% vs 227%, 1 vs 8 backfires. On explanation prompts ponytail goes above 100% (104%); Chisle stays at 87%.

Can I install caveman and ponytail together instead?

You can, but they fight over prose style and double the overhead. On one task the pair did worse (605 tokens) than Chisle alone (595).

Does it make answers wrong?

In the July re-verification every answer from every arm graded correct. The rule is necessary, not fewest characters, and safety is never cut.

Where's the raw data?

All transcripts are committed: benchmarks/results. Full tables: docs/benchmarks.md.

Your AI talks less, builds less, reads less, and says more. Like a senior dev who bills by the syllable.

The only tool in this class that publishes the runs where it lost. Here's why.

Built for Claude Code: coding answers come back 33% shorter and 24% cheaper, while caveman and ponytail make them longer · 12 agents · zero dependencies · one command

Chisle is a Claude Code plugin that makes Claude cheaper to run without making it dumber. It cuts what Claude writes: no filler, no hedging, no speculative abstractions, just the smallest code that works. It also cuts what Claude reads: a PostToolUse hook trims oversized tool output by ~46% before it re-enters the context window, where it would be re-billed on every later request. In agent-loop tests Claude with Chisle passes the same tasks as Claude without it. One npx chisle installs it, and the…

Site: chisle.jaypokale.me. If you run your own comparison against caveman or ponytail, I'd like to see it, especially where Chisle loses.

── more in #ai-tools 4 stories · sorted by recency
── more on @claude code 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/caveman-vs-ponytail-…] indexed:0 read:5min 2026-10-06 · —