{"slug": "caveman-vs-ponytail-vs-chisle-i-benchmarked-the-claude-code-token-saving-plugins", "title": "Caveman vs Ponytail vs Chisle: I benchmarked the Claude Code token-saving plugins on 20 tasks", "summary": "A developer built Chisle, a Claude Code plugin that trims token usage, and benchmarked it against the caveman and ponytail plugins plus a no-plugin baseline across 20 live tasks and 59+ model runs on Haiku and Sonnet. Chisle cut the total bill to 52% of baseline versus 80% for caveman and 68% for ponytail, with the smallest worst-case blowup (173% vs 424% and 227%) and only one backfire in 20 tasks, while all answers across every arm were graded correct in the July re-verification run.", "body_md": "If you use **Claude Code**, you've probably seen the two popular plugins that promise to cut your token bill: **[caveman](https://github.com/JuliusBrussee/caveman)**, which makes Claude talk like a caveman, and **[ponytail](https://github.com/dietrichgebert/ponytail)**, which pushes it to write less code. I built a third one, **[Chisle](https://chisle.jaypokale.me)**, and benchmarked all three against Claude with *no plugin at all*: **20 live tasks, 59+ model runs** on Haiku and Sonnet, with the arms differing only in the injected ruleset.\n\n**TL;DR:** caveman cut the bill to **80%**, ponytail to **68%**, and **Chisle to 52%**, nearly half. Chisle also had the smallest worst day (** 173%** vs **424%** and **227%**) and backfired once in 20 tasks, where caveman backfired 6 times and ponytail 8.\n\n| 20 live tasks, billed output vs no plugin | total bill | average task | worst case | backfires | \n|---|---|---|---|---|\n| no plugin (baseline) | 100% | 100% |  |  | \n| caveman | 80% | 98% | **424%** | 6 / 20 | \n| ponytail | 68% | 91% | 227% | 8 / 20 | \n| **Chisle** | **52%** | **69%** | **173%** | **1 / 20** | \n\nIn the July re-verification run, **every answer from every arm was graded correct**. None of these tools buys its savings with wrong answers. Every number comes from committed raw transcripts in the [Chisle repo](https://github.com/JayPokale/Chisle/tree/main/benchmarks/results).\n\n|  | tasks | caveman | ponytail | **Chisle** | \n|---|---|---|---|---|\n| **coding** (wants working code) | 12 | 74% | 59% | **44%** | \n| **explanation** (wants prose) | 8 | 103% | 104% | **87%** | \n\nOn coding, Chisle bills **44%** of a bare model, a third less than ponytail, the closest thing to a dedicated \"lazy code\" tool. On explanation prompts both specialists go **above 100%**: tools built to write less made Claude write *more* than using nothing. Chisle is the only one that stays under.\n\nSplit by answer length, the gap widens: on **long answers Chisle bills 45%**, caveman 79%, ponytail 59%.\n\n**caveman is genuinely good at compressing prose**, and on some short prose prompts it's a hair leaner than Chisle. But it has no judgment about *what* to build. Asked to *\"add caching\"*, it produced three implementations (**330 tokens**). Chisle gave one `@cache` decorator and a one-line upgrade path (** 151 tokens**). Its worst day cost **4.2×** a bare model.\n\n**ponytail has the right instinct**: smallest thing that works. But it pads prose so much that it backfires: on a \"retry logic\" prompt it ran **227%** of baseline. A tool whose whole job is writing less wrote more than twice as much.\n\nInstalling **both** to cover both axes doesn't fix it either: they fight over prose style and double per-session overhead. On one task the pair did *worse* (605 tokens) than Chisle alone (595).\n\n|  | prose | code judgment | input / context | publishes failures | \n|---|---|---|---|---|\n| caveman | ✅ | ❌ | ❌ | ❌ | \n| ponytail | ❌ | ✅ | ❌ | ❌ | \n| **Chisle** | ✅ | ✅ | ✅ | ✅ | \n\n`PostToolUse` hook trims oversized tool output by `Read`/` Edit`/` Write` are never touched).\n*\"Add debounce to a search input that currently fires an API call on every keystroke.\"* Verbatim committed output:\n\n`useDebounce<T>` hook in its own file, then Option 2, Option 3, a comparison table and caveats.`setTimeout` in the effect you already have, two lines on why, and `lodash.debounce` if it's already installed.\"\n\n``` js\nuseEffect(() => {\n  const timer = setTimeout(async () => {\n    if (query.trim()) { /* fetch */ }\n  }, 300);\n  return () => clearTimeout(timer);\n}, [query]);\n```\n\nNot golfed, just boring: one less file, one less abstraction, same behaviour.\n\n`cache` task where the bare model wrote an unusually long answer (4,910 tokens, against 375–813 in later runs), and a few cells that an older ruleset example may have primed. Without those cells, Chisle's 20-task total is Chisle is the only tool in this class that publishes the runs where it lost: [benchmarks/results](https://github.com/JayPokale/Chisle/tree/main/benchmarks/results).\n\n```\nnpx chisle             # installs for Claude Code and any other agents it finds\nnpx chisle --dry-run   # preview first\n```\n\nZero dependencies, MIT licensed. The same ruleset ships to **Cursor, Codex, Gemini CLI, GitHub Copilot, Windsurf, Cline, OpenCode, Kiro, Antigravity, Hermes and Pi**. Claude Code and Pi also get the input-side compressor and live modes (`lite`, `full`, `ultra`). `/chisle-audit` flags over-engineered code *and* bloated prose/docs in one ranked report; ponytail's audit is code-only and caveman has none.\n\n**Is Chisle better than caveman?**\n\nAcross 20 tasks: **52% vs 80%** of the bare-model bill, worst case **173% vs 424%**, **1 vs 6** backfires. On coding prompts **44% vs 74%**. caveman is a little leaner on some short prose prompts.\n\n**Is Chisle better than ponytail?**\n\nAcross 20 tasks: **52% vs 68%**, worst case **173% vs 227%**, **1 vs 8** backfires. On explanation prompts ponytail goes above 100% (104%); Chisle stays at 87%.\n\n**Can I install caveman and ponytail together instead?**\n\nYou can, but they fight over prose style and double the overhead. On one task the pair did worse (605 tokens) than Chisle alone (595).\n\n**Does it make answers wrong?**\n\nIn the July re-verification every answer from every arm graded correct. The rule is *necessary*, not *fewest characters*, and safety is never cut.\n\n**Where's the raw data?**\n\nAll transcripts are committed: [benchmarks/results](https://github.com/JayPokale/Chisle/tree/main/benchmarks/results). Full tables: [docs/benchmarks.md](https://github.com/JayPokale/Chisle/blob/main/docs/benchmarks.md).\n\n*Your AI talks less, builds less, reads less, and says more. Like a senior dev who bills by the syllable.*\n\n*The only tool in this class that publishes the runs where it lost. [Here's why.](https://jaypokale.me/writing/chisle-benchmarks-it-loses)*\n\n**Built for Claude Code: coding answers come back 33% shorter and 24% cheaper, while caveman and ponytail make them longer · 12 agents · zero dependencies · one command**\n\nChisle is a Claude Code plugin that makes Claude cheaper to run without making it dumber. It cuts what Claude writes: no filler, no hedging, no speculative abstractions, just the smallest code that works. It also cuts what Claude reads: a `PostToolUse` hook trims oversized tool output by ~46% before it re-enters the context window, where it would be re-billed on every later request. In agent-loop tests Claude with Chisle passes the same tasks as Claude without it. One `npx chisle` installs it, and the…\n\nSite: **[chisle.jaypokale.me](https://chisle.jaypokale.me)**. If you run your own comparison against caveman or ponytail, I'd like to see it, especially where Chisle loses.", "url": "https://wpnews.pro/news/caveman-vs-ponytail-vs-chisle-i-benchmarked-the-claude-code-token-saving-plugins", "canonical_source": "https://dev.to/jaypokale/caveman-vs-ponytail-vs-chisle-i-benchmarked-the-claude-code-token-saving-plugins-on-20-tasks-bg8", "published_at": "2026-10-06 20:14:43+00:00", "updated_at": "2026-10-06 20:18:16.327511+00:00", "lang": "en", "topics": ["ai-tools", "developer-tools", "large-language-models", "ai-agents"], "entities": ["Claude Code", "Chisle", "caveman", "ponytail", "Anthropic", "Haiku", "Sonnet", "Jay Pokale"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/caveman-vs-ponytail-vs-chisle-i-benchmarked-the-claude-code-token-saving-plugins", "markdown": "https://wpnews.pro/news/caveman-vs-ponytail-vs-chisle-i-benchmarked-the-claude-code-token-saving-plugins.md", "text": "https://wpnews.pro/news/caveman-vs-ponytail-vs-chisle-i-benchmarked-the-claude-code-token-saving-plugins.txt", "jsonld": "https://wpnews.pro/news/caveman-vs-ponytail-vs-chisle-i-benchmarked-the-claude-code-token-saving-plugins.jsonld"}}