# Caveman vs Ponytail vs Chisle: I benchmarked the Claude Code token-saving plugins on 20 tasks

> Source: <https://dev.to/jaypokale/caveman-vs-ponytail-vs-chisle-i-benchmarked-the-claude-code-token-saving-plugins-on-20-tasks-bg8>
> Published: 2026-10-06 20:14:43+00:00

If you use **Claude Code**, you've probably seen the two popular plugins that promise to cut your token bill: **[caveman](https://github.com/JuliusBrussee/caveman)**, which makes Claude talk like a caveman, and **[ponytail](https://github.com/dietrichgebert/ponytail)**, which pushes it to write less code. I built a third one, **[Chisle](https://chisle.jaypokale.me)**, and benchmarked all three against Claude with *no plugin at all*: **20 live tasks, 59+ model runs** on Haiku and Sonnet, with the arms differing only in the injected ruleset.

**TL;DR:** caveman cut the bill to **80%**, ponytail to **68%**, and **Chisle to 52%**, nearly half. Chisle also had the smallest worst day (** 173%** vs **424%** and **227%**) and backfired once in 20 tasks, where caveman backfired 6 times and ponytail 8.

| 20 live tasks, billed output vs no plugin | total bill | average task | worst case | backfires | 
|---|---|---|---|---|
| no plugin (baseline) | 100% | 100% |  |  | 
| caveman | 80% | 98% | **424%** | 6 / 20 | 
| ponytail | 68% | 91% | 227% | 8 / 20 | 
| **Chisle** | **52%** | **69%** | **173%** | **1 / 20** | 

In the July re-verification run, **every answer from every arm was graded correct**. None of these tools buys its savings with wrong answers. Every number comes from committed raw transcripts in the [Chisle repo](https://github.com/JayPokale/Chisle/tree/main/benchmarks/results).

|  | tasks | caveman | ponytail | **Chisle** | 
|---|---|---|---|---|
| **coding** (wants working code) | 12 | 74% | 59% | **44%** | 
| **explanation** (wants prose) | 8 | 103% | 104% | **87%** | 

On coding, Chisle bills **44%** of a bare model, a third less than ponytail, the closest thing to a dedicated "lazy code" tool. On explanation prompts both specialists go **above 100%**: tools built to write less made Claude write *more* than using nothing. Chisle is the only one that stays under.

Split by answer length, the gap widens: on **long answers Chisle bills 45%**, caveman 79%, ponytail 59%.

**caveman is genuinely good at compressing prose**, and on some short prose prompts it's a hair leaner than Chisle. But it has no judgment about *what* to build. Asked to *"add caching"*, it produced three implementations (**330 tokens**). Chisle gave one `@cache` decorator and a one-line upgrade path (** 151 tokens**). Its worst day cost **4.2×** a bare model.

**ponytail has the right instinct**: smallest thing that works. But it pads prose so much that it backfires: on a "retry logic" prompt it ran **227%** of baseline. A tool whose whole job is writing less wrote more than twice as much.

Installing **both** to cover both axes doesn't fix it either: they fight over prose style and double per-session overhead. On one task the pair did *worse* (605 tokens) than Chisle alone (595).

|  | prose | code judgment | input / context | publishes failures | 
|---|---|---|---|---|
| caveman | ✅ | ❌ | ❌ | ❌ | 
| ponytail | ❌ | ✅ | ❌ | ❌ | 
| **Chisle** | ✅ | ✅ | ✅ | ✅ | 

`PostToolUse` hook trims oversized tool output by `Read`/` Edit`/` Write` are never touched).
*"Add debounce to a search input that currently fires an API call on every keystroke."* Verbatim committed output:

`useDebounce<T>` hook in its own file, then Option 2, Option 3, a comparison table and caveats.`setTimeout` in the effect you already have, two lines on why, and `lodash.debounce` if it's already installed."

``` js
useEffect(() => {
  const timer = setTimeout(async () => {
    if (query.trim()) { /* fetch */ }
  }, 300);
  return () => clearTimeout(timer);
}, [query]);
```

Not golfed, just boring: one less file, one less abstraction, same behaviour.

`cache` task where the bare model wrote an unusually long answer (4,910 tokens, against 375–813 in later runs), and a few cells that an older ruleset example may have primed. Without those cells, Chisle's 20-task total is Chisle is the only tool in this class that publishes the runs where it lost: [benchmarks/results](https://github.com/JayPokale/Chisle/tree/main/benchmarks/results).

```
npx chisle             # installs for Claude Code and any other agents it finds
npx chisle --dry-run   # preview first
```

Zero dependencies, MIT licensed. The same ruleset ships to **Cursor, Codex, Gemini CLI, GitHub Copilot, Windsurf, Cline, OpenCode, Kiro, Antigravity, Hermes and Pi**. Claude Code and Pi also get the input-side compressor and live modes (`lite`, `full`, `ultra`). `/chisle-audit` flags over-engineered code *and* bloated prose/docs in one ranked report; ponytail's audit is code-only and caveman has none.

**Is Chisle better than caveman?**

Across 20 tasks: **52% vs 80%** of the bare-model bill, worst case **173% vs 424%**, **1 vs 6** backfires. On coding prompts **44% vs 74%**. caveman is a little leaner on some short prose prompts.

**Is Chisle better than ponytail?**

Across 20 tasks: **52% vs 68%**, worst case **173% vs 227%**, **1 vs 8** backfires. On explanation prompts ponytail goes above 100% (104%); Chisle stays at 87%.

**Can I install caveman and ponytail together instead?**

You can, but they fight over prose style and double the overhead. On one task the pair did worse (605 tokens) than Chisle alone (595).

**Does it make answers wrong?**

In the July re-verification every answer from every arm graded correct. The rule is *necessary*, not *fewest characters*, and safety is never cut.

**Where's the raw data?**

All transcripts are committed: [benchmarks/results](https://github.com/JayPokale/Chisle/tree/main/benchmarks/results). Full tables: [docs/benchmarks.md](https://github.com/JayPokale/Chisle/blob/main/docs/benchmarks.md).

*Your AI talks less, builds less, reads less, and says more. Like a senior dev who bills by the syllable.*

*The only tool in this class that publishes the runs where it lost. [Here's why.](https://jaypokale.me/writing/chisle-benchmarks-it-loses)*

**Built for Claude Code: coding answers come back 33% shorter and 24% cheaper, while caveman and ponytail make them longer · 12 agents · zero dependencies · one command**

Chisle is a Claude Code plugin that makes Claude cheaper to run without making it dumber. It cuts what Claude writes: no filler, no hedging, no speculative abstractions, just the smallest code that works. It also cuts what Claude reads: a `PostToolUse` hook trims oversized tool output by ~46% before it re-enters the context window, where it would be re-billed on every later request. In agent-loop tests Claude with Chisle passes the same tasks as Claude without it. One `npx chisle` installs it, and the…

Site: **[chisle.jaypokale.me](https://chisle.jaypokale.me)**. If you run your own comparison against caveman or ponytail, I'd like to see it, especially where Chisle loses.
