# Token Compression for Coding Agents: Fine-Tuned Middleware Cuts Codex Costs by 30%

> Source: <https://dev.to/mech_app_ai/token-compression-for-coding-agents-fine-tuned-middleware-cuts-codex-costs-by-30-40f7>
> Published: 2026-10-01 00:06:15+00:00

Coding agents hit a cost wall when tool-call output bloats context windows. A Show HN project tackles this with a fine-tuned compression model that sits between agent output and model input, trimming tokens by 29.6% without breaking KV cache or multi-turn reasoning. The project exposes a pattern: developers are building custom middleware layers to manage the economic pressure points in agentic workflows.

The team behind this project maxed out their Codex subscription and burned $700 per day per person on API calls. The culprit was not the agent's reasoning steps but the tool-call results that get fed back into the model. File retrieval, test output, and error traces accumulate fast. Each round trip inflates the input token count, and cache misses compound the cost.

Coding agents differ from conversational agents in context shape. A chat agent might reference a few messages. A coding agent drags in file trees, diff output, stack traces, and test results. The context window fills with structured data that the model needs for the next step but that also contains redundancy.

The solution is a local proxy that wraps Codex and intercepts tool-call results before they return to the model. The proxy runs a fine-tuned Qwen model trained to preserve agent trajectory while removing redundant information from tool outputs.

**Execution flow:**

The fine-tuning objective is trajectory preservation. The model learns which parts of tool output the agent needs for subsequent reasoning steps and which parts are noise. File retrieval accuracy and context-heavy tasks see the biggest gains.

The CLI is open source and installs as a shell wrapper around Codex. It runs on by default, and you can disable it with `codex --uncompress` when you need full output.

**Key design choices:**

`response.usage` to measure savings. You can run `savings` to see cumulative token reduction.
**Installation:**

```
curl -fsSL https://install.everestagi.com/install.sh | sh && \
  source ~/.config/everest/shell.sh
```

The proxy runs as a local service and intercepts API calls. You point your Codex client at the proxy endpoint instead of the OpenAI endpoint.

| Dimension | Benefit | Risk | 
|---|---|---|
| **Cost** | 29.6% token reduction, lower API spend | Compression model adds latency and local compute overhead | 
| **Fidelity** | Fine-tuned on agent trajectories to preserve reasoning steps | May remove information the agent needs for edge cases or complex tasks | 
| **Cache** | Operates before model input, so KV cache stays valid | If compression changes output shape, downstream tools may break | 
| **Privacy** | Local proxy, no data retention | Requires trust in the proxy binary and shell script installer | 
| **Portability** | Works with any Codex-compatible client | Tied to Codex and Astra workflows, not general-purpose | 

**Failure modes:**

The proxy exposes a `savings` command that shows cumulative token reduction. This is useful for tracking ROI, but it does not expose per-call compression ratios or failure cases.

**What you cannot see:**

For production use, you would want structured logs that capture input tokens, output tokens, compression ratio, and agent success rate per task type. You would also want a fallback mode that disables compression if the agent fails repeatedly on a specific task.

**Good fit:**

**Poor fit:**

This project demonstrates a practical response to the economic pressure points in agentic workflows. The architecture is sound: a local proxy with a fine-tuned compression model preserves KV cache and reduces input tokens without breaking multi-turn reasoning. The 29.6% token reduction is meaningful for teams burning hundreds of dollars per day on API calls.

The trade-off is fidelity. Fine-tuning helps, but compression always risks removing information the agent needs. For production use, you would want observability that tracks compression ratio per task type and a fallback mode that disables compression when the agent fails.

Use this if you are optimizing for cost and can tolerate occasional compression errors. Avoid it if you need exact tool-call output or operate in environments where local proxies are not allowed. The pattern is worth watching: as agents move from prototypes to daily workflows, cost-optimization middleware will become a standard layer in the stack.
