cd /news/developer-tools/tokencompress-a-sub-2ms-go-cli-and-m… · home topics developer-tools article
[ARTICLE · art-98486] src=github.com ↗ pub= topic=developer-tools verified=true sentiment=↑ positive

Tokencompress A sub-2ms Go CLI and MCP sidecar that prunes AI agent tool context

Tokencompress, a zero-dependency Go CLI and MCP sidecar, cuts AI agent tool context token consumption by 60% to 80% without a second LLM summarization turn, according to its developer. The tool prunes raw tool outputs such as JSON, terminal logs, and HTML before they enter an agent's context window, with benchmarks showing reductions of 93.2% for a 500-item JSON array, 87.1% for a Python stack trace, and 88.3% for an HTML web scrape. It is available under the MIT License and can be integrated with Claude Desktop via MCP.

read2 min views1 publishedAug 16, 2026
Tokencompress A sub-2ms Go CLI and MCP sidecar that prunes AI agent tool context
Image: Michielbdejong (auto-discovered)

A zero-dependency, sub-millisecond Go CLI and MCP sidecar that prunes raw tool outputs (JSON, terminal logs, HTML) before they enter your AI agent's context window.

Cut LLM context token consumption by 60% to 80% without using a second LLM summarization turn.

When AI agents execute tools (calling APIs, running terminal commands, or scraping web pages), they receive thousands of lines of raw, unparsed data:

Massive JSON Payloads: A 1,000-item array floods the prompt with 15,000+ tokens.Verbose Stack Traces: Framework noise and node_modules paths drown out the root exception.Raw HTML: CSS, scripts, and navigation menus bloat the context window.

This leads to high API costs, slower response times, and context rot—where agents hallucinate or repeat tool calls mid-task.

acts as a high-speed, deterministic filter between your tools and your model:

JSON Truncation: Keeps representative schema examples and replaces redundant array items with metadata (: N).Log Pruning: Strips internal framework paths and returns only the root error message, user file paths, and execution lines.HTML Cleaning: Strips scripts, styles, and navigation elements, converting content to readable text/markdown.Duplicate Loop Detection: Hashes tool outputs per session and prepends a warning header if an agent receives duplicate tool results twice in a row.

make build sudo mv tokencompress /usr/local/bin/

cat large_response.json | tokencompress --mode json

cat app.log | tokencompress --mode log --log-internal-marker "mycompany/internal"

curl -s [https://example.com](https://example.com) | tokencompress --mode html

Add directly to your Claude Desktop config ():

{ "mcpServers": { "tokencompress": { "command": "/usr/local/bin/tokencompress", "args": ["--mode", "mcp"] } } }

| Input Payload | Raw Token Count | Compressed Token Count | Token Reduction | Execution Time |

|---|---|---|---|---|
JSON Array (500 items) |

~12,500 tokens | ~850 tokens | 93.2% | | Python Stack Trace | ~3,200 tokens | ~410 tokens | 87.1% | | HTML Web Scrape | ~18,000 tokens | ~2,100 tokens | 88.3% |

Need a custom high-performance MCP proxy, specialized parsers for enterprise tools, or custom token-optimization setups for your agent infrastructure?

Custom Setup Gigs: 50 - 00 per custom integration setup.Custom Log/AST Parsers: 50 per custom domain parser.Contact: Open an issue on this repository or contact directly via GitHub profile.

MIT License. Free to use in open-source and commercial agent setups.

── more in #developer-tools 4 stories · sorted by recency
── more on @tokencompress 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/tokencompress-a-sub-…] indexed:0 read:2min 2026-08-16 ·