CaveCode – why input many code token when few do trick CaveCode, a pip-installable tool from developer Grimm67123, compresses source code before it reaches AI coding agents, cutting input token usage by an estimated 80%+ in its ultra mode while preserving AST structure, signatures and types. The tool ships three configurable compression tiers — lite (~5%–~10% estimated savings), medium (~15%–~20%) and ultra (~80%–~85%+) — with dedicated AST parsers for 9 languages including Python, JavaScript, TypeScript, Rust, Go, Java, C++, C# and C. CaveCode's documentation states compressed output is an information representation for agents, not executable source code, and that agents must still read and edit the original source files. why input many code when few code do trick LLMs and AI coding agents consume massive amounts of tokens reading boilerplate, repetitive syntax, and verbose formatting. CaveCode compresses source code before passing it to AI agents, stripping unnecessary token overhead while preserving critical information. It reduces input token usage by an estimated 80%+ with three configurable compression tiers all token counts and savings are estimates . This tool is mainly useful for large codebases, large multi-file agentic operations etc. - 🌐 Web Playground https://grimm67123.github.io/cavecode/ : Try compression modes directly in your browser. - 📸 Showcase https://github.com/Grimm67123/cavecode/blob/main/SHOWCASE.md : All 9 languages, ultra → medium → lite. - 🤖 Agent Protocol : Includes an AGENT.md https://github.com/Grimm67123/cavecode/blob/main/AGENT.md documentation file for AI coding agents. Compressed output is an information representation for agents, not executable source code. Agents continue to and must read and edit the original source files. Install via pip: pip install cavecode Or install directly from Git: pip install git+https://github.com/grimm67123/cavecode.git Or clone and install in editable development mode: git clone https://github.com/grimm67123/cavecode.git cd cavecode pip install -e . Verify installation: cavecode version | Mode | Estimated Token Savings | What it keeps | |---|---|---| | lite | ~5% – ~10% estimated | Full function bodies & code logic | | medium | ~15% – ~20% estimated | Function bodies with compressed syntax | | ultra | ~80% – ~85%+ estimated | AST structure, signatures & types | CaveCode supports 9 programming languages with dedicated AST parsers and syntax transformers: - Python .py - JavaScript .js , .jsx , .mjs , .cjs - TypeScript .ts , .tsx - Rust .rs - Go .go - Java .java - C++ .cpp , .cc , .cxx , .hpp - C .cs - C .c , .h Technical reference and operating protocol for AI coding agents using cavecode . To minimize context window usage, reduce token latency, and prevent token waste, agents follow this protocol: 1. Read reference & dependency code via cavecode read : When exploring the codebase, mapping architecture, or checking how external modules, classes, and functions are called, use cavecode read instead of reading large raw files: cavecode read path/to/file.py -m ultra Signatures, types, classes ~80% – ~85%+ estimated token savings cavecode read path/to/file.ts -m lite Full logic preserved for logic skims ~25% – ~30% estimated token savings cavecode read src/ -m ultra Scan all files in a directory directly to stdout cavecode read streams compressed code directly to stdout with zero disk modifications. 2. Read raw files when actively editing or debugging : cavecode is designed for gathering context on dependencies and interfaces. When you are actively writing code, creating diffs/patches, or performing step-by-step logic debugging in a target file , always use your standard native file read tool to inspect the raw file. This ensures exact indentation, byte-for-byte string matches, and authentic line numbers. 3. Never use compressed text as a patch or diff target : cavecode normalizes whitespace, formats syntax, and uses shorthand representations. These modified strings do not exist in the raw file. All edits, search-and-replace blocks, and unified diffs must target the original source files