why input many code when few code do trick
LLMs and AI coding agents consume massive amounts of tokens reading boilerplate, repetitive syntax, and verbose formatting. CaveCode compresses source code before passing it to AI agents, stripping unnecessary token overhead while preserving critical information. It reduces input token usage by an estimated 80%+ with three configurable compression tiers (all token counts and savings are estimates). This tool is mainly useful for large codebases, large multi-file agentic operations etc.
- 🌐 Web Playground : Try compression modes directly in your browser.
- 📸 Showcase : All 9 languages, ultra → medium → lite.
- 🤖 Agent Protocol : Includes anAGENT.md documentation file for AI coding agents.
Compressed output is an information representation for agents, not executable source code. Agents continue to (and must) read and edit the original source files.
Install via pip:
pip install cavecode
Or install directly from Git:
pip install git+https://github.com/grimm67123/cavecode.git
Or clone and install in editable development mode:
git clone https://github.com/grimm67123/cavecode.git
cd cavecode
pip install -e .
Verify installation:
cavecode version
| Mode | Estimated Token Savings | What it keeps |
|---|---|---|
lite |
~5% – ~10% (estimated) | Full function bodies & code logic |
medium |
~15% – ~20% (estimated) | Function bodies with compressed syntax |
ultra |
~80% – ~85%+ (estimated) | AST structure, signatures & types |
CaveCode supports 9 programming languages with dedicated AST parsers and syntax transformers:
- Python (
.py) - JavaScript (
.js,.jsx,.mjs,.cjs) - TypeScript (
.ts,.tsx) - Rust (
.rs) - Go (
.go) - Java (
.java) - C++ (
.cpp,.cc,.cxx,.hpp) - C# (
.cs) - C (
.c,.h)
Technical reference and operating protocol for AI coding agents using cavecode.
To minimize context window usage, reduce token latency, and prevent token waste, agents follow this protocol:
Read reference & dependency code via cavecode read :
When exploring the codebase, mapping architecture, or checking how external modules, classes, and functions are called, usecavecode read instead of reading large raw files:
cavecode read path/to/file.py -m ultra # Signatures, types, classes (~80% – ~85%+ estimated token savings)
cavecode read path/to/file.ts -m lite # Full logic preserved for logic skims (~25% – ~30% estimated token savings)
cavecode read src/ -m ultra # Scan all files in a directory directly to stdout
cavecode read streams compressed code directly to stdout with zero disk modifications.
2.
Read raw files when actively editing or debugging :cavecode is designed for gathering context on dependencies and interfaces. When you areactively writing code, creating diffs/patches, or performing step-by-step logic debugging in a target file , always use your standard native file read tool to inspect the raw file. This ensures exact indentation, byte-for-byte string matches, and authentic line numbers.
3.
Never use compressed text as a patch or diff target :cavecode normalizes whitespace, formats syntax, and uses shorthand representations. These modified strings do not exist in the raw file. All edits, search-and-replace blocks, and unified diffs must target the original source files (<filename>.<ext> ).
4.
Line numbers from -n are stream-relative :
The-n flag numbers the lines of thecompressed output stream , not the physical lines in the source file. Do not pass line numbers from compressed output into automated file patchers or linters.
5.
Optional on-disk companion files (cavecode compress) :
If you need persistent.cave companion files on disk across multiple turns or sub-commands:
cavecode compress . -m ultra # Generate .cave companion files across directory
cavecode compress file.py -m lite # Generate single file.cave.py
Compressed representations are strictly read-only reference artifacts. Compilers, linters, test runners, and debuggers run exclusively on original source files. When finished, clean them up:
cavecode revert .
Choose the compression tier using --mode or -m (default: ultra):
| Mode | Estimated Token Savings | Logic & Structure | Recommended Agent Use Case |
|---|---|---|---|
ultra(Default) |
~80% – ~85%+ (estimated) | 100% architectural and signature preservation. Strips comments and docstrings; collapses function and method bodies to structural shells (pass /{ ... } ). Keeps all class hierarchies, type annotations, and function definitions. |
Primary mode for agents. Ideal for large repository scanning, architecture discovery, mapping dependencies, and checking API contracts/parameters of files you are not actively modifying. |
lite |
~5% – ~10% (estimated) | 100% of function bodies & logic preserved. Removes docstrings and license blocks, normalizes indentation to 2 spaces, and condenses imports. | Skimming algorithms or internal data flow inside an external module when you need to understand how it works under the hood without burning full estimated token overhead. |
medium |
~15% – ~20% (estimated) | 100% of function bodies preserved. Uses compact keyword density (pub ,priv ,const ), strips debug/info logging calls, and summarizes comments. |
High-density reading across multiple interdependent files when you need a compact overview of logic. |
Reads source file(s) or directory on the fly with AST compression, printed directly to stdout. Zero disk modifications.
cavecode read path/to/file.py
cavecode read path/to/file.ts -m lite
cavecode read path/to/file.py -l 1:50
cavecode read path/to/file.go -n
cavecode read ./src -m ultra
Compresses a source file or an entire directory, creating <name>.cave.<ext> companion files alongside target files. Original files remain 100% untouched.
cavecode compress . -m ultra
cavecode compress ./src -m lite
cavecode compress path/to/file.py -m ultra
Removes generated .cave companion files from a file or directory, restoring a clean raw codebase state.
cavecode revert .
cavecode revert ./src
Calculates estimated token usage and potential savings without writing any files to disk.
cavecode estimate path/to/file.rs -m ultra
cavecode estimate ./src -m lite
Performs SHA-256 cryptographic verification proving target source files have not been modified.
cavecode verify .
CaveCode provides dedicated AST parsers for 9 programming languages:
- Python (
.py), JavaScript (.js,.jsx,.mjs,.cjs), TypeScript (.ts,.tsx) - Rust (
.rs), Go (.go), Java (.java), C++ (.cpp), C# (.cs), C (.c,.h)
Note on Markup & Configs: Markup and configuration formats (HTML, CSS, JSON, YAML, Markdown) do not have AST structures and are passed through untouched. Use standard file reading tools for these formats.