{"slug": "cavecode-why-input-many-code-token-when-few-do-trick", "title": "CaveCode – why input many code token when few do trick", "summary": "CaveCode, a pip-installable tool from developer Grimm67123, compresses source code before it reaches AI coding agents, cutting input token usage by an estimated 80%+ in its ultra mode while preserving AST structure, signatures and types. The tool ships three configurable compression tiers — lite (~5%–~10% estimated savings), medium (~15%–~20%) and ultra (~80%–~85%+) — with dedicated AST parsers for 9 languages including Python, JavaScript, TypeScript, Rust, Go, Java, C++, C# and C. CaveCode's documentation states compressed output is an information representation for agents, not executable source code, and that agents must still read and edit the original source files.", "body_md": "**why input many code when few code do trick**\n\nLLMs and AI coding agents consume massive amounts of tokens reading boilerplate, repetitive syntax, and verbose formatting. CaveCode compresses source code before passing it to AI agents, stripping unnecessary token overhead while preserving critical information. It reduces input token usage by an estimated **80%+** with three configurable compression tiers (all token counts and savings are estimates). This tool is mainly useful for large codebases, large multi-file agentic operations etc.\n\n- 🌐 **[Web Playground](https://grimm67123.github.io/cavecode/)** : Try compression modes directly in your browser.\n- 📸 **[Showcase](https://github.com/Grimm67123/cavecode/blob/main/SHOWCASE.md)** : All 9 languages, ultra → medium → lite.\n- 🤖 **Agent Protocol** : Includes an[AGENT.md](https://github.com/Grimm67123/cavecode/blob/main/AGENT.md) documentation file for AI coding agents.\n\n**Compressed output is an information representation for agents, not executable source code.** Agents continue to (and must) read and edit the original source files.\n\nInstall via pip:\n\n```\npip install cavecode\n```\n\nOr install directly from Git:\n\n```\npip install git+https://github.com/grimm67123/cavecode.git\n```\n\nOr clone and install in editable development mode:\n\n```\ngit clone https://github.com/grimm67123/cavecode.git\ncd cavecode\npip install -e .\n```\n\nVerify installation:\n\n```\ncavecode version\n```\n\n| Mode | Estimated Token Savings | What it keeps | \n|---|---|---|\n| **`lite`** | **~5% – ~10% (estimated)** | Full function bodies & code logic | \n| **`medium`** | **~15% – ~20% (estimated)** | Function bodies with compressed syntax | \n| **`ultra`** | **~80% – ~85%+ (estimated)** | AST structure, signatures & types | \n\nCaveCode supports 9 programming languages with dedicated AST parsers and syntax transformers:\n\n- Python (`.py` )\n- JavaScript (`.js` ,`.jsx` ,`.mjs` ,`.cjs` )\n- TypeScript (`.ts` ,`.tsx` )\n- Rust (`.rs` )\n- Go (`.go` )\n- Java (`.java` )\n- C++ (`.cpp` ,`.cc` ,`.cxx` ,`.hpp` )\n- C# (`.cs` )\n- C (`.c` ,`.h` )\n\nTechnical reference and operating protocol for AI coding agents using `cavecode`.\n\nTo minimize context window usage, reduce token latency, and prevent token waste, agents follow this protocol:\n\n1. \n**Read reference & dependency code via `cavecode read`** :\nWhen exploring the codebase, mapping architecture, or checking how external modules, classes, and functions are called, use`cavecode read` instead of reading large raw files:\n\n```\ncavecode read path/to/file.py -m ultra    # Signatures, types, classes (~80% – ~85%+ estimated token savings)\ncavecode read path/to/file.ts -m lite     # Full logic preserved for logic skims (~25% – ~30% estimated token savings)\ncavecode read src/ -m ultra               # Scan all files in a directory directly to stdout\n```\n\n *`cavecode read` streams compressed code directly to stdout with zero disk modifications.*\n2. \n**Read raw files when actively editing or debugging** :`cavecode` is designed for gathering context on dependencies and interfaces. When you are**actively writing code, creating diffs/patches, or performing step-by-step logic debugging in a target file** , always use your standard native file read tool to inspect the raw file. This ensures exact indentation, byte-for-byte string matches, and authentic line numbers.\n3. \n**Never use compressed text as a patch or diff target** :`cavecode` normalizes whitespace, formats syntax, and uses shorthand representations. These modified strings do not exist in the raw file. All edits, search-and-replace blocks, and unified diffs must target the original source files (`<filename>.<ext>` ).\n4. \n**Line numbers from `-n` are stream-relative** :\nThe`-n` flag numbers the lines of the*compressed output stream* , not the physical lines in the source file. Do not pass line numbers from compressed output into automated file patchers or linters.\n5. \n**Optional on-disk companion files (`cavecode compress`)** :\nIf you need persistent`.cave` companion files on disk across multiple turns or sub-commands:\n\n```\ncavecode compress . -m ultra      # Generate .cave companion files across directory\ncavecode compress file.py -m lite # Generate single file.cave.py\n```\n\n Compressed representations are strictly read-only reference artifacts. Compilers, linters, test runners, and debuggers run exclusively on original source files. When finished, clean them up: \n\n```\ncavecode revert .\n```\n\nChoose the compression tier using `--mode` or `-m` (default: `ultra`):\n\n| Mode | Estimated Token Savings | Logic & Structure | Recommended Agent Use Case | \n|---|---|---|---|\n| **`ultra`***(Default)* | **~80% – ~85%+ (estimated)** | **100% architectural and signature preservation.** Strips comments and docstrings; collapses function and method bodies to structural shells (`pass` /`{ ... }` ). Keeps all class hierarchies, type annotations, and function definitions. | **Primary mode for agents.** Ideal for large repository scanning, architecture discovery, mapping dependencies, and checking API contracts/parameters of files you are not actively modifying. | \n| **`lite`** | **~5% – ~10% (estimated)** | **100% of function bodies & logic preserved.** Removes docstrings and license blocks, normalizes indentation to 2 spaces, and condenses imports. | Skimming algorithms or internal data flow inside an external module when you need to understand how it works under the hood without burning full estimated token overhead. | \n| **`medium`** | **~15% – ~20% (estimated)** | **100% of function bodies preserved.** Uses compact keyword density (`pub` ,`priv` ,`const` ), strips debug/info logging calls, and summarizes comments. | High-density reading across multiple interdependent files when you need a compact overview of logic. | \n\nReads source file(s) or directory on the fly with AST compression, printed directly to stdout. **Zero disk modifications.**\n\n```\n# Read a single file in ultra mode (default: ~80%+ estimated savings)\ncavecode read path/to/file.py\n\n# Read a single file in lite mode (keeps full function implementations)\ncavecode read path/to/file.ts -m lite\n\n# Read a specific line range of compressed output\ncavecode read path/to/file.py -l 1:50\n\n# Read with line numbers (stream-relative)\ncavecode read path/to/file.go -n\n\n# Scan all supported files in a directory to stdout\ncavecode read ./src -m ultra\n```\n\nCompresses a source file or an entire directory, creating `<name>.cave.<ext>` companion files alongside target files. Original files remain **100% untouched**.\n\n```\n# Generate .cave files across entire repository\ncavecode compress . -m ultra\n\n# Generate .cave files for a directory in lite mode\ncavecode compress ./src -m lite\n\n# Compress a single file to a .cave companion file\ncavecode compress path/to/file.py -m ultra\n```\n\nRemoves generated `.cave` companion files from a file or directory, restoring a clean raw codebase state.\n\n```\n# Clean up all .cave files in the project\ncavecode revert .\n\n# Clean up .cave files in a specific folder\ncavecode revert ./src\n```\n\nCalculates estimated token usage and potential savings without writing any files to disk.\n\n```\ncavecode estimate path/to/file.rs -m ultra\ncavecode estimate ./src -m lite\n```\n\nPerforms SHA-256 cryptographic verification proving target source files have not been modified.\n\n```\ncavecode verify .\n```\n\nCaveCode provides dedicated AST parsers for 9 programming languages:\n\n- Python (`.py` ), JavaScript (`.js` ,`.jsx` ,`.mjs` ,`.cjs` ), TypeScript (`.ts` ,`.tsx` )\n- Rust (`.rs` ), Go (`.go` ), Java (`.java` ), C++ (`.cpp` ), C# (`.cs` ), C (`.c` ,`.h` )\n\n*Note on Markup & Configs:* Markup and configuration formats (HTML, CSS, JSON, YAML, Markdown) do not have AST structures and are passed through untouched. Use standard file reading tools for these formats.", "url": "https://wpnews.pro/news/cavecode-why-input-many-code-token-when-few-do-trick", "canonical_source": "https://github.com/Grimm67123/cavecode", "published_at": "2026-09-30 12:30:58+00:00", "updated_at": "2026-09-30 12:49:35.897373+00:00", "lang": "en", "topics": ["ai-agents", "developer-tools", "ai-tools", "large-language-models"], "entities": ["CaveCode", "Grimm67123", "Python", "JavaScript", "TypeScript", "Rust", "Go", "Java"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/cavecode-why-input-many-code-token-when-few-do-trick", "markdown": "https://wpnews.pro/news/cavecode-why-input-many-code-token-when-few-do-trick.md", "text": "https://wpnews.pro/news/cavecode-why-input-many-code-token-when-few-do-trick.txt", "jsonld": "https://wpnews.pro/news/cavecode-why-input-many-code-token-when-few-do-trick.jsonld"}}