cd /news/ai-agents/show-hn-cai-218x-fewer-tokens-than-g… · home › topics › ai-agents › article
[ARTICLE · art-146783] src=daidocs.com ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Show HN: Cai:218x fewer tokens than Graphify, 2976x than raw files

Kerneta's .Cai plain-text code graph answers a module-import query in 22 tokens versus Graphify's 4,696 — 218x fewer — and 64,075 tokens for reading httpx's own source, or 2,976x fewer, according to the Show HN post. On the make-this-change task, .Cai used 43x fewer tokens than Graphify while matching or beating Graphify on accuracy across six suites, with the gap widening on larger real repositories. The tool parses source with tree-sitter and stores the graph as flat text files with edges written literally as "caller -> callee", rather than in a graph database.

read14 min views1 publishedOct 7, 2026
Show HN: Cai:218x fewer tokens than Graphify, 2976x than raw files
Image: source

2,976x fewer than reading the code.

A coding agent answers "who calls this, what imports that, what breaks if I change it" by re-reading source. .Cai keeps the whole code graph as plain text and answers from an index instead. Ask what a module imports and it spends 22 tokens where Graphify spends 4,696: 218x fewer, and where reading httpx's own source is 64,075: 2,976x fewer. On the everyday make-this-change task it is 43x leaner than Graphify. Same answers, a fraction of the budget, every figure from a run you can repeat.

What if a code graph were just plain text? #

I kept coming back to one thought. A graph is the right shape for code: who calls what, what imports what, what breaks if I touch this function. But every tool that builds one hides it in a database. And the thing reading it is an LLM, which was trained on text and is happiest reading text. So I tried the stubborn version: keep the whole graph as plain text you can open, grep, and diff in git, and then measure whether that actually holds up against a real graph database. It did, by more than I expected. Here is what I found, with every number from runs you can repeat.

If you only read one paragraph #

Graphify is the popular take on this idea, and it went viral claiming up to about 79x fewer tokens, though that figure is measured on a mixed pile of code, docs and images; on pure code the honest number is smaller. Keeping the graph as plain text, .Cai matched or beat Graphify on accuracy across six suites while cutting the token cost many times over, by 43x on the make-this-change task I care about most and by hundreds of times on structural queries, with the gap widening on bigger real repos. The headline numbers are up top; everything below here is the working, from offline runs you can repeat.

Plain text, shaped like a graph #

.Cai reads your code with a parser (tree-sitter) and writes the structure down as text. A function becomes a line. A call becomes an edge, written literally as caller -> callee. That is the whole trick. The graph is real, the same nodes and edges a graph database would hold, but it lives in flat files you can cat, grep and diff. When the agent asks who calls cart_total, it gets back the edge list, not the file.

The mechanism is simple. A plain-text slice of the graph drops into context more cleanly than a serialized subgraph and costs fewer tokens, because it carries only the edges that answer the question and nothing else. That is the whole bet, and the measurements further down are where it either holds or does not.

Text is the one format every model has seen a trillion times. So maybe the best memory for an LLM is not a database it has to be taught to read, but the plain text it already thinks in.

What the two tools are #

The comparison only means something if both tools start from the same place, and they do: each parses the same source into a graph of functions and calls. What differs is what they hand back and how they store it, which is what the table lays out. The third column throughout is the naive baseline, the raw source files an agent reads when it has no tool at all.

Kerneta .Cai Graphify raw files
What it is plain-text code graph with targeted ops code knowledge graph, queried by traversal the whole source, no index
Store readable text files plus JSONL graph database the files themselves
Answer shape the exact edge list for the question a subgraph of nodes every file, read in full
Needs a model? no, for code (yes only for prose and images) no, for code no

A note on fairness, since it is the first thing a careful reader asks: both tools parse the same source, so neither is graded against its own parser. Gold answers come from the Python standard-library ast (for structural questions) or from the AST-derived set of functions a change touches. That ground truth is independent of both tools.

The headline runs #

Synthetic corpora are easy to game, so the results that matter are on real projects. psf/requests was graded by the ast oracle; the summary suites include a 9-language set, a deep-dependency stress test, and a real-code retrieval test. Accuracy is correct answers; tokens are the average a query puts in front of the model.

Suite Q .Cai Graphify raw files .Cai tok Graphify tok raw tok .Cai vs Graphify .Cai vs raw
Local, 9 languages 94 94/94 83/94 80/94 29 105 215 4x less 7x less
Deep multi-hop (constructed) 11 11/11 10/11 8/11 27 118 188 4x less 7x less
Real-code retrieval 10 10/10 10/10 10/10 24 126 4,818 5x less 202x less
psf/requests (ast oracle) 16 16/16 13/16 16/16 55 689 9,023 13x less 164x less

On requests, Graphify missed three questions a real project exercises that a toy one does not: what the package entry point imports, which module is imported most (it has no "count across the graph" operator), and the full transitive dependency set (its traversal stops at depth two). The raw-file column is the naive baseline an agent falls back to: always correct, because it reads everything, at 9,023 tokens a question against .Cai's 55.

The 9-language set, per language

Language Q .Cai Graphify raw files .Cai tok Graphify tok raw tok .Cai vs Graphify .Cai vs raw
python 11 11/11 9/11 10/11 33 160 344 5x less 10x less
js 11 11/11 9/11 10/11 33 161 258 5x less 8x less
ts 11 11/11 9/11 10/11 33 161 258 5x less 8x less
ruby 11 11/11 6/11 10/11 33 70 196 2x less 6x less
go 10 10/10 10/10 8/10 25 75 148 3x less 6x less
rust 10 10/10 10/10 8/10 25 76 163 3x less 6x less
java 10 10/10 10/10 8/10 26 78 195 3x less 8x less
csharp 10 10/10 10/10 8/10 26 76 194 3x less 8x less
cpp 10 10/10 10/10 8/10 26 77 161 3x less 6x less

go, rust, java, csharp and cpp use their single-file shop call-graph corpora; python, js, ts and ruby use the multi-file shopcart corpora. Tokens are chars/4, averaged per query and applied identically to every system.

A closer look: encode/httpx #

The same repo asked every kind of structural question, with Graphify and the raw-file baseline measured on each. Recall is how much of the true answer set the tool surfaced; the small number is average tokens. The raw column reads every source file, so it always has the answer (100%), at the cost of the whole repo in context. Green is 90% and up, amber is 50 to 89%, red is under 50%.

Structural queries (.Cai vs Graphify vs raw)

Query type n .Cai recall Graphify recall raw recall .Cai tok Graphify tok raw tok .Cai vs Graphify .Cai vs raw
imports 15 100% 91% 100% 22 4,696 64,075 218x less 2,976x less
callers (who calls X) 12 100% 56% 100% 135 1,738 64,075 13x less 474x less
callees (what X calls) 12 100% 45% 100% 139 1,087 64,075 8x less 462x less
transitive deps 10 100% 90% 100% 228 4,736 64,075 21x less 281x less
subclasses 10 100% 100% 100% 40 2,879 64,075 73x less 1,622x less
most-imported module 1 100% 0% 100% 115 7 64,075 GF 0% recall 557x less

.Cai returns the exact edge list, so a caller query is about 135 tokens; Graphify returns node dumps and runs into the thousands, and the raw baseline is the whole repo, 64,075 tokens, on every question. Graphify's recall drops on callers and callees, the two questions an agent asks most when editing, and its most-imported op scores zero because it has no count-across-the-graph operator (it is cheap but wrong). On imports a structural .Cai query is 218x leaner than Graphify and 2,976x leaner than reading the source.

Semantic and lookup retrieval (.Cai vs Graphify)

Query type n .Cai recall .Cai top-1 Graphify recall .Cai tok Graphify tok .Cai vs Graphify
semantic (paraphrase) 40 95% 32% 100% 722 1,841 3x less
lookup 25 92% 56% 100% 425 3,698 9x less

semantic: a behaviour described in natural language, answer is the implementing symbol. lookup: where a symbol is defined. Recall is the gold file appearing in the retrieved set; top-1 is it ranking first. These two use .Cai's hybrid lexical + MiniLM retrieval, so the token counts depend on the embedding model. Graphify leads on recall here; .Cai leads on tokens and matches it closely on recall.

"Make this change" #

The everyday job of a coding agent is editing existing code, and the step that breaks a build is missing a call site. So the real test is: given a change, does the tool surface every function that must be edited? Six real changes in httpx. A blind agent, told nothing in advance, lists the edit set from each tool's output alone; we grade it against the true set of affected functions. A missed caller is a broken build, so recall and complete edit sets are what count.

Tool avg recall complete edit sets avg F1 avg tokens
.Cai 1.00 6 / 6 0.98 38
Graphify 0.64 3 / 6 0.71 1,651
raw files 0.61 1 / 6 0.72 64,075

.Cai found every call site on all six tasks, at 38 tokens each, 43x cheaper than Graphify and 1,686x cheaper than reading the raw files, because its callers op returns the exact edge list rather than a passage dump. Graphify missed call sites on three tasks; reading the raw source got a complete edit set on only one of the six, because the agent has to spot every call site by eye in 64,075 tokens of code.

Hard paraphrase questions #

Twelve behaviour questions about httpx that never name the target symbol, so the agent has to reason from retrieved code. Gold-in-pack is the ceiling (was the right code even retrieved); solved is a blind agent's answer graded against the parser's gold.

Tool gold in pack solved avg tokens
.Cai 12 / 12 12 / 12 907
Graphify 11 / 12 10 / 12 2,746

Gold-in-pack is the solvability ceiling: .Cai retrieved the right code for all twelve, Graphify for eleven. .Cai solved two more of the twelve while spending about 3x fewer tokens.

Scale: .Cai and Graphify on the CPython standard library #

The same "who calls this" question, now on a corpus about 35x bigger than httpx: the CPython standard library, 153 modules, 7,469 symbols. Both tools build the graph offline with no model (.cai: 7,469 symbols, 26,214 edges; Graphify: 11,262 nodes, 19,347 links). The question is whether the per-query cost grows with the repo. It does not, for either, because each answer is the queried symbol's own neighborhood, so cost tracks fan-in rather than store size. Numbers are tokens for one "callers" query, run with each tool's native reverse-traversal op.

callers of (stdlib symbol) .Cai tok Graphify tok .Cai vs Graphify
partial 107 642 6x less
wraps 141 311 2x less
ArgumentParser 233 386 2x less
reduce 64 232 4x less
total_ordering 21 66 3x less
lru_cache 45 62 1x less
namedtuple 20 57 3x less
OrderedDict 20 10 GF smaller
Counter 19 9 GF smaller
deque 19 8 GF smaller
average, 10 symbols 69 178 3x less

Across ten symbols .Cai averages 69 tokens a query to Graphify's 178, about 2.6x leaner, and the gap widens on the high fan-in symbols an agent actually asks about, such as partial and wraps. On symbols with almost no callers (OrderedDict, Counter, deque) Graphify is a few tokens smaller, because .Cai carries a small fixed header per answer; those rows are in the table rather than dropped. The point is the shape, not the winner: at 35x the symbols both still answer in tens to low hundreds of tokens, because the cost is the queried symbol's own edges and not the size of the store. This pass compares the two graph tools only, using each tool's direct callers op, a lighter path than the natural-language queries in the httpx table above, so read the table as a scale curve rather than a restatement of those rows.

The honest margins #

Things a skeptical reader should know, stated plainly rather than buried.

| Gold is tool-independent. Answers come from the Pythonast or the AST impact set, not from either tool, so nothing is graded against itself. | | No API spend. Structural parsing is offline; the blind end-to-end answerer ran on a Claude subscription, not a paid API. | | The tools do different jobs. Some of the token gap is that .Cai returns an exact edge list while Graphify returns a subgraph and the raw baseline returns whole files. That is the point, but it is not a like-for-like payload. | | Where .Cai is weaker: paraphrase recall under many distractors at scale (recovers by widening the search), and it does not yet extract external-library usage structurally ("where does my project call library X"). | | Scope: the constructed deep-hop suite is a stress test, not an official benchmark, and the real-code retrieval suite measures retrieval efficiency, not an issue-resolution agent metric. Corpora run from ~200 to 7,469 symbols; a multi-repo sweep would harden the numbers further. | | Prose and images still need a model, the same boundary Graphify has. Code and structured data are deterministic and offline; document meaning routes to .dai, images to a vision model. |

What is next #

Two things are in motion. It parses 10 languages today (Python, JavaScript, TypeScript, Go, Rust, Java, C, C++, Ruby, C#), and the plan is to expand toward the coverage a graph database like Graphify offers, since each new language is mostly a parser away. And because the store already is the graph, nodes and edges written as text, we are building an interactive visualization that draws the structure straight from the plain .cai files. It exports to mermaid, dot and graphml now, with a proper graph view in progress, so you get the readable source and the picture from one file rather than choosing between them.

Repeat it yourself #

Every figure here comes from result files produced offline. The full replication kit is public: the corpora, the questions and gold, the runners, and the result files, so you can clone it, build the stores, and rebuild every number yourself. The document-memory results, scored against LongMemEval, are on the other tab.

── more in #ai-agents 4 stories · sorted by recency
── more on @kerneta 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/show-hn-cai-218x-few…] indexed:0 read:14min 2026-10-07 · —