Show HN: Sonde, a local code graph for AI agents that refuses to guess Sonde, a local code-context engine for AI coding agents, claims to match agentic search recall on structural tasks while using 3x less context, 8x fewer tool calls, and 147x less wall-clock time, based on benchmarks on a 19,409-line TypeScript repository. The tool, published on GitHub by Cheppu Labs, indexes repositories into a symbol-level graph in SQLite and exposes three MCP tools, but it scores 0.00 on behavioral queries without shared vocabulary. A local code-context engine for AI coding agents. Sonde indexes a TypeScript, Python, or Swift repository into a symbol-level graph in SQLite and exposes three MCP tools — find symbols , query graph , and get impact radius — so an agent can answer who calls this , what breaks if I change it , and which tests relate to it in one call instead of a search loop. The honest claim is not "finds what grep cannot". A competent agentic search loop finds the same structural evidence — we measured it, on a real 19,409-line repository, and it scored 1.000 recall on every task. The claim is the same answers for a fraction of the cost, inside a budget : | On real production TypeScript | Sonde | Agentic search | |---|---|---| | Recall on structural tasks | 1.00 | 1.00 | | Tool calls | 1.0 | 8.0 | | Context tokens | 1,262 | 3,621 | | Latency | 263 ms | 38,602 ms | | Runs that blew the token budget | 0 of 6 | 3 of 6 | Sonde matches the baseline's answers on every structural task while using ~3× less context, 8× fewer calls, and ~147× less wall-clock time — and it never exceeds the caller's token budget, because the packer truncates to it by construction. That is the trade being offered. Where it loses. Behavioural queries with no shared vocabulary — "where is the retry backoff decided?" — score 0.00. Local semantic retrieval was built and measured and does not fix it see the design doc §2.2 ; the capability is therefore not claimed. If your questions are mostly of that shape, an agentic search loop is the better tool today. Numbers are reproducible: npm run bench:fixture && npm run bench:large . Full results in BENCHMARK-LARGE.md /anishmoncivarghese/sonde/blob/main/BENCHMARK-LARGE.md and BENCHMARK.md /anishmoncivarghese/sonde/blob/main/BENCHMARK.md . npm install -g @cheppulabs/sonde cd your-project sonde init sonde init indexes the repository and registers sonde as an MCP server in this project's .mcp.json , asking before it writes anything skip the prompt with sonde init --yes . It never touches an .mcp.json it can't safely merge into — an existing sonde entry that differs from what init would write is left alone and reported, not overwritten. For Python repositories, use sonde init --resolve ; the default tree-sitter tier did not pass the project's placement gate for structural queries. Equivalent by hand, if you'd rather see every step: sonde index . then add to .mcp.json: { "mcpServers": { "sonde": { "command": "sonde", "args": "mcp", "serve", "." } } } No account or hosted service is required. sonde doc writes an ARCHITECTURE.md /anishmoncivarghese/sonde/blob/main/ARCHITECTURE.md describing the repository's modules, how they depend on each other, and what each one exposes — generated from the graph, so it reports what the code actually does rather than what someone remembered. sonde doc write ARCHITECTURE.md sonde doc --stdout print it instead sonde doc --check fail if it is out of date for CI sonde doc --module src/store symbol-level detail, never committed It is meant to be committed and regenerated, not hand-edited. Regeneration is byte-identical when nothing changed, so it does not churn your diffs; when two branches both regenerate it, resolve the conflict by running sonde doc again rather than merging by hand. It refuses to overwrite an ARCHITECTURE.md it did not generate. Three things it deliberately does not do: It does not draw a dependency it cannot evidence. Module pairs that merely share symbol names are excluded and counted separately. On this repository two adapters share the filenames symbols.ts , parser.ts and references.ts , which manufactured 62 heuristic "references" between modules that never import each other — once the second-heaviest arrow in the diagram. It does not pretend the diagram is complete. The diagram shows the heaviest dependencies and states how many it omitted; the table below it goes further, and --module has the rest. A diagram containing every dependency is unreadable and therefore shows nothing. It does not claim to be current when it is not. The header names the commit it describes and warns when files have changed since. Never returns stale source bytes. Whenever a response includes source, Sonde re-reads and re-hashes the indexed byte range before returning it spec §8.1, Guarantee A . Always reports structural drift , rather than claiming completeness it cannot verify spec §8.1, Guarantee B . sonde status shows the same drift and tier distribution carried by tool response envelopes. Every edge is tier-labelled by how it was found — COMPILER resolved exactly by a bundled type checker under --resolve : the TypeScript compiler for TypeScript, pyright for Python , LEXICAL resolved through an import binding or lexical scope , HEURISTIC member access or another relationship requiring type inference , EXTERNAL target outside the indexed repository , or UNRESOLVED genuinely unplaceable, with a reason . Never fabricates an edge. An unresolved reference becomes EXTERNAL or UNRESOLVED — never a guessed target and never a silently dropped reference. Sonde measures its zero-setup tree-sitter path against the TypeScript compiler on a pinned fixture and publishes the result, unflattering numbers included spec §12 . COMPILER edges use that compiler directly, so comparing them back to the same authority would not be an independent accuracy test. Generated: 2026-08-23T18:16:41.557Z TypeScript: 5.9.3 bundled; repository TypeScript is never loaded What these numbers cover. The oracle measures the tree-sitter resolution path — the zero-setup default, and the only tier whose accuracy is in question. COMPILER-tier edges come from the TypeScript compiler itself, so scoring them against the same compiler would measure nothing; they are exact by construction and excluded from these figures. Run sonde index --resolve to produce them. The oracle is filtered to in-repo targets; node modules and .d.ts declarations are excluded. Type-only references, JSX intrinsics, export = , decorators, and declaration merging are known expected divergences spec §10 . Tier rows compare that tier alone with the complete oracle, making each tier's independent contribution visible; ALL is the combined result. These divergences are structural, so reading a precision figure as "how often Sonde is wrong" overstates the error rate: Ambiguous member calls emit every candidate. For x.foo with two visible foo declarations, Sonde emits both as confidence-weighted HEURISTIC edges. At most one matches the compiler, so the other counts as a false positive by construction. The alternative is guessing a single target, which invariant 1 forbids — a wrong resolved-looking edge is worse than two honestly heuristic ones. Precision is therefore capped below 1.000 wherever the fixture contains an ambiguous call. Constructor calls are ours alone. Sonde emits CALLS for new Foo ; the oracle does not model them, so each one is a false positive against ground truth that omits it. Member-level IMPLEMENTS is ours alone. Sonde derives an IMPLEMENTS edge from RegExpRouter.add to Router.add once the class declares it implements the interface. tsc reports heritage clauses at the type level only, so every member-level edge counts as a false positive against ground truth that does not model them. The capability is the reason impact on an interface method works at all, so the precision cost is disclosed rather than removed. Counts are absolute, not percentages of a large corpus. Fixture edge totals appear below so a single edge's effect on each figure is visible. Fixture config SHA-256: e02e2d5003f96d1ad22519f04e10d687fe689cf9298e7fcbc588eab525dce1ad Oracle edges: 9 · Sonde edges: 7 · one oracle edge moves recall by 11.1% | Edge kind | Tier | Precision | Recall | TP | FP | FN | |---|---|---|---|---|---|---| | CALLS | ALL | 0.500 | 1.000 | 2 | 2 | 0 | | CALLS | LEXICAL | 0.500 | 0.500 | 1 | 1 | 1 | | CALLS | HEURISTIC | 0.500 | 0.500 | 1 | 1 | 1 | | IMPLEMENTS | ALL | 0.500 | 1.000 | 1 | 1 | 0 | | IMPLEMENTS | LEXICAL | 0.500 | 1.000 | 1 | 1 | 0 | | IMPLEMENTS | HEURISTIC | 1.000 | 0.000 | 0 | 0 | 1 | | INHERITS | ALL | 1.000 | 1.000 | 1 | 0 | 0 | | INHERITS | LEXICAL | 1.000 | 1.000 | 1 | 0 | 0 | | INHERITS | HEURISTIC | 1.000 | 0.000 | 0 | 0 | 1 | | REFERENCES | ALL | 0.571 | 0.800 | 4 | 3 | 1 | | REFERENCES | LEXICAL | 0.800 | 0.800 | 4 | 1 | 1 | | REFERENCES | HEURISTIC | 0.000 | 0.000 | 0 | 2 | 5 | Overall: precision 0.571, recall 0.889 Regenerate with npm run bench:oracle . sonde init path --resolve --yes index and register project MCP config sonde index path --resolve full index; optional compiler pass sonde update path --resolve update; optional compiler pass sonde status path freshness and tier distribution sonde search