Leviathan indexes million-row datasets for AI agents without filling context windows Joshua Gunn open-sourced Leviathan, a Rust command-line tool that indexes large record collections for AI agents, on October 5th under the Apache-2.0 license. In a synthetic million-row benchmark, Leviathan returned a relevant record in the top five results for 99% of questions at a median of 436 tokens per query and 33 milliseconds of median search latency, versus a median of 107,122 tokens for the tested grep strategy. The benchmark documentation covers 1,500 invented machines and 1,200 questions across six dataset sizes, and the repository states the accuracy figures are an upper bound for real data because the test measures retrieval only, not whether a model produces a correct final answer. Leviathan indexes million-row datasets for AI agents without filling context windows Joshua Gunn's open-source Rust tool returned relevant records in 99% of a synthetic million-row test, using far fewer tokens than the tested grep approach. By Ryan Merket https://runtimewire.com/author/ryan-merket ยท Published Primary source: X https://x.com/joshuagunnn/status/2107108554192941493 Why it matters Agent workflows can waste context and model calls reading raw records before reasoning begins. Leviathan offers a local retrieval layer, with an open benchmark that quantifies token savings while clearly leaving real-world accuracy and end-to-end answer quality untested. Joshua Gunn @joshuagunnn https://x.com/joshuagunnn open-sourced Leviathan https://github.com/elstongun/leviathan on October 5th, a Rust command-line tool that indexes large collections of records and returns short search results for AI agents. The project targets a common problem in agent workflows: retrieving useful details from data too large to load into a model's context window. Gunn's background is in high-frequency-trading infrastructure; his public profiles say he now works on intelligent manufacturing. Leviathan's benchmark uses synthetic factory-maintenance records, a practical connection to that work: agents are asked questions such as what fixed a machine previously, then must find the relevant entry among a large history. At one million records, Leviathan's benchmark reports a median of 436 tokens per query and a relevant record in the top five results for 99% of questions. The comparison grep strategy used a median of 107,122 tokens. Gunn's launch thread on X https://x.com/joshuagunnn/status/2107108554192941493 also cites 33 milliseconds of median search latency and 602 tokens for Leviathan's largest answer. The benchmark has limits. The benchmark documentation https://github.com/elstongun/leviathan/blob/main/docs/BENCHMARKS.md describes a generated dataset of 1,500 invented machines and 1,200 questions across six dataset sizes. The benchmark measures whether retrieval returns a relevant record; it does not test whether an AI model can use that record to produce a correct final answer. The repository itself calls the accuracy figures an upper bound for real data, where language and record structure may differ. The comparison also favors neither a casual nor a fully realistic reading of grep. Its tested strategies were given the exact machine key, even when a person would have used the machine's name. Grep still returned a relevant record within its output cap for 96% of the million-row questions, but produced much larger outputs for an agent to inspect. Leviathan's advantage in the test is compact, ranked retrieval, not proof that ordinary text search fails on every large dataset. Leviathan works as a local index rather than a hosted database service. It accepts JSONL, JSON, CSV or TSV files, SQLite databases, and output piped from database command-line tools. The repository says it builds a SQLite full-text index, then returns result cards with record identifiers, matching text and citations. For the million-row synthetic dataset, the documented index build took 52 seconds and occupied 1.2 GB, compared with 678 MB for the source data. A developer can export records from PostgreSQL, MySQL, DuckDB or MongoDB and pipe them into Leviathan; the tool says it does not handle database credentials. Its command-line interface is the default, with an optional Model Context Protocol server. Gunn argues that registering MCP tools can impose a token cost in every agent session, while a command-line skill is only used when called. The benchmark measures the MCP schema at 638 tokens per session for its test index. The code is available under the Apache-2.0 license, and the repository provides installation instructions and prebuilt binaries. Leviathan uses conventional full-text indexing and ranking to reduce how much raw history an agent must read. Whether that remains reliable beyond the project's synthetic maintenance-log test will depend on how the search performs on messier records and questions that span whole datasets rather than one identifiable entity.