David Zhang's jevgrep saves 29% on agent costs, excluding its own bill David Zhang released jevgrep, a command-line tool that searches code repositories for coding agents, and its current repository reports a 28.6% reduction in coding-agent costs — from $7.62 to $5.44 across ten Python tasks — while both the jevgrep and baseline runs solved eight of 10 tasks. The figure excludes the Jev API charges required to run jevgrep, and a separate evaluation file reports at least $1.57 in observed Jev charges for those tasks with the complete Jev bill unknown. Zhang's September 26th post on X had cited roughly 40% lower cost but only seven of 10 tasks solved with jevgrep versus eight in the baseline, and the repository notes the ten-task test is a tuned Python subset with a frozen package and agent skill and that the saved baseline was not rerun. David Zhang's jevgrep saves 29% on agent costs, excluding its own bill The CLI searches code repositories for coding agents; its latest 10-task test kept the baseline's solve rate, while the launch post cited a different 40% result. By Ryan Merket https://runtimewire.com/author/ryan-merket · Published Primary source: X https://x.com/dzhng/status/2103920741481848861 Why it matters Jevgrep's updated test reports 28.6% lower agent-model costs with the same 8-of-10 solve rate, but excludes at least $1.57 in observed Jev charges. The result points to a real optimization target while showing why retrieval savings need full-cost accounting and broader tests. David Zhang released jevgrep, a command-line tool that helps coding agents find relevant code before they start making changes. In a September 26th post on X https://x.com/dzhng/status/2103920741481848861 , Zhang said an early SWE-bench test cut agent costs by about 40%. The current jevgrep repository https://github.com/dzhng/jevgrep reports a different comparison: 28.6% lower coding-agent costs, with both versions solving eight of 10 tasks. https://x.com/dzhng/status/2103920741481848861 https://x.com/dzhng/status/2103920741481848861 That newer result is the more useful measure. The README says the agent's cost fell from $7.62 to $5.44 across ten Python tasks, including failed attempts. The figure excludes the Jev API charges needed to run jevgrep. A separate evaluation file reports at least $1.57 in observed Jev charges for those tasks and says the complete Jev bill is unknown. The headline saving therefore does not include the full cost of the retrieval system. The ten-task sample also sets a tight limit on what can be claimed. The same eight tasks passed with and without jevgrep, and the two failures remained failures. The repository identifies the test as a tuned Python subset, with a frozen package and agent skill, and says the saved baseline was not rerun. This is evidence that the tool reduced the tested agent's model bill without lowering its solve rate in that run. It does not establish the same result across larger repositories, other programming languages, models or coding-agent products. Zhang's first post gave a less favorable comparison: roughly 40% lower cost, but seven of 10 tasks solved with jevgrep against eight in the baseline. The repository's present README describes the later equal-success run. The changing results illustrate why a single percentage is a poor summary of retrieval tools: less context can lower inference spending while also depriving an agent of code it needs. The relevant measure is cost at a comparable completion rate, with the retrieval model's bill included. Jevgrep addresses the preliminary work that happens before code generation. A developer can ask jg a question such as where authentication is checked before a request reaches a handler. The tool searches folders, files and code declarations, then returns relevant paths, reading leads and source excerpts for the coding agent to inspect. Zhang says coding agents spend 30% to 60% of their tokens collecting context; that is his estimate, not a general industry measurement. The command-line design is intentional. Zhang said in his post that jevgrep's output is meant for agents, not human readers, and told users to install its built-in skill so an agent knows when to call jg . The repository says the skill works with Claude Code, Codex, OpenCode and other compatible agents. The CLI itself is available through npm; it requires Node.js 22 or newer, macOS or Linux, and credentials for a supported model provider, including TypeSafe, Vercel AI Gateway, OpenRouter https://runtimewire.com/models/fal/openrouter-router or OpenCode Zen. The tool runs on Jev, a decision model from TypeSafe AI. TypeSafe opened Jev to early access on September 15th, 11 days before Zhang's post. The lab describes Jev as a model for structured decisions rather than general text generation. Jevgrep applies it to a practical retrieval problem: ranking which parts of a repository are likely to answer an agent's question. The sequence gives TypeSafe a concrete developer-tool use case soon after its model launch, while giving Zhang a way to outsource part of codebase research to a specialized service. Zhang's earlier work includes co-founding Amity, where he served as CTO, and NockNock, a venue-booking marketplace. He later founded Aomni, an AI sales-research company, and his GitHub profile identifies him as an engineer and links him to Duet. The jevgrep project continues his work on agent-oriented software: his other public repositories include an AI research assistant and an open-source CRM designed for agents. Jevgrep is open source under an MIT license. Its repository cautions that eligible source code is sent to Jev through the selected provider; default filters respect ignore files and exclude obvious credential files, but are not a guarantee that sensitive material has been removed. That makes the tool's economic test only one part of adoption: teams also have to decide whether their code can be sent to the provider they configure.