{"slug": "grep-launches-agentrun-says-code-first-harness-cuts-agent-costs-91", "title": "Grep launches AgentRun, says code-first harness cuts agent costs 91%", "summary": "Grep co-founders AJ Asver and Miguel Rios Berrios launched AgentRun, a code-first harness that converts repetitive AI agent work into typed decisions and ordinary code, and Asver said Grep's internal 100-alert evaluation cut compliance alert review costs from $2.89 with Opus 5 to $0.25 with AgentRun, a roughly 91% reduction that extrapolates from about $289,000 to $25,000 across 100,000 alerts. The test, conducted by Grep rather than an independent benchmark, reported that the routed workflow avoided 30% of profile research, stopped early on 32 of the 100 alerts and invoked a more capable judge agent for identity questions in eight cases. AgentRun targets the cost structure that keeps enterprise agents in pilots by routing exceptions to frontier models while code and specialized decision models handle routine cases.", "body_md": "# Grep launches AgentRun, says code-first harness cuts agent costs 91%\n\n**AJ Asver's internal test reduced compliance model costs from $2.89 to $0.25 per alert, with uncertain cases routed to a frontier agent.**\n\n        By [Ryan Merket](https://runtimewire.com/author/ryan-merket)\n        · Published \n\nPrimary source: [X - AJ Asver](https://x.com/_aj/status/2102061534956662818)\n\n## Why it matters\n\nAgentRun targets the cost structure blocking enterprise agents from moving beyond pilots: frontier models handle exceptions, while code and specialized decision models process routine cases.\n\n[Grep](https://grep.ai/) co-founders [AJ Asver (@_aj)](https://x.com/_aj/status/2102061534956662818) and [Miguel Rios Berrios (@MiguelriosEN)](https://x.com/MiguelriosEN/status/2101029313906987422) have launched AgentRun, a harness that watches an AI agent complete repetitive work and converts parts of that job into cheaper typed decisions and ordinary code.\n\nAsver said Grep's internal test cut the model cost of reviewing a compliance alert from $2.89 with Opus 5 to $0.25 with AgentRun, a reduction of about 91%. Extrapolated across 100,000 alerts, that lowers the bill from roughly $289,000 to $25,000. His September 21st post rounded those figures to more than $290,000 and less than $26,000.\n\nThe cost claim rests on a 100-alert internal evaluation conducted by Grep, rather than an independent benchmark. Grep says all 100 alerts completed and that its leaner workflows produced fewer unsafe clearances than the full-agent approach, though the reported figures primarily document model costs, tool calls and workflow completion.\n\n### A $2.89 alert becomes $0.25\n\nAgentRun grew out of a problem Asver and Rios Berrios encountered after building compliance agents for enterprise customers: a capable frontier model could investigate each case, but repeatedly running a long agent session over every record made the economics difficult at production volume.\n\nThe founders know the workload firsthand. Before starting Parcha Labs, which later developed Grep, Asver worked in product roles at Brex, Coinbase and Google. Rios Berrios previously led platform engineering at Brex and consumer data science at Twitter. Their original product focused on know-your-customer and anti-money-laundering reviews, where analysts must investigate large numbers of screening alerts and record the evidence behind each conclusion.\n\nIn Grep's test, a full Opus 5 agent researched every profile attached to an alert and cost $2.89 per case. A research agent using Gemini Flash with code-based judgment cost $2.02. Moving research to [DeepSeek V4.1 Flash](https://runtimewire.com/models/azure/deepseek-v4.1-flash) and decisions to Jev lowered the figure to $0.39.\n\nAgentRun reached $0.25 by changing the procedure itself. Grep says the routed workflow avoided 30% of profile research, stopped early on 32 of the 100 alerts and invoked a more capable judge agent for identity questions in eight cases. About 30 Jev questions per alert accounted for $0.003 of the total cost.\n\nOne representative alert shows where those savings came from. The original agent opened 24 potential matches, made 826 tool calls, read 271 pages and took 51 minutes. The AgentRun procedure prioritized the profiles that could change the final risk decision, opened two cards, made about 30 tool calls and finished in roughly three minutes.\n\n### The agent writes itself out of routine work\n\nAgentRun begins with a frontier agent completing a job from a standard operating procedure. During that run, the agent records which tools worked, which steps failed and which pieces of evidence determined the outcome. Grep then turns the repeatable procedure into a workflow composed of model calls, typed questions and deterministic code.\n\nMechanical operations such as comparing dates, counting discrepancies or applying a fixed stopping rule move into code. Research remains with an agent when the workflow has to search the web, inspect a registry or operate a browser. Closed questions, including whether a profile matches a policy criterion, move to [TypeSafe AI's Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev).\n\nJev does not generate prose. It receives structured state and returns typed choices, scores or true-or-false probabilities that software can use directly. Grep uses those probabilities to set confidence thresholds. A high-confidence answer continues through the workflow, while an uncertain answer returns to a frontier agent or goes to human review when the procedure requires it.\n\nThat escalation path matters in compliance. A cheaper classifier can lower costs only when the workflow preserves the asymmetric risk of the job: a false positive wastes analyst time, while incorrectly clearing a real sanctions or money-laundering match can become a regulatory failure. Asver said the appropriate confidence threshold depends on the stakes, suggesting 80% to 90% may be acceptable for autocomplete or generative interfaces while higher-risk work requires stricter gates.\n\nIndependent framework documentation carries the same warning. [Pydantic's Jev integration guide](https://pydantic.dev/docs/ai/models/typesafe/) recommends measuring accuracy and handoff rates on labeled data from the intended task. It also advises keeping people involved in consequential actions and pairing model judgments with deterministic controls.\n\n### Grep's bet on cheaper repetition\n\nAsver and Rios Berrios founded Parcha Labs in 2023 after working together at Brex. Parcha raised a [$5 million seed round](https://kindredventures.com/announcement/our-investment-in-parcha-ai-agents-for-the-enterprise/) from Kindred Ventures and Initialized Capital to build enterprise agents for compliance and operations. The founders later expanded the system into Grep, a broader platform for research, due diligence and recurring knowledge work.\n\nAgentRun pushes Grep further into the infrastructure beneath those agents. Grep is selling the idea that a frontier model should discover and document a procedure once, then leave routine executions to cheaper models, typed decisions and code. Expensive reasoning remains available for exceptions instead of being purchased again for every ordinary case.\n\nThe timing follows TypeSafe's September 15th release of Jev after two years in stealth. TypeSafe founder Diogo Almeida, who previously worked on instruction-following research at OpenAI, designed Jev around structured decisions rather than text generation. TypeSafe prices Jev at $42 per billion input tokens and says responses typically arrive in 70 to 500 milliseconds. Those performance figures come from TypeSafe's own evaluations.\n\nGrep began rolling AgentRun out to enterprise customers during the week of September 21st, with access for Pro customers scheduled afterward. The product's central claim is measurable: repetitive agent work should become progressively cheaper as more steps are captured in software. Grep's 100-alert test supports that thesis inside one tightly defined compliance workflow. Production performance across different policies, data sources and failure conditions will determine how far the approach travels beyond it.", "url": "https://wpnews.pro/news/grep-launches-agentrun-says-code-first-harness-cuts-agent-costs-91", "canonical_source": "https://runtimewire.com/article/grep-agentrun-compliance-agent-costs-jev", "published_at": "2026-09-21 20:11:21+00:00", "updated_at": "2026-09-21 20:25:47.358925+00:00", "lang": "en", "topics": ["ai-agents", "ai-products", "ai-startups", "ai-tools"], "entities": ["Grep", "AgentRun", "AJ Asver", "Miguel Rios Berrios", "Opus 5", "Gemini Flash", "DeepSeek V4.1 Flash", "Jev"], "alternates": {"html": "https://wpnews.pro/news/grep-launches-agentrun-says-code-first-harness-cuts-agent-costs-91", "markdown": "https://wpnews.pro/news/grep-launches-agentrun-says-code-first-harness-cuts-agent-costs-91.md", "text": "https://wpnews.pro/news/grep-launches-agentrun-says-code-first-harness-cuts-agent-costs-91.txt", "jsonld": "https://wpnews.pro/news/grep-launches-agentrun-says-code-first-harness-cuts-agent-costs-91.jsonld"}}