{"slug": "show-hn-halv-cut-ai-agent-cost-by-57-1-using-jev", "title": "Show HN: Halv cut AI agent cost by 57.1% using Jev", "summary": "Halv cut recorded model cost by 57.1% versus vanilla Codex across 14 repository tasks, spending $144.67 against $337.50 while both arms passed 25 of 42 runs, according to Halv's own interim benchmark checkpoint. The Halv workflow combined JEV agent selection, an investigate-l1-l2-v1 hierarchy with Astra medium as coordinator and Sol 6 or Luna 6 workers, plus Crux 0.10.3, RTK and Headroom context tools on Codex 0.155.1. Halv reported the strongest single result on ArcadeDB-4455 at 81.7% lower cost ($14.97 versus $81.85) with 3/3 passes in both arms, and stated the experiment cannot isolate how much of the savings came from JEV, model selection, context tools or coordination policy.", "body_md": "For the main results in plain language, read [the short version](https://halv.ai/blog/halv-57-percent-lower-model-cost).\n\n**Halv recorded less than half the model cost while matching vanilla’s aggregate verifier pass count.**\n\nAcross 14 repository tasks, we ran three paired repetitions per task.\nEach pair compared vanilla Codex with a Halv workflow that combined model routing and context tools.\nBoth arms passed **25 of 42 runs**.\n\nHalv recorded **$144.67**, compared with **$337.50** for vanilla.\nThat is **57.1% lower recorded model cost**, including all coordinator and worker sessions in the selected runs.\n\nThis is an interim checkpoint from our own benchmark campaign. It is a workflow comparison with different worker models, not a controlled test of one component.\n\n## The result at a glance\n\n| Metric | Vanilla Codex | Codex + Halv | \n|---|---|---|\n| Verifier passes | 25/42 | 25/42 | \n| Pass rate | 59.5% | 59.5% | \n| Recorded model cost | $337.50 | **$144.67** | \n| Cost per verifier pass | $13.50 | **$5.79** | \n| Total input + output tokens | 48,755,656 | 170,114,775 | \n| Distinct tasks | 14 | 14 | \n| Repetitions per task | 3 | 3 | \n\nWe calculate savings from unrounded costs:\n\n```\n1 − ($144.666190 / $337.501056) = 57.1361%\n```\n\nThe numerator includes selected runs that failed verification. The denominator uses the same accounting boundary for vanilla. Neither arm gets to remove a valid run because its patch failed.\n\nEqual aggregate pass counts do not establish equal quality on every task. They also do not turn 42 repetitions into 42 independent repository tasks.\n\n## The strongest results\n\n**ArcadeDB-4455: 81.7% lower cost, with 3/3 passes in both arms.**\nVanilla recorded $81.85 across its three runs.\nHalv recorded $14.97 and passed all three verifiers.\nThis is the strongest task result that combines lower cost with the same pass count.\n\n**HugeGraph-3037: 67.1% lower cost, with 3/3 Halv passes versus 2/3 vanilla passes.**\nHalv recorded $12.21, compared with $37.05 for vanilla.\n\n**ArcadeDB-4411: 46.1% lower cost, with 3/3 passes in both arms.**\nHalv recorded $7.45, compared with $13.80 for vanilla.\n\nThese examples show where the workflow worked well. The full table also shows where it cost more or produced fewer passes.\n\n## What changed in the workflow\n\nBoth arms used Codex **0.155.1**, an **Astra medium** coordinator, and the **default** service tier.\nEach pair used the same task and repository verifier.\n\nVanilla asked Astra to complete the task alone.\nHalv used the experimental `investigate-l1-l2-v1` hierarchy:\n\n1. Astra assigns an investigation to Luna 6 with medium reasoning.\n2. The investigation informs the task brief and difficulty assessment.\n3. JEV selects the L1 worker model and reasoning effort.\n4. L1 can assign bounded work to a cheaper L2 worker when useful.\n5. Astra reviews and integrates the result.\n\nEvery selected worker used the Codex lane. The recorded routes selected Sol 6 with medium or xhigh reasoning, and Luna 6 with high reasoning. The separate investigation stage used Luna 6 medium.\n\nL2 work appeared in **4 of the 42 Halv runs**.\nMost runs used the investigation and L1 stages without adding L2.\nThe hierarchy permits delegation; it does not require every task to create another layer.\n\nHalv also enabled **Crux 0.10.3**, **RTK**, and **Headroom**.\nCrux supplied code navigation, RTK filtered command output, and Headroom processed model context.\nThe external benchmark monitor supplied status without asking Astra to generate routine progress updates.\n\nThese components changed together. This experiment cannot isolate how much of the result came from JEV, model selection, context tools, or coordination policy.\n\n## Where Halv’s recorded cost went\n\n**JEV agent selection is included in the Halv subscription price.**\nHalv users pay no separate fee for JEV to select their agents.\nThe table below measures model execution costs; it does not allocate the Halv subscription fee across benchmark runs.\n\nThe selected Halv runs include all recorded roles:\n\n| Role | Recorded cost | Share of Halv cost | \n|---|---|---|\n| Astra coordinator | $64.50 | 44.6% | \n| Investigation | $1.06 | 0.7% | \n| L1 workers | $76.03 | 52.6% | \n| L2 workers | $3.07 | 2.1% | \n| **Total** | **$144.67** | **100%** | \n\nThe coordinator still accounted for a substantial share. The cheap investigation stage did not make coordination free. Measuring the complete workflow prevents a short coordinator conversation from hiding expensive worker activity.\n\nHalv used **3.49 times as many total tokens** as vanilla.\nLower recorded cost came with a different mix of model usage and caching.\nThe result supports a cost claim for this workflow, not a claim that this hierarchy saves tokens.\n\n[Try Halv and measure the complete workflow on your own verified tasks.](https://halv.ai/download)\n\n## Every task, including the regressions\n\nEach row sums three repetitions per arm. Positive savings mean Halv recorded lower cost. Negative savings mean Halv recorded higher cost. Costs in this table are rounded to cents; aggregate calculations use unrounded values.\n\n| Task | Vanilla passes | Halv passes | Vanilla cost | Halv cost | Halv savings | \n|---|---|---|---|---|---|\n| ArcadeDB-4281 | 2/3 | 3/3 | $11.26 | $7.50 | 33.4% | \n| ArcadeDB-4411 | 3/3 | 3/3 | $13.80 | $7.45 | 46.1% | \n| ArcadeDB-4455 | 3/3 | 3/3 | $81.85 | $14.97 | **81.7%** | \n| gentle-ai-595 | 3/3 | 3/3 | $8.13 | $7.19 | 11.6% | \n| nanobot-4048 | 3/3 | 3/3 | $4.70 | $5.71 | −21.4% | \n| nanobot-4129 | 3/3 | 3/3 | $4.40 | $5.48 | −24.7% | \n| nanobot-4274 | 0/3 | 0/3 | $11.17 | $14.56 | −30.4% | \n| lossless-claw-814 | 0/3 | 0/3 | $20.66 | $11.97 | 42.1% | \n| Perry-3982 | 3/3 | 1/3 | $71.20 | $17.90 | 74.9% | \n| codex-lb-744 | 3/3 | 3/3 | $11.46 | $10.77 | 6.0% | \n| agno-8148 | 0/3 | 0/3 | $32.24 | $8.85 | 72.6% | \n| dubbo-go-3357 | 0/3 | 0/3 | $7.60 | $9.84 | −29.4% | \n| hugegraph-3037 | 2/3 | 3/3 | $37.05 | $12.21 | 67.1% | \n| pulsar-25793 | 0/3 | 0/3 | $21.99 | $10.27 | 53.3% | \n| **Total** | **25/42** | **25/42** | **$337.50** | **$144.67** | **57.1%** | \n\nHalv cost more on **four of 14 tasks**: the three nanobot tasks and dubbo-go-3357.\nThose rows show that routing overhead can outweigh the cheaper worker mix.\nThe measurements alone do not identify the cause of each regression.\n\nPerry-3982 deserves separate attention. Halv recorded 74.9% lower cost, but passed only one repetition while vanilla passed all three. That is a correctness regression, not an unqualified success.\n\nThe two additional Halv passes on ArcadeDB-4281 and HugeGraph-3037 offset the two missing Perry passes. Several tasks failed in both arms. Spending less on a failed task still does not complete that task.\n\n## Selection, exclusions, and the checkpoint\n\nThe campaign plan contains **111 tasks**.\nThis report covers its first **14 tasks**, with three repetitions per task, in fixed plan order.\nThe runner alternated Halv and vanilla within each repetition.\n\nThe campaign paused after Pulsar-25793 at the user’s request on **September 30, 2026, at 17:42 UTC**.\nWe did not select these 14 tasks by sorting their savings.\nHowever, this was a discretionary interim stop, not a preregistered final endpoint.\n\nThe selector takes the **first valid run** for each task, arm, and repetition.\nIt requires complete usage, a valid verifier reward, and the required runtime checks.\nInfrastructure-invalid runs can be retried.\nA valid reward of zero remains in the comparison.\n\nOperational repairs occurred during the campaign. They included indexing launcher repairs and a Gradle stack adjustment before the Pulsar retries. We also tried a native-completion instruction change, then restored the original waiting policy at the user’s request. The ten recorded runs from that policy experiment were discarded before continuation.\n\nThe headline excludes:\n\n- **13 infrastructure-invalid attempts** in the retained results log, with**$25.53 in known cost** and**nine unpriced attempts** ;\n- **10 recorded runs** from the rolled-back policy experiment, with**$40.79 in known cost** ;\n- diagnostics and interrupted work, including an unmeasured cancelled attempt.\n\nThese exclusions matter.\nThe $144.67 versus $337.50 comparison is **selected valid-run cost**, not the entire campaign bill.\nThe [attempt ledger](https://halv.ai/evidence/swe-rebench-astra-42-pairs/attempt-ledger.json) preserves the excluded records and unknown costs.\n\n## How to inspect the evidence\n\nThe [evidence browser](https://halv.ai/evidence/swe-rebench-astra-42-pairs/) links every task repetition to both run summaries.\nEach summary includes the recorded verifier reward, usage by session, model configuration, cost, and source fingerprint.\nHalv summaries also include routing decisions and cost by role.\n\nThe public export omits credentials, local paths, prompts, and model transcripts. It contains sanitized summaries, not a complete transcript archive. Source hashes identify the private source snapshot without publishing those private files.\n\nDownload the [evidence archive](https://halv.ai/evidence/swe-rebench-astra-42-pairs/halv-swe-rebench-astra-42-pairs-evidence.tar.gz), extract it, and run:\n\n```\nnode verify-selected-pairs.mjs\n```\n\nThe script checks file hashes, unique pairs, session totals, recorded rewards, task totals, and aggregate arithmetic. It does not rerun repository tests or independently authenticate provider billing.\n\nThe [machine-readable results](https://halv.ai/evidence/swe-rebench-astra-42-pairs/selected-pairs.json) contain the full configuration summary and all 42 pairs.\nThe [checksum manifest](https://halv.ai/evidence/swe-rebench-astra-42-pairs/MANIFEST.sha256) covers each published evidence file.\n\n## What this result supports\n\nThis checkpoint supports a specific claim: **the Halv workflow recorded 57.1% lower model cost with the same aggregate number of verifier passes**.\nIt does not establish statistical equivalence, universal savings, or Claude Code performance.\nThree repetitions per task do not establish a stable estimate for every repository.\n\nRecorded costs are token-based model estimates, not invoices. They include cached input accounting and every recorded coordinator and worker session in the selected runs. They exclude the Halv subscription, hosting, indexing compute, developer time, and unmeasured operational work. A fixed subscription price does not fall when this metric falls.\n\nOur [earlier 20-pair report](https://halv.ai/blog/halv-swe-rebench-20-pairs) remains available with its original evidence.\nIt used a different model configuration and task sample.\nComparing the two headline percentages does not measure improvement between Halv versions.\n\nFor this checkpoint, the useful result is already concrete.\nHalv spent **$192.83 less in recorded model cost** across the selected runs and matched vanilla’s **25 verifier passes**.\n\n## Frequently asked questions\n\n### What does the 57.1% savings measure?\n\nIt compares recorded model cost across 84 selected runs: 42 Halv runs and 42 vanilla runs. Costs include coordinator sessions, worker sessions, and valid runs that failed verification. They exclude infrastructure-invalid attempts and the rolled-back policy experiment.\n\n### Did Halv solve more tasks?\n\nBoth arms passed 25 of 42 runs. They did not solve the same set of repetitions. Halv gained passes on ArcadeDB-4281 and HugeGraph-3037, but lost two Perry passes.\n\n### Did Halv use fewer tokens?\n\nNo. Halv used 170,114,775 total tokens versus 48,755,656 for vanilla. The routed workflow recorded lower model cost while using 3.49 times as many tokens.\n\n### Did both arms use the same model?\n\nBoth coordinators used Astra with medium reasoning. Vanilla worked alone. Halv added Luna and Sol workers with different reasoning efforts, plus Crux, RTK, and Headroom.\n\n### Does this reduce my subscription price by 57.1%?\n\nNo. These are recorded token-based cost estimates for the selected benchmark runs. They are not invoices or a reduction in fixed subscription prices.", "url": "https://wpnews.pro/news/show-hn-halv-cut-ai-agent-cost-by-57-1-using-jev", "canonical_source": "https://halv.ai/blog/halv-swe-rebench-astra-42-pairs/", "published_at": "2026-09-30 19:16:09+00:00", "updated_at": "2026-09-30 19:19:37.987959+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "ai-products", "developer-tools", "mlops"], "entities": ["Halv", "JEV", "Codex", "Astra", "Luna 6", "Sol 6", "Crux", "ArcadeDB"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-halv-cut-ai-agent-cost-by-57-1-using-jev", "markdown": "https://wpnews.pro/news/show-hn-halv-cut-ai-agent-cost-by-57-1-using-jev.md", "text": "https://wpnews.pro/news/show-hn-halv-cut-ai-agent-cost-by-57-1-using-jev.txt", "jsonld": "https://wpnews.pro/news/show-hn-halv-cut-ai-agent-cost-by-57-1-using-jev.jsonld"}}