cd /news/ai-tools/show-hn-benzi-a-code-intillegence-ha… · home topics ai-tools article
[ARTICLE · art-126426] src=benzi.fly.dev ↗ pub= topic=ai-tools verified=true sentiment=↑ positive

Show HN: Benzi – A Code Intillegence/Harness Beating Claude Code and CodeGraph

Benzi, a new code-intelligence harness, resolved 78.2% of 500 real issues on SWE-bench Verified at under 10 cents per fix, according to benchmarks published on the project's site. In a 24-bug, 10-language comparison against Claude Code Sonnet, Benzi Sonnet, Benzi DeepSeek, and DeepSeek Harness DeepSeek, Benzi Sonnet read 9,125 source lines versus Claude Code Sonnet's 20,704 and DeepSeek Harness DeepSeek's 43,598. Benzi's per-bug costs ranged from $0.19 to $0.57 on Sonnet and $0.015 to $0.068 on DeepSeek, with the project claiming Claude Code's cost climbs with difficulty faster than any other series tested.

read5 min views4 publishedSep 11, 2026

Benzi · benchmarks

Apples to apples comparison: Benzi harness vs mainstream harnesses, 24 GitHub issues, 10 languages.

Benzi harness on SWE-bench Verified. 78.2% of 500 real issues resolved at under 10¢ a fix.

Apples to Apples comparison: Code Graph's code intelligence vs Benzi's AI-native code intelligence

Every task, every attempt, verbatim — nothing held back.

Learn more about Benzi: benzi.fly.dev/about Every harness opens more source as bugs get harder. The question is the slope. Each point is one bug; the 24 are laid out easiest to hardest, left to right.

Difficulty is Claude Code's turn count on that bug — a third-party yardstick, so no harness sets its own position on the axis. Lines read counts only what came back from

    file-read calls; grep and shell output are search, not reading. The figure beside each line is its
    slope: how many extra lines that harness opens per step of difficulty. Hover any point for the bug
    and its count.
source lines read · 24 bugs · lowest of the four marked
bug Benzi Sonnet Benzi DeepSeek Claude Code Sonnet DeepSeek Harness DeepSeek
--- --- --- --- ---
mux 64 64 383 1,372
commons-cli 39 235 456 771
addressable 180 240 200 1,603
jsoup 61 170 60 616
yaml-cpp 120 383 194 461
cJSON 17 187 105 854
dayjs 64 336 307 610
gson 78 158 80 516
CsvHelper 42 390 335 805
semver 353 468 563 1,857
hashie 153 172 300 458
money 681 572 661 1,892
rich 255 514 736 1,421
fmt 604 315 264 1,097
sqlglot 227 115 800 777
scrapy 386 1,580 1,259 2,127
marked 905 807 2,167 4,091
http-parser 449 1,201 1,270 2,635
zod 841 1,476 1,956 3,608
quartznet 488 1,707 1,897 4,197
sqlparser 620 1,007 1,225 2,334
nlohmann/json 379 929 1,764 5,201
ts-pattern 1,130 1,232 1,139 1,910
nats-server 989 2,149 2,583 2,385
all 24 9,125 16,407 20,704 43,598

Lines read counts only what came back from file-read calls; grep and shell output are search, not reading. The green figure in each row is the lowest of the four.

The same 24 bugs in the same order, with wall clock in place of lines read.

Wall clock is raw here — unlike the tables above, Benzi's per-repo index build is not subtracted, so these seconds run slightly higher than the warm figures quoted elsewhere on this page.

    Each point is that harness's most recent solved run for that bug; unsolved and unfinished runs are left
    out rather than plotted as fast. Benzi on Sonnet never solved http-parser and the DeepSeek harness never
    ran nats-server, so those two points are absent and neither enters its fit.

And the same again with dollars on the vertical axis.

Priced at the published per-token rates, same run selection as the chart above it. The two DeepSeek series run an order of magnitude cheaper than the two Sonnet ones, so at this scale they sit close to the baseline — the per-bug figures behind them are in the DeepSeek table further down.

    What the axis does show is the **slope**: Claude Code's cost climbs with difficulty faster than any
    other series here.
cost per fix · USD at list price · lowest of the four marked
bug Benzi Sonnet Benzi DeepSeek Claude Code Sonnet DeepSeek Harness DeepSeek
--- --- --- --- ---
mux $0.39 $0.036 $0.25 $0.023
commons-cli $0.26 $0.030 $0.29 $0.014
addressable $0.57 $0.043 $0.35 $0.042
jsoup $0.23 $0.025 $0.30 $0.051
yaml-cpp $0.25 $0.068 $0.37 $0.024
cJSON $0.27 $0.036 $0.44 $0.073
dayjs $0.21 $0.038 $0.55 $0.040
gson $0.19 $0.015 $0.45 $0.022
CsvHelper $0.25 $0.038 $0.58 $0.042
semver $0.62 $0.10 $1.43 $0.097
hashie $0.69 $0.051 $0.92 $0.043
money $0.79 $0.036 $0.95 $0.052
rich $0.33 $0.11 $1.30 $0.053
fmt $0.54 $0.14 $1.12 $0.071
sqlglot $0.62 $0.033 $1.23 $0.059
scrapy $0.73 $0.10 $1.35 $0.19
marked $1.08 $0.20 $3.59 $0.13
http-parser $0.36 $3.44 $0.23
zod $1.53 $0.13 $2.47 $0.33
quartznet $0.95 $0.23 $3.99 $0.31
sqlparser $0.72 $0.14 $3.22 $0.33
nlohmann/json $1.03 $0.17 $3.68 $0.44
ts-pattern $3.74 $0.31 $4.33 $0.053
nats-server $1.99 $0.23 $2.94
all 24 $17.96 $2.66 $39.54 $2.70

Priced at published per-token rates. The green figure in each row is the lowest of the four; the two DeepSeek columns are cheaper largely because that model costs roughly twenty times less per token. Blank cells are the two runs that never produced a fix.

── more in #ai-tools 4 stories · sorted by recency
── more on @benzi 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/show-hn-benzi-a-code…] indexed:0 read:5min 2026-09-11 ·