The July benchmark tested GPT-5.6 Sol Ultra against V8 security fixes; Hacktron's September post recirculated the results, which came from a controlled lab setup.
By [Ryan Merket](https://runtimewire.com/author/ryan-merket)
· Published
Primary source: [X](https://x.com/HacktronAI/status/2104965022464069763)
Why it matters #
Hacktron's controlled test shows one model completing a difficult browser exploit-development chain at a reported compute cost below $1,600, while leaving real-world repeatability and total costs unproven.
Hacktron co-founder and CTO Mohan Pedhapati reported that GPT-5.6 Sol Ultra built a working Chrome/V8 exploit chain in a controlled test, with Hacktron putting the model-compute cost at $1,596.89. The benchmark was published on July 14th, 2026. Hacktron's September 29th post on X recirculated the results; it did not describe a newly discovered compromise of Chrome users.
https://x.com/HacktronAI/status/2104965022464069763 Pedhapati's full benchmark write-up says Hacktron gave three models a V8 source tree, a sandbox-enabled JavaScript shell, and publicly available security-fix commits to analyze. Sol Ultra was the only one of the three to finish the chain. GPT-5.6 Sol Medium and Grok 4.5 reached partial stages before Hacktron stopped their runs.
The test ended with code execution in the lab environment, demonstrated by launching Calculator. That result is evidence of a model completing a multi-stage exploit-development exercise against a particular browser-engine build. It is not evidence that the model remotely compromised deployed Chrome installations, or that the same chain works against current, patched browsers.
Hacktron reported that the successful run took three days, processed 2.096 billion tokens, made 14,062 requests and used 74 subagents. The $1,596.89 figure is the company's reported model-compute cost for that run. It is not a measure of the full cost of producing a dependable real-world exploit: the benchmark supplied the source, a test harness and security fixes as starting material, and the published result does not establish how much human research, infrastructure or preparation would be required to repeat it under other conditions.
That distinction is central to the result. Hacktron's report describes a controlled experiment, while its X post turns the cost figure into a warning about the falling price of vulnerability discovery. The benchmark supports a narrower point: given a defined target, an instrumented environment and enough inference compute, one tested model persisted through a long series of steps that the other two did not complete. It does not show that exploit development has become routine, or that every software team faces the same immediate risk.
Pedhapati's own background is part of Hacktron's case for why the test matters. The company's team page identifies him as a former senior security researcher at Cure53 and says he previously founded a security-auditing business that generated about €1.5 million in revenue. Hacktron is applying offensive-security research experience to a commercial product: its current site describes code review that validates vulnerabilities in pull requests and white-box penetration tests that use source-code context. Those are the workflows through which Hacktron argues software teams can check their own code before release.
That commercial pitch also puts the benchmark in context. Hacktron is selling security testing and publishing research that argues the cost of attack is dropping. The company disclosed a $2.9 million pre-seed round in May, led by Crane Venture Partners, with participation from Project Europe, Vercel Ventures, Plug and Play Ventures, Cambridge Enterprise Ventures and others. In the same announcement, Hacktron said it had generated about $240,000 in revenue during its first nine months. Those are company-reported figures, not independently audited results.
The Chrome test is consequential as a demonstration of capability, but its economic headline needs a precise denominator. The $1,597 covers compute for one reported run, not a commercial penetration test, a live browser attack, or proof that attackers can reliably convert this workflow into compromises at scale. Hacktron's report itself notes that the other tested models did not finish the chain. The evidence is a sharp advance in one controlled task, rather than a general measure of how quickly software can now be broken.