# Devin AI and GPT-5.6 Sol Ultra crack math problems that stumped researchers for half a century

> Source: <https://startupfortune.com/devin-ai-and-gpt-56-sol-ultra-crack-math-problems-that-stumped-researchers-for-half-a-century/>
> Published: 2026-07-23 19:52:31+00:00

*Devin and GPT-5.6 Sol Ultra have turned graph theory into a stress test for AI agents, but you should treat the biggest claim as a claim until the checking is finished.*

AI agents did not just write another demo app in July. They walked into graph theory, a field where a proof is either correct or it is not, and started producing results on problems people had left sitting for decades. That is the live story.

The first burst came from Devin, Cognition's autonomous software engineer. Jared Zoneraich, a builder in residence at Cognition, said on X that Devin had refuted Graffiti Conjecture 154, proved Graffiti Conjectures 39 and 40, and refuted Brandt's Regular Supergraph Problem from Douglas West's open problems list. The Graffiti program itself is real history, not AI marketing copy. It was introduced by Siemion Fajtlowicz in 1986 to generate graph theory conjectures, and old lists of its problems still sit on university pages. Devin's reported counterexample to Conjecture 154 was a lollipop graph, a 50-clique joined to a 70-edge path, with 120 vertices in all. That is the sort of detail you want in this story, because it gives mathematicians something to attack.

Then OpenAI pushed the story much further. On July 10, several outlets reported that OpenAI had posted a PDF claiming that GPT-5.6 Sol Ultra produced a proof of the Cycle Double Cover Conjecture, a problem associated with George Szekeres's 1973 work and Paul Seymour's 1979 formulation. The conjecture asks whether every bridgeless graph has a collection of cycles that covers each edge exactly twice. It sounds neat. It has not been neat to prove.

The proof claim came one day after OpenAI's July 9 launch of GPT-5.6, where the company described Sol as its flagship model and Ultra as a setting that coordinates multiple agents across parallel workstreams. Reports on the proof say the run used 64 subagents in under an hour. The Decoder reported that University of Manchester mathematician Thomas Bloom praised the proof's simplicity but criticized the missing citations to earlier work, including a 1983 paper by Bermond, Jackson, and Jaeger. Do not skip that caveat. The Cycle Double Cover Conjecture has attracted claimed proofs before, and the graveyard of plausible-looking mathematics is larger than most investors think.

## The agent story is bigger than the proof

Cognition's valuation story was already moving fast before Devin's graph theory week. In its May 27 funding announcement, Cognition said it had raised more than $1 billion at a $26 billion valuation, led by Lux Capital, General Catalyst, and 8VC. The company also said its run-rate revenue had grown to $492 million, named Citi, Mercedes-Benz, Goldman Sachs, Dell, Santander, the U.S. Army and others as customers, and claimed that 89% of code committed by its own engineers is committed by Devin.

Those numbers are not made true by a lollipop graph. But they do change how you read the company. Cognition is not selling a faster autocomplete box. It is selling the idea that an agent can choose a problem, work through a strategy, run checks, and hand you something that survives inspection. If that holds beyond software maintenance, the market has been valuing the wrong part of the stack.

There is still a hard distinction here. Devin is an orchestration product, not a foundation model lab in the same sense as OpenAI, Anthropic or Google DeepMind. Cognition's own site says Devin works inside existing tools. It evaluates model performance across many software engineering tasks. If the product is choosing the right model, spawning the right agents, testing the output and packaging the work, that is still valuable. It just means you should not confuse the agent brand with the underlying model capability. That is not a small distinction.

## Verification is the product now

The most important update since the first OpenAI PDF is not the applause. It is the checking. MathWorld's Cycle Double Cover page now lists OpenAI's proof, OpenAI's prompt, and an OpenAI cdc-lean GitHub repository in its references. DeepWiki's index of that repository describes a Lean 4 formalization for finite loopless bridgeless multigraphs, with top-level theorems named for the cycle double cover result. That does not automatically end the debate, because mathematicians still have to inspect whether the formal definitions match the conventional conjecture. But it moves the argument into a better room.

That is where you should focus your attention. A polished three-page proof can be persuasive and wrong. A machine-checked repository can still encode the wrong statement, but at least you can inspect the statement. For AI-generated science, the proof is no longer just the paper. It is the paper, the prompt, the formalization, the dependency list, and the audit trail.

The spillover goes well beyond graph theory. Google DeepMind said in 2023 that GNoME found 2.2 million crystal structures, including about 380,000 stable candidates for experimental synthesis. Isomorphic Labs says its drug design engine lets scientists run in-silico experiments across vast GPU fleets, and Recursion says its platform is built on more than 50 petabytes of biological and chemical data. These are not the same problem as graph theory. They are harder to verify cleanly. That is exactly why the math story matters.

Frankly, the striking part is not that one model may have solved one famous conjecture. It is the workflow. Zoneraich showed Devin one result and asked it to find similar open problems. OpenAI's prompt reportedly pushed Sol Ultra to send agents down different mathematical routes and assign others to hunt for errors. You can see the shape of future research in that pattern: generate, attack, formalize, repeat.

For founders and investors, the lesson is plain. Do not buy every proof claim the day it appears on a PDF. Do not dismiss it either. The agent era will be measured by what survives verification, and July 2026 gave you a cleaner measuring stick than another leaderboard ever could.

**Also read:** [Patreon cuts 93 jobs and names AI transformation as the forcing function](https://startupfortune.com/patreon-cuts-93-jobs-and-names-ai-transformation-as-the-forcing-function/) • [AegisAI bets the email security incumbents can't keep up with AI-generated phishing](https://startupfortune.com/aegisai-bets-the-email-security-incumbents-cant-keep-up-with-ai-generated-phishing/) • [Runway bets on being the infrastructure layer under every generative media model](https://startupfortune.com/runway-bets-on-being-the-infrastructure-layer-under-every-generative-media-model/)
