cd /news/artificial-intelligence/cognition-ceo-scott-wu-says-ai-bench… · home topics artificial-intelligence article
[ARTICLE · art-75951] src=cryptobriefing.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Cognition CEO Scott Wu says AI benchmarks are losing their meaning as models saturate every test

Cognition CEO Scott Wu argues that traditional AI benchmarks are losing their meaning as models saturate every test, with the company's AI coding agent Devin approaching near-perfect scores on SWE-Bench. Cognition, now valued at $26 billion after raising $1 billion in May 2026, has developed its own proprietary evaluation called FrontierCode 1.1 and uses internal 'junior dev' benchmarks to measure real-world output instead of leaderboard scores.

read2 min views1 publishedJul 27, 2026
Cognition CEO Scott Wu says AI benchmarks are losing their meaning as models saturate every test
Image: Cryptobriefing (auto-discovered)

Via colossus.com

The company behind Devin, now valued at $26 billion, argues the industry needs to stop obsessing over leaderboards and start measuring real-world output

Here’s a fun paradox for the AI industry: the better your models get, the less your scorecard matters. That’s essentially what Cognition CEO Scott Wu is arguing as the company’s flagship AI coding agent, Devin, approaches near-perfect scores on the very benchmarks that once defined the competitive landscape.

Wu’s position is that traditional AI benchmarks are becoming less meaningful because frontier models can now solve essentially any well-defined task thrown at them. When everyone’s acing the test, the test stops telling you anything useful.

From 13% to 90%, and now what #

When Devin launched in March 2024, it scored just 13% on SWE-Bench, a widely used benchmark for evaluating AI software engineering capabilities. By mid-2026, that number had climbed to roughly 90% on the original SWE-Bench and approximately 80% on SWE-Bench Pro.

Instead of chasing public benchmark scores, Cognition has developed its own proprietary evaluation called FrontierCode 1.1. The company also uses internal “junior dev” benchmarks designed to simulate real-world coding tasks rather than the kind of cleanly defined problems that traditional benchmarks tend to favor.

Wu has been particularly pointed about metrics like token usage, arguing they can actually mislead evaluations of what an AI system is genuinely producing.

A $26 billion bet on outcomes over scores #

Cognition raised $1 billion in May 2026 at a $26 billion valuation. The company was founded in November 2023 by Wu and his co-founders.

In certain contexts, Devin reportedly contributes about 89% of the committed code produced by Cognition’s own engineering team.

Cognition also recently acquired Windsurf, a rival AI coding entity, in a move that bolsters its competitive position as the market for autonomous coding tools heats up.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our

Editorial Policy.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @cognition 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/cognition-ceo-scott-…] indexed:0 read:2min 2026-07-27 ·