cd /news/artificial-intelligence/computer-use-agents-hit-85-on-osworl… · home topics artificial-intelligence article
[ARTICLE · art-90572] src=cryptobriefing.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Computer-use agents hit 85% on OSWorld benchmark, up from 12% in two years

Anthropic's Claude Mythos Preview, Fable 5, and Opus 4.8 top the OSWorld-Verified leaderboard with scores of 85.4%, 85.0%, and 83.4%, up from 12% in April 2024, surpassing the human baseline of 72%. However, on the harder OSWorld 2.0 benchmark, the best agents complete only 20.6% of tasks, and successful agents require 1.4 to 2.7 times more steps than humans, indicating limitations in open-ended tasks.

read2 min views1 publishedAug 10, 2026
Computer-use agents hit 85% on OSWorld benchmark, up from 12% in two years
Image: Cryptobriefing (auto-discovered)

Rapid gains on a popular AI benchmark mask a harder truth: real-world tasks still trip up the best agents by a wide margin

Two years ago, AI agents that operate computers the way humans do, by looking at a screen and deciding where to click, were completing roughly one in eight assigned tasks correctly. Today, the best systems are clearing 85% on the same test.

The benchmark in question is OSWorld, currently the most widely cited evaluation platform for multimodal computer-use agents (CUAs). These are AI systems that navigate a desktop environment using screenshots, mouse clicks, and keystrokes rather than direct API calls, essentially watching a screen and acting on what they see.

The numbers behind the leap #

In April 2024, leading agents were scoring around 12% on OSWorld. By mid-2025 that figure had climbed into the mid-30s. By June 2026, three Anthropic models sit at the top of the OSWorld-Verified leaderboard: Claude Mythos Preview at 85.4%, Fable 5 at 85.0%, and Opus 4.8 at 83.4%.

For context, the human baseline on OSWorld sits at approximately 72%. When Simular’s Agent S3 crossed that threshold in December 2025, reaching 72.6%, it was the first widely noted instance of a CUA surpassing average human performance on the test. Anthropic’s current crop has now pushed well past it. Anthropic is not alone in this race. Simular, OpenAI, and Google DeepMind have all contributed competitive systems to the leaderboard.

Why the 85% number needs a footnote #

OSWorld was designed when agents were scoring in the teens. As top systems close in on a human baseline of 72% and then surpass it, the benchmark stops being a useful discriminator.

OSWorld 2.0 is the harder test, and the numbers it produces are sobering. On this updated benchmark, which includes longer, more complex tasks averaging 1.6 human hours to complete, the best agents are finishing only about 20.6% of assignments correctly. That is not a typo in the research sense, just a jarring contrast: 85% on the old test, roughly one in five on the new one.

There is also a step-count problem. Even on the annotated versions of benchmarks where top agents succeed, they require between 1.4 and 2.7 times more steps than a human would take to complete the same task.

What this means for the technology’s trajectory #

The OSWorld 2.0 result at 20.6% suggests the technology is genuinely capable for a defined class of tasks, specifically shorter, repeatable, well-structured workflows, while still falling short on the kind of open-ended work that fills most knowledge workers’ afternoons.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our

Editorial Policy.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/computer-use-agents-…] indexed:0 read:2min 2026-08-10 ·