{"slug": "computer-use-agents-hit-85-on-osworld-benchmark-up-from-12-in-two-years", "title": "Computer-use agents hit 85% on OSWorld benchmark, up from 12% in two years", "summary": "Anthropic's Claude Mythos Preview, Fable 5, and Opus 4.8 top the OSWorld-Verified leaderboard with scores of 85.4%, 85.0%, and 83.4%, up from 12% in April 2024, surpassing the human baseline of 72%. However, on the harder OSWorld 2.0 benchmark, the best agents complete only 20.6% of tasks, and successful agents require 1.4 to 2.7 times more steps than humans, indicating limitations in open-ended tasks.", "body_md": "# Computer-use agents hit 85% on OSWorld benchmark, up from 12% in two years\n\nRapid gains on a popular AI benchmark mask a harder truth: real-world tasks still trip up the best agents by a wide margin\n\nTwo years ago, AI agents that operate computers the way humans do, by looking at a screen and deciding where to click, were completing roughly one in eight assigned tasks correctly. Today, the best systems are clearing 85% on the same test.\n\nThe benchmark in question is OSWorld, currently the most widely cited evaluation platform for multimodal computer-use agents (CUAs). These are AI systems that navigate a desktop environment using screenshots, mouse clicks, and keystrokes rather than direct API calls, essentially watching a screen and acting on what they see.\n\n## The numbers behind the leap\n\nIn April 2024, leading agents were scoring around 12% on OSWorld. By mid-2025 that figure had climbed into the mid-30s. By June 2026, three Anthropic models sit at the top of the OSWorld-Verified leaderboard: Claude Mythos Preview at 85.4%, Fable 5 at 85.0%, and Opus 4.8 at 83.4%.\n\nFor context, the human baseline on OSWorld sits at approximately 72%. When Simular’s Agent S3 crossed that threshold in December 2025, reaching 72.6%, it was the first widely noted instance of a CUA surpassing average human performance on the test. Anthropic’s current crop has now pushed well past it.\n\nAnthropic is not alone in this race. Simular, OpenAI, and Google DeepMind have all contributed competitive systems to the leaderboard.\n\n## Why the 85% number needs a footnote\n\nOSWorld was designed when agents were scoring in the teens. As top systems close in on a human baseline of 72% and then surpass it, the benchmark stops being a useful discriminator.\n\nOSWorld 2.0 is the harder test, and the numbers it produces are sobering. On this updated benchmark, which includes longer, more complex tasks averaging 1.6 human hours to complete, the best agents are finishing only about 20.6% of assignments correctly. That is not a typo in the research sense, just a jarring contrast: 85% on the old test, roughly one in five on the new one.\n\nThere is also a step-count problem. Even on the annotated versions of benchmarks where top agents succeed, they require between 1.4 and 2.7 times more steps than a human would take to complete the same task.\n\n## What this means for the technology’s trajectory\n\nThe OSWorld 2.0 result at 20.6% suggests the technology is genuinely capable for a defined class of tasks, specifically shorter, repeatable, well-structured workflows, while still falling short on the kind of open-ended work that fills most knowledge workers’ afternoons.\n\n**Disclosure:** This article was edited by Editorial Team. For more information on how we create and review content, see our\n\n[Editorial Policy](https://cryptobriefing.com/editorial-policy/).", "url": "https://wpnews.pro/news/computer-use-agents-hit-85-on-osworld-benchmark-up-from-12-in-two-years", "canonical_source": "https://cryptobriefing.com/computer-use-agents-85-percent-osworld-benchmark/", "published_at": "2026-08-10 14:50:09+00:00", "updated_at": "2026-08-10 15:13:41.710384+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-agents"], "entities": ["Anthropic", "Claude Mythos Preview", "Fable 5", "Opus 4.8", "OSWorld", "Simular", "OpenAI", "Google DeepMind"], "alternates": {"html": "https://wpnews.pro/news/computer-use-agents-hit-85-on-osworld-benchmark-up-from-12-in-two-years", "markdown": "https://wpnews.pro/news/computer-use-agents-hit-85-on-osworld-benchmark-up-from-12-in-two-years.md", "text": "https://wpnews.pro/news/computer-use-agents-hit-85-on-osworld-benchmark-up-from-12-in-two-years.txt", "jsonld": "https://wpnews.pro/news/computer-use-agents-hit-85-on-osworld-benchmark-up-from-12-in-two-years.jsonld"}}