{"slug": "can-ai-agents-actually-use-a-computer-here-s-the-real-answer", "title": "Can AI Agents Actually Use a Computer? Here's the Real Answer", "summary": "Claude Fable 5 ranked top of OSWorld-Verified with 85% in June 2026, according to llm-stats, against a human baseline of 72.36, up from the best model's 42% a year earlier. By August 29, BenchLM's tracking of the same benchmark had Qwen3.8 Max in front at 86.1%, with Claude Mythos 5 and Claude Fable 5 tied at 85. The article argues that benchmark scores on the 369-task desktop suite do not translate into finished business processes, since agents can complete a task on screen without verifying the outcome.", "body_md": "# **The Verification Bottleneck**\n\nAn agent files an insurance claim. The portal accepts it, the screen confirms, and a dashboard somewhere turns the task green.\n\nTwo days later an adjuster calls the office to verify a policy number before the claim can move, and the phone rings in an empty room.\n\nSo the claim stalls until a human notices the payment that never arrived, which can take a quarter, if anyone notices at all.\n\nThe agent did everything right. It read the screen, filled every field and clicked the right buttons, and it still had no way to learn that correct and finished are different things.\n\nThe industry is still debating whether models can operate a desktop at all, and the answer is yes. Companies pay for finished work, though, delivered without someone standing behind the machine checking every line, and the gap between clicking and finishing is where this year’s deployments either hold up or fall apart.\n\n*together with Plaid*\n\n**An agent that can file a legitimate claim can file a fake one just as fast. Most [identity checks](https://plaid.com/new-identity-crisis-ai-fraud-report/?utm_source=AICorner&utm_medium=PaidNewsletter&utm_campaign=AICorner_Paid_Newsletter_Ad_Buy&utm_content=EvolutionIdentityFraud_Primary) still assume a human at the keyboard.**\n\nPlaid’s latest white paper traces [how fraud evolved](https://plaid.com/new-identity-crisis-ai-fraud-report/?utm_source=AICorner&utm_medium=PaidNewsletter&utm_campaign=AICorner_Paid_Newsletter_Ad_Buy&utm_content=EvolutionIdentityFraud_Primary) across each internet wave, and why this one needs a new approach. \n\nInside:\n\n▫️ Why [point-in-time verification](https://plaid.com/new-identity-crisis-ai-fraud-report/?utm_source=AICorner&utm_medium=PaidNewsletter&utm_campaign=AICorner_Paid_Newsletter_Ad_Buy&utm_content=EvolutionIdentityFraud_Primary) is breaking down\n\n▫️ How [AI-driven fraud](https://plaid.com/new-identity-crisis-ai-fraud-report/?utm_source=AICorner&utm_medium=PaidNewsletter&utm_campaign=AICorner_Paid_Newsletter_Ad_Buy&utm_content=EvolutionIdentityFraud_Primary) is scaling faster than traditional controls\n\n▫️ Why financial behavior offers stronger, harder-to-fake context\n\n▫️ How broader, [cross-platform visibility](https://plaid.com/new-identity-crisis-ai-fraud-report/?utm_source=AICorner&utm_medium=PaidNewsletter&utm_campaign=AICorner_Paid_Newsletter_Ad_Buy&utm_content=EvolutionIdentityFraud_Primary) reveals patterns others miss\n\n## **Table of Contents**\n\n1. What The Current Benchmarks Are\n\n2. A Benchmark Score and a Business Process Are Different Units\n\n3. The Best Architecture in Production Is Fifteen Years Old\n\n4. The Number That Settles This\n\n5. The Attack Surface Is the Product\n\n6. Why the Back Offices Are Still Hiring\n\n## **1. What The Current Benchmarks Are**\n\n[Claude Fable 5](https://www.the-ai-corner.com/p/claude-fable-5-guide) ranked top of OSWorld-Verified with 85%, sourced to llm-stats in June 2026. Human testers score around 72% on the same task set.\n\n**A year earlier the best model managed just 42%.**\n\nBy the 29th of August, [BenchLM’s tracking of the same benchmark](https://benchlm.ai/benchmarks/osworld-verified) had Qwen3.8 Max in front at 86.1%, with Claude Mythos 5 and Claude Fable 5 tied behind at 85.\n\nEleven weeks. New leader, different lab, different continent.\n\nNormalised against the **human baseline of 72.36** and the arc still deserves respect. Last year’s best model sat at 58% of human performance and this year’s is at 119%. Something real crossed over in early 2026.\n\n### **What a score of 85 contains**\n\nOSWorld-Verified is a set of **369 desktop tasks**, and a model passes one by actually **finishing it** rather than describing how.\n\nEight of the tasks can be left out under the official rules. Some results get re-run by the people who maintain the leaderboard and others are simply reported by the labs, [using different time limits, different operating systems and different tools](https://leaderboard.steel.dev/leaderboards/osworld/).\n\nThere is also a newer version, OSWorld 2.0, whose scores are not comparable with the older one.\n\n**So an 85 belongs to a model, plus the software wrapped around it**, plus how many steps it was allowed to take. Two companies both claiming 85 may not be running the same thing at all.\n\n[Andreessen Horowitz published a piece in August](https://a16z.com/can-agents-use-a-computer-yet-weve-got-the-data/), where they interviewed the people actually running these systems, and one of them turned out to be **running millions of automated tasks a month without knowing which model does the work**. He had never needed to find out. His supplier swaps models underneath him the way a cloud provider swaps servers.\n\nThat detail says more about where computer-use agents have got to than any leaderboard does. When the heaviest users stop checking which model they are on, the model has stopped being the thing that decides whether the work gets done.\n\n## **2. Per Task and Per Step Are Different Maths**\n\nWhat a benchmark measures and what a company actually buys are two completely different things. Here’s the explanation.\n\n### **Task success and step success**\n\n**A benchmark scores whole tasks,** while a real office job is a **chain of steps**, and every step has to work for the job to count as done.\n\nSo if a job has 20 steps and each step works 95% of the time, the job works **36% of the time**. That is 0.95 multiplied by itself twenty times. Here is what the same sum looks like across a few reliability levels.\n\nThis is simplified. **A well-built system retries and recovers**, so real numbers land better than the table. Read it as the shape of the problem rather than a prediction.\n\n### **Why a better model buys less than it looks**\n\n**Agents hold up on short jobs that follow a fixed script and come apart on long ones.** That is exactly the pattern their interviews describe, without ever naming the reason underneath it.\n\nGoing from 95 to 98 per step nearly doubles end-to-end success at 20 steps, and still leaves you failing two thirds of the time at 50.\n\nModel improvements arrive in a straight line, but **failure compounds**. That is an unfair fight, and it is why the last few points on a leaderboard matter far less at work than they do in a launch, and why the people running this at volume have all landed on the same answer.\n\n## **3. Why the Winning Architecture Is Fifteen Years Old**\n\nThe design pattern serious teams advocate for is being seen as an emerging AI-native idea. But that is not the case. More like far from it.\n\n### **What run-caching does**\n\nThe a16z piece describes the pattern well. The agent runs a workflow once, the system caches it as deterministic code, every subsequent run executes as cheap code, and the model is only called back when something breaks so it can diagnose, fix and re-cache.\n\nThey frame this as **cost optimisation**.\n\nBut it’s more like **reliability architecture,** and what it does is collapse N stochastic steps into one checkable event. You basically stop rolling the dice 50 times and start rolling it once, at the moment the portal changes.\n\n### **What the scraper case shows**\n\nTheir flagship case, reported first-hand, is a **CPG data platform running 15 to 20 million automated portal interactions a month**, which halved the engineering team assigned to scraper maintenance. Excellent outcome, real money.\n\nHowever, this is actually using **hand-coded scrapers** to do the work. The agent is the repair crew that fixes them when a retailer changes the page.\n\nThat is not an agent using a computer. That is an agent fixing the thing that uses the computer.\n\nThe conclusion is that software that clicks through screens on a fixed script has existed for 15 years, under the name **Robotic Process Automation**, and it always broke the moment a website changed.\n\nIf the honest story right now is that AI models are the fix for that 15-year-old flaw, that is a genuinely valuable business worth building. It is also a much smaller claim than the one printed on the tin.\n\nThe same tension runs through the speed fix they point to. **Reading a page’s underlying structure** instead of looking at a picture of it is faster, and it ties you straight back to how that particular application is built. Screenshots removed that dependency.\n\nThis puts it back. Speed bought with fragility is not free.\n\n## **4. The Number That Settles This**\n\nHere’s the metric that does not exist, and its absence is the most interesting thing about this whole conversation.\n\nThe a16z cost comparison puts an agent at **$6-8 an hour,** against roughly 10 for offshore BPO and 30 to 45 for a US back-office worker, the last built up from a BLS median of 20.59 with benefits at about 30% of total comp. Fine as far as it goes.\n\nBut an hour of agent time is not the thing a buyer purchases. A completed, **correct task is.**\n\nLet’s assume **$7** per hour, **9 minutes** per agentic task, gives you about **$1.05 an attempt**. At **85%** first-pass success with independent retries you’re looking at roughly **1.18 attempts**, so $1.24. Now price the 15% that need a person. 10 minutes of a US operator at 35 fully loaded is $5.83, and 0.15 of that adds 87 cents. You land near $2.10 per completed task.\n\n**Double the headline**. And that’s the optimistic version, because it assumes every failure gets caught.\n\n**Nobody is claiming $2.10 is the real number.** The real number doesn’t exist yet and neither does anybody quoting $6-8 an hour. However, cost per verified completed task, including escalation, **is the only figure that decides whether any of this works**, and I cannot find a single vendor publishing it.\n\nSame for the second missing metric. Nobody reports a silent error rate.\n\nThe a16z piece names the failure mode precisely and then walks past it. An agent transcribing payment terms into an **ERP reads net 60 as net 30**.\n\nThe record looks plausible as it passes every visual check. It surfaces weeks later when an invoice goes out wrong. An agent files an insurance claim, the screen confirms receipt, and two days later an adjuster phones about a policy number the agent will never know was asked for.\n\nA loud failure costs you a retry. A silent one costs you the invoice, the relationship and the audit. Those are not the same 15%, and no benchmark on earth distinguishes them.\n\n## **5. The Attack Surface Is the Product**\n\n### **What prompt injection is**\n\nAn agent reads the instructions from its owner and the text on the webpage as the same kind of thing: **words on a screen**. It has no reliable way of telling which words came from the person paying for it. So anyone who can get text in front of the agent can give it orders.\n\nPicture a temp who does whatever any note left on the desk says, because nobody explained which notes come from the boss. Hide a note on a page the agent is going to read, and the note becomes an instruction.\n\n### **Three conditions, all of them shipped on purpose**\n\nThe recurring finding across [serious prompt injection incidents](https://www.sysdig.com/learn-cloud-native/prompt-injection) is one configuration where an **agent has access to private data, exposure to untrusted content, and the ability to act externally**.\n\nThe use case is that an agent holding live enterprise credentials, logging into government and insurance portals it does not control, taking actions on the other side of them.\n\nAll three conditions. By design. As the product.\n\nIn May 2026 the Five Eyes intelligence agencies published joint guidance naming **prompt injection as a main route for manipulating these systems**, and told organisations to assume their agents will sometimes behave in ways nobody planned for. Attackers who adjust their approach get past almost every published defence.\n\nScreenshot-driven agents add a channel that text-only systems simply do not have. Brave’s security team [demonstrated it against Perplexity Comet](https://arxiv.org/pdf/2511.19477) using white text on white backgrounds and instructions buried in HTML comments, which got the agent [fetching one-time passwords out of email](https://plaid.com/new-identity-crisis-ai-fraud-report/?utm_source=AICorner&utm_medium=PaidNewsletter&utm_campaign=AICorner_Paid_Newsletter_Ad_Buy&utm_content=EvolutionIdentityFraud_Primary) when the user had only asked it to summarise a page.\n\n### **The same blind spot, with far worse consequences**\n\nThe conditions under which a **silent error goes undetected** are precisely the conditions under which a successful injection goes undetected. Same gap in the process, wildly different bill at the end.\n\nTwo further questions sit alongside it.\n\nThe first is **who pays when an agent files** **the wrong thing** with a regulator? Is it the supplier, the buyer, or the lab that built the model?\n\nNobody has settled that, and it stalls deals in legal review long after the trial run went well.\n\nThe second is **permission** from the other side. The main use case here is automating someone else’s website, and plenty of them forbid automated access in their terms and run software to spot it. Some will push back.\n\n*together with Plaid*\n\n**The agents that click through portals can also spoof the signals identity checks rely on.** \n\nPlaid’s white paper explains why [point-in-time verification is breaking down](https://plaid.com/new-identity-crisis-ai-fraud-report/?utm_source=AICorner&utm_medium=PaidNewsletter&utm_campaign=AICorner_Paid_Newsletter_Ad_Buy&utm_content=EvolutionIdentityFraud_Primary), and what holds up better.\n\n## **6. What the Hiring Data Shows**\n\nIf agents at $7 an hour genuinely substituted for offshore labour at $10, **the offshore numbers would be falling**. They are not, and that disagreement is the most useful data in this entire debate.\n\n### **What the employment numbers show**\n\nNone of what follows appears in their piece.\n\n[Global contact-centre employment is forecast to grow from 15.3 million in 2025 to 16.8 million by 2029](https://news.outsourceaccelerator.com/outsourcing-keeps-growing/), even while automation suppresses an estimated **1.9 million individual roles** across the same window.\n\n[The Philippines produced $40 billion in BPO export revenue](https://bpoai.ai/news/philippine-bpo-industry-hits-40b-in-2025-eyes-59b-by-2028) with **1.9 million workers in 2025** and is tracking toward **$42.3 billion and 1.96 million in 2026**.\n\n[Sam Altman said in May 2026](https://www.the-ai-corner.com/p/sam-altman-startup-school-2026) that he was delighted to have **been wrong about customer support roles disappearing**, reversing his own position from ten months earlier. Against him, [Vinod Khosla ahead of the India AI Impact Summit](https://m.economictimes.com/tech/technology/exciting-to-see-this-much-interest-tech-giant-vinod-khosla-on-india-ai-impact-summit/articleshow/128556260.cms): “IT and BPO services will disappear, almost certainly within the next five years.”\n\nOne of those readings will look foolish in three years. I do not know which, and I would be suspicious of anybody who claims they do.\n\n### **What the industry is repositioning around**\n\nThe firms being valued upward are the ones **building the checking layer**. That includes **supervision, quality control, and handling the cases that go wrong**.\n\nThat surviving human function is the **verification layer**, sitting in plain sight in the labour data, which is their own argument turning up in the employment statistics.\n\nWhich brings the moat claim into focus. **Context is the moat**. That means the written procedures, the unwritten know-how, who to escalate to.\n\nAgainst a model provider that defends beautifully. Against another startup it defends less well, and their own examples suggest why.\n\nIf the knowledge lives in one recorded video of somebody doing the job once, that video is cheap to make and any competitor can use it. Procedures and logins belong to the customer, not to whoever holds them this quarter.\n\nWhat they are describing is a **switching cost.** It makes leaving annoying rather than impossible, and it gets weaker as setting up a new supplier gets cheaper, which is exactly what they forecast will happen.\n\nThe more durable version is the one their buyers gave them directly. Being the vendor an enterprise is permitted to run in production. Security review, audit trail, liability terms, procurement sign-off. Slow to earn, slow to lose, and impossible to show off on a stage.\n\n**So, can agents use a computer.** Yes, and that question closed months ago without much ceremony at all.\n\nThe one that replaced it is whether you can tell when they got it wrong, cheaply enough and fast enough for it to matter. Where you can, the economics already beat every human alternative and the deployments are real and growing. Where you cannot, a smarter model changes absolutely nothing, because the model was never the thing that failed.\n\nDoing the work stopped being the hard part somewhere around February.\n\nVouching for it never was easy.", "url": "https://wpnews.pro/news/can-ai-agents-actually-use-a-computer-here-s-the-real-answer", "canonical_source": "https://www.the-ai-corner.com/p/can-ai-agents-use-a-computer", "published_at": "2026-10-08 14:52:23+00:00", "updated_at": "2026-10-08 15:16:57.119638+00:00", "lang": "en", "topics": ["ai-agents", "artificial-intelligence", "large-language-models"], "entities": ["Claude Fable 5", "OSWorld-Verified", "Qwen3.8 Max", "Claude Mythos 5", "llm-stats", "BenchLM", "Plaid"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/can-ai-agents-actually-use-a-computer-here-s-the-real-answer", "markdown": "https://wpnews.pro/news/can-ai-agents-actually-use-a-computer-here-s-the-real-answer.md", "text": "https://wpnews.pro/news/can-ai-agents-actually-use-a-computer-here-s-the-real-answer.txt", "jsonld": "https://wpnews.pro/news/can-ai-agents-actually-use-a-computer-here-s-the-real-answer.jsonld"}}