The YC S26 founders claim $60,000 in early revenue, though CoArena's site currently says new tasks and judging are d.
By Ryan Merket · Published
Primary source: Y Combinator
Why it matters #
Computer-use agents are improving faster than fixed benchmarks can remain useful. CoArena is betting that free access can generate the live tasks, failure traces and human labels AI labs will buy.
Prateek Jannu and Nitish Kovuru have launched CoArena, a service that sends two computer-use agents through the same user-submitted task and asks people to judge the results without knowing which model produced them.
The two-person San Francisco operation is part benchmark, part free agent product and part data collection business. Users supply live work such as researching flights or assembling expense reports. CoArena runs the task twice, reveals the model identities after a vote and uses the preference data to update its leaderboard.
Jannu, CoArena's CEO, dropped out of Stanford to build the product. Kovuru, who serves as CTO and COO, studied computer science at Columbia after earning a computer engineering degree from Purdue. Before CoArena, the pair built Coasty, a computer-use agent for automating work in legacy desktop software and third-party portals.
That earlier product shaped their case against fixed benchmarks. Coasty reported a score of 82.81% on OSWorld Verified, a benchmark covering desktop applications such as browsers, office software and operating-system tools. Jannu and Kovuru argue that building for OSWorld showed them how a published task set can become less useful as agents improve, developers optimize against it and benchmark items circulate through training data.
CoArena replaces that fixed test set with an incoming stream of tasks from users. Both agents receive the same sandbox, action vocabulary and step budget. Their positions are randomized, and vendor names are scrubbed before judging. A battle that exposes a model's identity in its visible output is excluded from the rating calculation.
The approach is designed to measure performance on work that a model provider did not select in advance. It also produces the asset CoArena intends to sell: paired trajectories containing agent actions, observations, reasoning when available, errors and human preference labels.
The free benchmark feeds the paid product
CoArena's data page describes this output plainly: "The exhaust is the product." Coasty Systems plans to license preference data, trajectories and evaluation access under contracts priced by consumption. A public sample includes voted battles and the first three actions from each run, while the commercial tiers retain longer traces and additional battle metadata.
Screenshots are excluded from the licensed datasets. CoArena says submitted text and URLs are redacted before delivery, although raw screen captures still pass to the model provider operating each agent during a live run. That distinction matters because user-generated computer tasks can expose account details, private documents or third-party content before a redaction system processes the resulting record.
The commercial structure also explains why CoArena is free to use. Each task supplies fresh evaluation material, and each vote adds a preference label that could make the dataset more useful to model developers. Users receive the agents' output; Coasty Systems retains an expanding corpus for licensing.
In an August 23rd launch post, the founders said CoArena had been available for about two weeks, that its user base had doubled each week and that it had generated $60,000 in revenue. Those figures are company-supplied. The post does not break out paying customers, signed contract value, recognized revenue or the share attributable specifically to CoArena rather than other Coasty Systems work.
CoArena also reported that roughly 66% of agent runs completed their tasks, with the strongest model finishing about 87% and the weakest about 42%. Human judges marked both agents as failures in roughly 7% of battles, according to the launch post. CoArena did not attach battle counts or model-level denominators to those launch statistics, limiting comparisons with results from larger, fixed benchmarks.
CoArena is benchmarking its suppliers
CoArena's neutrality problem is unusually direct: Coasty Systems operates the leaderboard while also building a computer-use agent. Its governance policy excludes Coasty's agent from new matchmaking and from the ranked roster. Three earlier Coasty battles remain in the underlying record, including one that received a vote, but they do not appear in the ranking.
The published leaderboard uses a Bradley-Terry model fitted across eligible blind votes rather than treating a running Elo score as the final ranking. CoArena says it refits the ratings from the corpus so retracted votes and judgments from accounts later marked untrusted can be removed. Models with fewer than 30 rated battles receive provisional treatment, and the service publishes confidence intervals intended to show how thin records affect placement.
The methodology is more explicit than the usual benchmark landing page, but CoArena still depends on the size and behavior of its crowd. In an August 7th methodology post, the founders acknowledged that the early ranking rested on a small number of judges and that many battles were labeled by their submitters. CoArena clusters rating uncertainty by judge and prioritizes independent votes when they exist, but a young arena cannot manufacture an independent annotator base through statistical design alone.
As of August 30th, CoArena's homepage says posting and judging are d while recorded battles remain available to watch. The site does not give a reason for the . Its leaderboard remains online, led by Claude Fable 5, followed by Gemini 3.7 Flash and GPT-5.6 Sol at the time of publication.
That leaves CoArena's core loop temporarily incomplete. Its pitch depends on a continuing supply of unseen tasks and blind votes, while its business depends on turning those interactions into evaluation and training data that AI labs will pay to use. The founders have built detailed machinery for explaining what each score means. Their next constraint is keeping enough people inside the machinery to make the scores matter.