cd /news/ai-agents/how-ora-benchmarks-every-major-ai-ag… · home topics ai-agents article
[ARTICLE · art-106328] src=vercel.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

How Ora benchmarks every major AI agent on Vercel

Ora, a platform that benchmarks AI agents on live websites, has tested every major agent including Claude Code, ChatGPT, Gemini, Hermes, OpenClaw, and Vercel's eve, and found that 99% of the web isn't agent-ready. In a comparison of eve against Claude Code using the same models, eve achieved 7% fewer steps, 2x native success, and 9% more valid endpoints, leading Ora to build on eve. The platform runs entirely on Vercel, including the agent runtime.

read4 min views1 publishedAug 21, 2026

Front end, back end, and agent runtime on one platform

Every major agent tested side by side on live sites

Hundreds of commits a day from a 16-person engineering team

Ora sends agents onto live websites with instructions to sign up for a product, integrate with it, and pay for it. Agents often fail, and by Ora's estimate, 99% of the web isn't agent-ready. The platform shows customers where and why agents fail, and what to change.

Assaf Elovic, co-founder of Ora, spent years helping agents discover the web. His previous company, Tavily, built a web search engine for AI agents and was acquired by Nebius earlier this year. Search solved half the problem, but an agent that finds a product still has to actually use it. He and co-founder Liad Yosef started Ora to measure how ready the web is for agents, and to fix the parts that aren't.

Today that means spawning agents against live customer sites from journey.ora.ai, where Ora runs a journey and records the cost, latency, and steps an agent needs to finish a task. The platform runs on Vercel, including the agent runtime.

Agents decompose into two parts: a model, which does the reasoning, and a harness, the software that gives the model its tools and drives it from step to step. Ora's lineup covers the agents customers use most: Claude Code, ChatGPT, Gemini, Hermes, OpenClaw, and eve, Vercel's agent framework. Ora runs each agent on a customer's website and watches how it handles common workflows.

No two harnesses want the same infrastructure. Each expects its own environment and exposes its steps differently, so Ora runs a separate runtime for every harness and traces every step. Ido Finder, who leads engineering at Ora, calls that side-by-side coverage one of the most valuable things ora brings to its customers.

When an agent stalls in a signup flow, the customer sees which step and what it tried. Without the trace, the result is a score with nothing behind it.

Ora built this testing system entirely on Vercel. Front end, back end, and the agent runtime share the same deployment path, logs, and authentication. The agent runtime isn't separate infrastructure to operate; it lives where the rest of the product lives.

When Vercel launched eve, Ora gave it no special treatment. eve went through the same benchmark, under the same conditions, as every other harness in the lineup. Ora works with Vercel Engineering as a design partner, and Finder gave the eve team direct access to the platform to dig into the results.

The initial test put eve against Claude Code across hundreds of real journeys on multiple domains. Both harnesses ran the same models, Claude Fable 5 and Haiku 4.5, and every run gave the agent the same job: integrate with a product.

Ora published three numbers from the comparison:

7% fewer steps to reach the goal

2x native success: twice as many tasks finished on the customer's own site instead of falling back to web search

9% more valid endpoints: more of the endpoints the agent found were ones it could actually call

The benchmark fed back into eve, too. One run surfaced a prompt-caching issue, the eve team shipped a fix, and Ora's next round of results measured roughly 15% lower total cost.

For a company that benchmarks every major harness for a living, this is not a casual choice. After those results, Ora builds on eve. Because eve follows the Next.js paradigm, there was little to configure, and tools, skills, and connectors take little code. The feature that sealed it was the sandbox override. An agent framework like eve ships with its own sandbox, the isolated environment where the agent executes, runs tools, and touches files. That's a good default for most teams, because you get safe execution for free. But an agent in the framework's own sandbox runs outside the instrumented environment where Ora traces every step. The override lets the team swap that environment in, so eve agents get recorded like every other harness, with nothing new built.

Journey.ora.ai now has eve on both sides, eve is one of the harnesses it tests, and eve is what it runs on.

The engineering team is 16 people shipping hundreds of commits a day, and day-to-day infrastructure work belongs to their coding agents. A stack that sits in one place lets those agents run it end to end.

Finder puts the time saved at a few hours a week at least. Elovic credits a similar amount to how well coding agents build with Vercel's libraries.

Ora is adding more products, and the architecture is growing with them. The team is splitting its platform into microservices, all of them on Vercel. New services deploy to the same infrastructure and talk to each other with no extra configuration, and the internal agents built on eve will run as one more service.

By Ora's own measure, 99% of the web still can't handle an agent that shows up to sign up, integrate, and pay.

About Ora: Ora makes the web agent-ready. Companies use ora to benchmark how AI agents discover, navigate, and interact with their websites, and to build the infrastructure agents need to find, use, and transact with them.

── more in #ai-agents 4 stories · sorted by recency
── more on @ora 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-ora-benchmarks-e…] indexed:0 read:4min 2026-08-21 ·