cd /news/ai-agents/how-we-benchmark-ai-agents-and-tools… · home › topics › ai-agents › article
[ARTICLE · art-145430] src=arize.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

How we benchmark AI agents and tools with Harbor and Arize Phoenix

Arize AI detailed a benchmarking setup that pairs the open-source Harbor framework with its Phoenix observability platform to evaluate AI agents, MCP servers, skills, and CLIs in reproducible sandboxes. Harbor runs each attempt in a fresh sandbox using local Docker or supported cloud providers, while a Phoenix plugin records Harbor tasks as versioned datasets, each agent configuration as an experiment, and verifier rewards as Phoenix annotations, plus an infra_ok annotation flagging execution errors. Arize built the approach after its single-turn evals for PXI, the coding agent in Phoenix, failed to supply frontend context like the current page and time and could not test what the agent did with real tool results.

by read4 min views5 publishedOct 5, 2026
How we benchmark AI agents and tools with Harbor and Arize Phoenix
Image: Arize (auto-discovered)

Why is agent benchmarking so hard? #

Offline experimentation is one of the hardest parts of agent development. Once an agent is deployed, we can observe what it does in the wild. Before deployment, benchmarking requires a reproducible setup: the same tasks, data, tools, and environment for every version, with state reset between runs.

We ran into this problem when trying to evaluate PXI, the coding agent built into Phoenix. We started with simple agent evals that tested PXI’s tool calling behavior on single turns, but we kept running into limitations. Our fixtures were missing context that the frontend normally supplied, like the current page and time. Some tools needed the frontend in order to execute at all. Adding an MCP server for skills and docs meant another thing to mock. And without a real Phoenix database, we could check whether PXI called a tool, but barely test what it did with the result.

Expanding that suite meant rebuilding or mocking more and more parts of the system. It was getting harder to trust that this Frankenstein version of PXI behaved like the one in production.

Run AI agent benchmarks in reproducible sandboxes with Harbor #

Harbor is the open-source framework behind coding-agent benchmarks like Terminal-Bench, but you don’t have to work at an AI lab to use it. Give it a task, an environment with your application and data, and a verifier to check success. Harbor runs each attempt in a fresh sandbox, using local Docker or your choice of supported cloud sandbox providers.

For PXI, this meant running the real agent endpoint against a Phoenix server pre-seeded with traces. We could test what PXI did with real tool results and inspect the database afterwards.

Benchmark MCP servers, skills, and CLIs with Harbor #

You don't have to build your own agent to need agent experiments, though. If you ship an MCP server, skills, or a CLI, you want to know whether agents can use them effectively. Harbor natively supports coding agents like Claude Code and Codex alongside custom agents.

With Harbor, we can reuse the same tasks to test PXI, our MCP server, and the Phoenix CLI. If we ask an agent to annotate failing traces, we can check that it wrote the annotations to the right records in the database. We can also inspect the agent’s trajectory to evaluate its response and measure how many turns or tool calls it took. The same verifiers work across agents, measuring correctness and efficiency without prescribing exactly how each agent should get the job done.

Record Harbor benchmarks as Phoenix experiments and traces #

The Phoenix plugin records Harbor tasks as versioned datasets, each agent configuration as an experiment, and verifier rewards as Phoenix annotations. It also turns Harbor's recorded trajectory into Phoenix traces, so we can inspect the model calls, tool use, and available usage data behind a reward without instrumenting the agent.

The plugin records all rewards produced by Harbor’s verifiers as Phoenix annotations. It also adds an infra_ok annotation to every run to flag execution errors.

When tasks change, the plugin creates a new dataset version, while earlier experiments stay tied to the version they ran against. That helps us distinguish an agent improvement from a change to the test.

Those results are also accessible through Phoenix’s MCP server and CLI. We can ask a coding agent to compare experiments or investigate traces, then review the evidence behind its findings ourselves.

Use traces to improve agents even when the benchmark passes #

In our benchmark, we tested PXI, Claude Code, and Codex on the same questions about a Phoenix project. Almost every answer was correct, but the traces showed meaningful differences in how each agent reached the answer.

In one trial, Codex correctly diagnosed a failing tool but followed our error-analysis skill into a much longer workflow than the question needed. The trace showed us where to focus: help the agent distinguish a quick diagnosis from a full investigation. With the benchmark in place, we can test that change across tasks and check whether it reduces turns and tool calls without sacrificing correctness.

Run a baseline, change something, repeat #

Here’s how to get started with Harbor and the Phoenix plugin. Pick one task your users care about, give the agent a realistic environment, and check the outcome. Run a baseline, inspect the trace, then change a prompt, skill, or tool and rerun the same task. For us, Harbor makes that loop practical without rebuilding PXI's world in mocks.

Once that loop is in place, we can automate parts of it. Optimizers like GEPA use evaluation feedback to propose and test prompt changes. The same Harbor tasks can measure those changes, while Phoenix traces help us track and understand their effects.

The Harbor integration guide explains more about the Phoenix plugin, and our internal benchmark suite may provide some inspiration for how to set this up for your own systems.

We're walking through this setup live, including the PXI benchmark, at Benchmarking AI agents & tool use with Harbor and Arize Phoenix on Oct 8, 2026.

── more in #ai-agents 4 stories · sorted by recency
── more on @arize ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-we-benchmark-ai-…] indexed:0 read:4min 2026-10-05 · —