cd /news/ai-agents/what-agent-frameworks-cost-on-the-wi… · home › topics › ai-agents › article
[ARTICLE · art-142665] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

What agent frameworks cost on the wire: measurements from agentic-arena

A developer built agentic-arena, a benchmark that holds the model, tools, datasets and iteration budget fixed to measure how seven agent frameworks differ on the wire. Six of seven frameworks stayed within a 1.15x band of a hand-rolled stdlib loop on prompt tokens, while smolagents used 3.90x because it restates tool schemas in prose, and five of seven adapters failed to pass a shared request timeout to their client, hanging 20 seconds on a one-second budget.

by read2 min views1 publishedSep 30, 2026

Most framework comparisons argue from feature lists. I wanted numbers, so agentic-arena holds the model, tools, datasets and iteration budget fixed and measures what each framework does differently.

Read this first: everything below is measured against a scripted (mock) model, so the turns are byte-identical for every framework. That makes these wire and behaviour measurements. They say nothing about which framework writes better answers.

Mean prompt tokens per item on a 15-item tool-use task:

framework prompt tokens vs baseline
vanilla (stdlib loop) 753.5 1.00x
langgraph 753.5 1.00x
pydantic_ai 794.0 1.05x
microsoft_af 802.0 1.06x
google_adk 836.1 1.11x
openai_agents 856.9 1.14x
smolagents 2935.5 3.90x

Six of seven sit inside a 1.15x band, and none is leaner than the hand-rolled loop. smolagents' 3.90x is a 4,207-character system prompt where the arena asked for 384, including a prose restatement of tools it already sent as a schema. The tools are transmitted, and billed, twice.

That is a worst case. The same conversation run to 30 turns shrinks the gap from 8.83x on the first request to 1.27x on the 31st, because a fixed overhead decays as the conversation grows.

I scripted 429, 500 and 400 responses. The hand-rolled baseline has no retry at all, so one 429 loses the item. Every framework survives a single 429. Only smolagents survives three in a row, by quietly sleeping roughly two to four minutes on one item. The item passes, so nothing in a scorecard shows it, but in a batch your throughput quietly collapses.

A three-role pipeline doubles the LLM calls and multiplies prompt tokens by 2.5x, whether you build it with a graph library or a or loop. The structure costs that, not the framework.

Model-decided handoffs cost about 10% more prompt than the same pipeline wired structurally. Of that gap, 94% is the ransfer_to_* tool schemas, which ride on every request whether or not anyone delegates. You pay for the options you offer, not the ones you use.

Five of the seven adapters never passed the shared request timeout to their client, so they waited out their library's default through a 20-second hang on a one-second budget. A hanging mock provider exposed it. It is fixed, and gated in CI.

Everything above has a command that regenerates it, and CI regenerates each number on a clean install:

git clone https://github.com/code-with-rashid/agentic-arena
cd agentic-arena
python -m pip install -e .
python -m arena run --arena tool_use --framework all --mode mock --no-scorecard

Full findings and the reasoning behind each: https://code-with-rashid.github.io/agentic-arena/findings/

── more in #ai-agents 4 stories · sorted by recency
── more on @agentic-arena 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/what-agent-framework…] indexed:0 read:2min 2026-09-30 · —