Most framework comparisons argue from feature lists. I wanted numbers, so agentic-arena holds the model, tools, datasets and iteration budget fixed and measures what each framework does differently.
Read this first: everything below is measured against a scripted (mock) model, so the turns are byte-identical for every framework. That makes these wire and behaviour measurements. They say nothing about which framework writes better answers.
Mean prompt tokens per item on a 15-item tool-use task:
| framework | prompt tokens | vs baseline |
|---|---|---|
| vanilla (stdlib loop) | 753.5 | 1.00x |
| langgraph | 753.5 | 1.00x |
| pydantic_ai | 794.0 | 1.05x |
| microsoft_af | 802.0 | 1.06x |
| google_adk | 836.1 | 1.11x |
| openai_agents | 856.9 | 1.14x |
| smolagents | 2935.5 | 3.90x |
Six of seven sit inside a 1.15x band, and none is leaner than the hand-rolled loop. smolagents' 3.90x is a 4,207-character system prompt where the arena asked for 384, including a prose restatement of tools it already sent as a schema. The tools are transmitted, and billed, twice.
That is a worst case. The same conversation run to 30 turns shrinks the gap from 8.83x on the first request to 1.27x on the 31st, because a fixed overhead decays as the conversation grows.
I scripted 429, 500 and 400 responses. The hand-rolled baseline has no retry at all, so one 429 loses the item. Every framework survives a single 429. Only smolagents survives three in a row, by quietly sleeping roughly two to four minutes on one item. The item passes, so nothing in a scorecard shows it, but in a batch your throughput quietly collapses.
A three-role pipeline doubles the LLM calls and multiplies prompt tokens by 2.5x, whether you build it with a graph library or a or loop. The structure costs that, not the framework.
Model-decided handoffs cost about 10% more prompt than the same pipeline wired structurally. Of that gap, 94% is the ransfer_to_* tool schemas, which ride on every request whether or not anyone delegates. You pay for the options you offer, not the ones you use.
Five of the seven adapters never passed the shared request timeout to their client, so they waited out their library's default through a 20-second hang on a one-second budget. A hanging mock provider exposed it. It is fixed, and gated in CI.
Everything above has a command that regenerates it, and CI regenerates each number on a clean install:
git clone https://github.com/code-with-rashid/agentic-arena
cd agentic-arena
python -m pip install -e .
python -m arena run --arena tool_use --framework all --mode mock --no-scorecard
Full findings and the reasoning behind each: https://code-with-rashid.github.io/agentic-arena/findings/