13:27
2026-08-29
dev.to
ai-agents
We ran 160 agent tasks across two frameworks. The frameworks tied. Then we changed the model.
A benchmark of 160 agent tasks found LangGraph and Pydantic AI tied at 100% correctness with gpt-4o, but swapping to gpt-4o-mini dropped overall correctness to 75%, with one date-arithmetic task failiβ¦