cd /news/artificial-intelligence/arc-agi-leaderboard · home topics artificial-intelligence article
[ARTICLE · art-73075] src=arcprize.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

ARC-AGI Leaderboard

The ARC-AGI-3 leaderboard, released by the ARC Prize team, ranks AI systems on their ability to adapt to novel interactive environments, measuring performance against cost-per-task. The leaderboard shows reasoning systems like o1-pro and Gemini 3 Pro achieving high accuracy but at higher costs, while Kaggle systems demonstrate efficient solutions under a $50 compute budget. Only systems costing less than $10,000 to run are included, with some results marked as preview.

read1 min views1 publishedJul 25, 2026
ARC-AGI Leaderboard
Image: source

Understanding the Leaderboard

ARC-AGI has evolved from its first versions (ARC-AGI-1 and 2) which measured passive fluid intelligence, to ARC-AGI-3 which challenges AI agents to adapt on the fly to novel interactive environments.

The scatter plot above visualizes the critical relationship between cost-per-task and performance - a key measure of efficiency. True intelligence isn't just about solving problems, but solving them efficiently with minimal resources.

Interpreting the data

Reasoning Systems Trend Line solutions display connected points representing the same model at different reasoning levels. These trend lines illustrate how increased reasoning time affects performance, typically showing asymptotic behavior as thinking time increases.Base LLMs solutions represent single-shot inference from standard language models like GPT-4.5 and Claude 3.7, without extended reasoning capabilities. These points demonstrate raw model performance without additional reasoning enhancements.Kaggle Systems solutions showcase competition-grade submissions from the Kaggle challenge, operating under strict computational constraints ($50 compute budget for 120 evaluation tasks). These represent purpose-built, efficient methods specifically designed for the ARC Prize.

Verification Policy

For more information, see our testing policy.

Leaderboard Breakdown

Notes

Only systems which required less than $10,000 to run are shown.

For models that were not able to produce full test out puts, remaining tasks were marked as incorrect. Results marked as "preview" are unofficial and may be based on incomplete testing.

1 ARC-AGI-2 score estimate based on partial testing results and o1-pro pricing.

2 Provisional cost estimates based on Gemini 3 Pro pricing. Model to be retested once released.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arc-agi-3 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/arc-agi-leaderboard] indexed:0 read:1min 2026-07-25 ·