Understand the AI landscape to choose the best model and provider for your use case
Highlights
Intelligence #
Intelligence of leading AI models based on our independent evaluations
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index by Open Weights / Proprietary
Cost per Intelligence Index Task
Intelligence Index vs. Cost per Intelligence Index Task
Intelligence Index vs. Cost per Intelligence Index Task, by Model Release
Frontier Language Model Intelligence, Over Time
Performance, cost, and execution time for leading coding agents on end-to-end software engineering tasks
Artificial Analysis Coding Agent Index
Image & Video #
Top models from our Image Arena and Video Arena leaderboards, with 95% confidence intervals
Text to Image Leaderboard
See the full leaderboard here.
Speech #
Top models from our Text to Speech Arena, Speech to Text and Speech to Speech evaluations
Provider Voice Arena Quality Elo
Measures the performance of models on specific capabilities and industries
Artificial Analysis Finance & Accounting Index
Intelligence Evaluations
Agentic real-world work tasks, (Elo-500)/2000 Agentic tool use
Agentic coding & terminal use
Coding
Reasoning & knowledge
Scientific reasoning
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Agentic knowledge work, Elo
Agentic SaaS workflows
Legal agentic work, criterion pass rate
Agentic business operations
Quantitative analysis on spreadsheets & documents
Instruction following
Long-horizon agentic tasks
Kubernetes incident root-cause analysis
Visual reasoning
AA-Briefcase AA-Briefcase is a frontier agentic evaluation for long-horizon knowledge work, testing agents on realistic business workflows that require deliverables such as spreadsheets, presentations, and memos
AA-Briefcase Elo
AA-AnalystAgent AA-AnalystAgent is a benchmark for end-to-end quantitative analysis on real-world spreadsheets and documents — the kind of work business and data analysts do every day
AA-AnalystAgent pass^5
AA-Omniscience AA-Omniscience is a knowledge and hallucination benchmark that rewards accuracy, punishes bad guesses and provides a comprehensive view of which models produce factually reliable outputs across different domains
AA-Omniscience Index
GDPval-AA v2 GDPval-AA v2 evaluates AI models on real-world, economically valuable tasks across a wide range of occupations
GDPval-AA v2 Leaderboard
Artificial Analysis Openness Index assesses how 'open' models are on the basis of their availability and transparency across different components.
Artificial Analysis Openness Index: Components
Artificial Analysis Openness Index vs. Artificial Analysis Intelligence Index
Output Tokens #
Output tokens of leading AI models based on our independent evaluations
Output Tokens per Intelligence Index Task
Cost #
Price and real-world costs of leading AI models based on our independent evaluations
Cost per Intelligence Index Task
Cost to Run Artificial Analysis Intelligence Index
Pricing: Cache Hit, Input, and Output
Speed & Latency #
Comparison of first-party API performance