Artificial Analysis Capability Indices v1.1
Artificial Analysis released Capability Indices v1.1 on September 14, 2026, adding Agentic Tool Use and Agentic Knowledge Work evaluations while removing Agentic Customer Interaction (𝜏³-Banking) acro…
Artificial Analysis released Capability Indices v1.1 on September 14, 2026, adding Agentic Tool Use and Agentic Knowledge Work evaluations while removing Agentic Customer Interaction (𝜏³-Banking) acro…
Moonshot AI's Kimi K3 ranks second on the AA-Briefcase benchmark, trailing only Anthropic's Claude Fable 5, but faces high operational costs with nearly an hour per task and approximately ten times th…
Kimi K3 from Moonshot AI has surpassed OpenAI's GPT-5.6 Sol in agentic knowledge work, achieving an Elo rating of 1547 compared to GPT-5.6 Sol's 1495 on the AA-Briefcase benchmark, which evaluates mod…
Anthropic's Claude Sonnet 5 achieves a score of 53 on the Artificial Analysis Intelligence Index, matching GPT-5.5 with high reasoning but costing $2.29 per task—15% more than Opus 4.8—due to increase…
A new evaluation benchmark, AA-Briefcase, measures frontier knowledge work performance, with models like Claude Opus 4.8 averaging 24 minutes per task and achieving an Elo of 1356, while MiniMax-M3 ta…