Artificial Analysis Capability Indices v1.1
Artificial Analysis released Capability Indices v1.1 on September 14, 2026, adding Agentic Tool Use and Agentic Knowledge Work evaluations while removing Agentic Customer Interaction (𝜏³-Banking) acro…
Artificial Analysis released Capability Indices v1.1 on September 14, 2026, adding Agentic Tool Use and Agentic Knowledge Work evaluations while removing Agentic Customer Interaction (𝜏³-Banking) acro…
Anthropic introduced Claude Fable 5.1 and Claude Mythos 5.1, the same model with different safeguard levels, claiming they are the world's most advanced models for coding and knowledge work. Fable 5.1…
Anthropic's Claude Sonnet 5, priced at $2/$10 per million tokens during an introductory period (standard $3/$15 after August 2026), matches or nearly matches the flagship Claude Opus 4.8 on knowledge …
Artificial Analysis launched an Agentic Index, a weighted average of agentic benchmarks including GDPval-AA v2 and τ³-Banking, to measure AI capability signals. The index's top score was not disclosed…
Meta released Muse Spark 1.2 and Muse Code, a terminal coding agent co-trained with the model, which scored #5 on GDPval-AA v2 and showed cost-efficient performance, with gains concentrated in agentic…
Anthropic released Claude Sonnet 5, which outperforms its predecessor Sonnet 4.6 across all benchmarks and surpasses the larger Opus 4.8 on the GDPval-AA v2 knowledge work test with a score of 1,618. …