Artificial Analysis Capability Indices v1.1 Artificial Analysis released Capability Indices v1.1 on September 14, 2026, adding Agentic Tool Use and Agentic Knowledge Work evaluations while removing Agentic Customer Interaction (𝜏³-Banking) across its Finance & Accounting, Strategy & Ops, Legal, and Healthcare & Medical indices. The update maps O*NET occupation tasks to benchmarks weighted by capability frequency, incorporating Intelligence Index v4.2 and v4.3 slices; Strategy & Ops now weights Agentic Tool Use at 30%, up from 0% in v1.0, while Finance & Accounting weights it at 10% and Legal at 5%. The Engineering index swaps Terminal-Bench v2.1 for Terminal-Bench v4.0 and drops GPQA Diamond, and Healthcare & Medical adds Long-Context Reasoning from MLCR-AA. All articles https://artificialanalysis.ai/articles September 14, 2026 Announcing Artificial Analysis Capability Indices v1.1 We are updating the Capability Indices v1.1 with stronger domain tuning, combining slices of core Intelligence Index v4.3 evaluations alongside specialized evaluations. The Capability Indices map tasks from O NET occupations to the benchmarks that represent them, weighting each benchmark by how often its capability appears across tasks. Capability Indices v1.1 incorporates updates to Intelligence Index v4.2 and v4.3, slicing by domain when available. Changelog | Index | Added | Deleted | |---|---|---| | Finance & Accounting | Agentic Tool Use AutomationBench-AA, Finance Agentic Knowledge Work AA-Briefcase Long-Context GDP.pdf | Agentic Customer Interaction 𝜏³-Banking | | Strategy & Ops | Agentic Tool Use AutomationBench-AA, Operations Agentic Knowledge Work AA-Briefcase Long-Context GDP.pdf | Agentic Customer Interaction 𝜏³-Banking | | Legal | Agentic Tool Use AutomationBench-AA, Operations and Support Agentic Knowledge Work AA-Briefcase Long-Context GDP.pdf | Agentic Customer Interaction 𝜏³-Banking | | Healthcare & Medical | Agentic Tool Use AutomationBench-AA, Operations and Support Long-Context Reasoning MLCR-AA Agentic Knowledge Work AA-Briefcase | Agentic Customer Interaction 𝜏³-Banking | | Engineering | Agentic Terminal Use Terminal-Bench v4.0 Agentic Knowledge Work AA-Briefcase | Agentic Terminal Use Terminal-Bench v2.1 Reasoning GPQA Diamond | | Economics | Agentic Knowledge Work AA-Briefcase | - | Changes by Index Finance & Accounting Agentic Tool Use added, sourced from the Finance slice of AutomationBench-AA. Agentic Customer Interaction removed. AA-Briefcase added alongside GDPval-AA v2 in Agentic Knowledge Work, and GDP.pdf added alongside LCR in Long-Context. | Capability | Evaluations | v1.0 | v1.1 | |---|---|---|---| | Business Knowledge | AA-Omniscience | 30% | 30% | | Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | 30% | 30% | | Reasoning | HLE | 20% | 20% | | Agentic Tool Use | AutomationBench-AA | 0% | 10% | | Long-Context | LCR, GDP.pdf | 5% | 5% | | Non-Hallucination | AA-Omniscience | 5% | 5% | Strategy & Ops Agentic Tool Use added, sourced from the Operations slice of AutomationBench-AA to reflect the importance of tool use in day-to-day operational tasks. Agentic Customer Interaction removed. AA-Briefcase added alongside GDPval-AA v2 in Agentic Knowledge Work, and GDP.pdf added alongside LCR in Long-Context. | Capability | Evaluations | v1.0 | v1.1 | |---|---|---|---| | Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | 35% | 35% | | Business Knowledge | AA-Omniscience | 30% | 30% | | Agentic Tool Use | AutomationBench-AA | 0% | 30% | | Long-Context | LCR, GDP.pdf | 5% | 5% | Legal Agentic Tool Use added, sourced from the Operations and Support slices of AutomationBench-AA to reflect the role of tool use in legal workflows. Agentic Customer Interaction removed. AA-Briefcase added alongside GDPval-AA v2 in Agentic Knowledge Work, and GDP.pdf added alongside LCR in Long-Context. | Capability | Evaluations | v1.0 | v1.1 | |---|---|---|---| | Legal Knowledge | AA-Omniscience | 35% | 35% | | Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | 25% | 25% | | Reasoning | HLE | 15% | 15% | | Long-Context | LCR, GDP.pdf | 10% | 10% | | Non-Hallucination | AA-Omniscience | 10% | 10% | | Agentic Tool Use | AutomationBench-AA | 0% | 5% | Healthcare & Medical Long-Context Reasoning added, sourced from MLCR-AA. This evaluates reasoning across lengthy clinical records. Agentic Tool Use added, sourced from the Operations and Support slices of AutomationBench-AA to cover operational workflows relevant to healthcare. Agentic Customer Interaction removed. AA-Briefcase added alongside GDPval-AA v2 in Agentic Knowledge Work. | Capability | Evaluations | v1.0 | v1.1 | |---|---|---|---| | Medical & Health Knowledge | AA-Omniscience | 35% | 30% | | Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | 25% | 25% | | Long-Context Reasoning | MLCR-AA | 0% | 15% | | Non-Hallucination | AA-Omniscience | 15% | 10% | | Reasoning | HLE | 15% | 10% | | Agentic Tool Use | AutomationBench-AA | 0% | 10% | Engineering Agentic Terminal Use updated to Terminal-Bench v4.0. GPQA Diamond removed from Reasoning. AA-Briefcase added alongside GDPval-AA v2 in Agentic Knowledge Work. | Capability | Evaluations | v1.0 | v1.1 | |---|---|---|---| | Engineering Knowledge | AA-Omniscience | 35% | 35% | | Reasoning | HLE, CritPt | 35% | 30% | | Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | 25% | 20% | | Agentic Terminal Use | Terminal-Bench 4.0 | 5% | 15% | Economics AA-Briefcase added alongside GDPval-AA v2 in Agentic Knowledge Work. | Capability | Evaluations | v1.0 | v1.1 | |---|---|---|---| | Economics Knowledge | AA-Omniscience | 35% | 35% | | Reasoning | HLE | 35% | 35% | | Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | 15% | 25% | | Long-Context Reasoning | LCR | 15% | 5% | Full results and methodology - Finance & Accounting: https://artificialanalysis.ai/models/capabilities/finance-and-accounting https://artificialanalysis.ai/models/capabilities/finance-and-accounting - Strategy & Ops: https://artificialanalysis.ai/models/capabilities/strategy-and-ops https://artificialanalysis.ai/models/capabilities/strategy-and-ops - Legal: https://artificialanalysis.ai/models/capabilities/legal https://artificialanalysis.ai/models/capabilities/legal - Healthcare & Medical: https://artificialanalysis.ai/models/capabilities/healthcare-and-medical https://artificialanalysis.ai/models/capabilities/healthcare-and-medical - Engineering: https://artificialanalysis.ai/models/capabilities/engineering https://artificialanalysis.ai/models/capabilities/engineering - Economics: https://artificialanalysis.ai/models/capabilities/economics https://artificialanalysis.ai/models/capabilities/economics Explore all Capability Indices: https://artificialanalysis.ai/models/capabilities https://artificialanalysis.ai/models/capabilities Read the methodology: https://artificialanalysis.ai/methodology/capability-indices https://artificialanalysis.ai/methodology/capability-indices Read the latest Ant Group releases finance-focused Ling-3.0-flash-Fin Ant Group has released their finance-focused flash model Ling-3.0-flash-Fin September 16, 2026 Benchmarking GPT-6 Astra GPT-6 Astra ties leadership with Claude Fable 5.1 in both of our flagship Indices, at lower cost. Astra equals Fable 5.1 in the Intelligence Index at ~40% of the cost, and in the Coding Agent Index at ~60% of the cost. September 9, 2026 Announcing the Artificial Analysis Intelligence Index v4.3 We are upgrading Terminal-Bench to 4.0 and adding AutomationBench-AA, an agentic workflow automation benchmark with a private test set. This is a continuation of our rollout of Intelligence Index v5. September 7, 2026