cd /news/ai-research/artificial-analysis-capability-indic… · home topics ai-research article
[ARTICLE · art-132406] src=artificialanalysis.ai ↗ pub= topic=ai-research verified=true sentiment=· neutral

Artificial Analysis Capability Indices v1.1

Artificial Analysis released Capability Indices v1.1 on September 14, 2026, adding Agentic Tool Use and Agentic Knowledge Work evaluations while removing Agentic Customer Interaction (𝜏³-Banking) across its Finance & Accounting, Strategy & Ops, Legal, and Healthcare & Medical indices. The update maps O*NET occupation tasks to benchmarks weighted by capability frequency, incorporating Intelligence Index v4.2 and v4.3 slices; Strategy & Ops now weights Agentic Tool Use at 30%, up from 0% in v1.0, while Finance & Accounting weights it at 10% and Legal at 5%. The Engineering index swaps Terminal-Bench v2.1 for Terminal-Bench v4.0 and drops GPQA Diamond, and Healthcare & Medical adds Long-Context Reasoning from MLCR-AA.

read5 min views2 publishedSep 17, 2026
Artificial Analysis Capability Indices v1.1
Image: source

All articles September 14, 2026

We are updating the Capability Indices v1.1 with stronger domain tuning, combining slices of core Intelligence Index v4.3 evaluations alongside specialized evaluations.

The Capability Indices map tasks from O*NET occupations to the benchmarks that represent them, weighting each benchmark by how often its capability appears across tasks. Capability Indices v1.1 incorporates updates to Intelligence Index v4.2 and v4.3, slicing by domain when available.

| Index | Added | Deleted |

|---|---|---|
| Finance & Accounting | Agentic Tool Use (AutomationBench-AA, Finance) Agentic Knowledge Work (AA-Briefcase) Long-Context (GDP.pdf) | Agentic Customer Interaction (𝜏³-Banking) | 
| Strategy & Ops | Agentic Tool Use (AutomationBench-AA, Operations) Agentic Knowledge Work (AA-Briefcase) Long-Context (GDP.pdf) | Agentic Customer Interaction (𝜏³-Banking) | 
| Legal | Agentic Tool Use (AutomationBench-AA, Operations and Support) Agentic Knowledge Work (AA-Briefcase) Long-Context (GDP.pdf) | Agentic Customer Interaction (𝜏³-Banking) | 

| Healthcare & Medical | Agentic Tool Use (AutomationBench-AA, Operations and Support) Long-Context Reasoning (MLCR-AA) Agentic Knowledge Work (AA-Briefcase) | Agentic Customer Interaction (𝜏³-Banking) |

| Engineering | Agentic Terminal Use (Terminal-Bench v4.0) Agentic Knowledge Work (AA-Briefcase) | Agentic Terminal Use (Terminal-Bench v2.1) Reasoning (GPQA Diamond) | 
| Economics | Agentic Knowledge Work (AA-Briefcase) | - | 

Finance & Accounting

Agentic Tool Use added, sourced from the Finance slice of AutomationBench-AA. Agentic Customer Interaction removed. AA-Briefcase added alongside GDPval-AA v2 in Agentic Knowledge Work, and GDP.pdf added alongside LCR in Long-Context.

Capability Evaluations v1.0 v1.1
Business Knowledge AA-Omniscience 30% 30%
Agentic Knowledge Work GDPval-AA v2, AA-Briefcase 30% 30%
Reasoning HLE 20% 20%
Agentic Tool Use AutomationBench-AA 0% 10%
Long-Context LCR, GDP.pdf 5% 5%
Non-Hallucination AA-Omniscience 5% 5%

Strategy & Ops

Agentic Tool Use added, sourced from the Operations slice of AutomationBench-AA to reflect the importance of tool use in day-to-day operational tasks. Agentic Customer Interaction removed. AA-Briefcase added alongside GDPval-AA v2 in Agentic Knowledge Work, and GDP.pdf added alongside LCR in Long-Context.

Capability Evaluations v1.0 v1.1
Agentic Knowledge Work GDPval-AA v2, AA-Briefcase 35% 35%
Business Knowledge AA-Omniscience 30% 30%
Agentic Tool Use AutomationBench-AA 0% 30%
Long-Context LCR, GDP.pdf 5% 5%

Legal

Agentic Tool Use added, sourced from the Operations and Support slices of AutomationBench-AA to reflect the role of tool use in legal workflows. Agentic Customer Interaction removed. AA-Briefcase added alongside GDPval-AA v2 in Agentic Knowledge Work, and GDP.pdf added alongside LCR in Long-Context.

Capability Evaluations v1.0 v1.1
Legal Knowledge AA-Omniscience 35% 35%
Agentic Knowledge Work GDPval-AA v2, AA-Briefcase 25% 25%
Reasoning HLE 15% 15%
Long-Context LCR, GDP.pdf 10% 10%
Non-Hallucination AA-Omniscience 10% 10%
Agentic Tool Use AutomationBench-AA 0% 5%

Healthcare & Medical

Long-Context Reasoning added, sourced from MLCR-AA. This evaluates reasoning across lengthy clinical records. Agentic Tool Use added, sourced from the Operations and Support slices of AutomationBench-AA to cover operational workflows relevant to healthcare. Agentic Customer Interaction removed. AA-Briefcase added alongside GDPval-AA v2 in Agentic Knowledge Work.

Capability Evaluations v1.0 v1.1
Medical & Health Knowledge AA-Omniscience 35% 30%
Agentic Knowledge Work GDPval-AA v2, AA-Briefcase 25% 25%
Long-Context Reasoning MLCR-AA 0% 15%
Non-Hallucination AA-Omniscience 15% 10%
Reasoning HLE 15% 10%
Agentic Tool Use AutomationBench-AA 0% 10%

Engineering

Agentic Terminal Use updated to Terminal-Bench v4.0. GPQA Diamond removed from Reasoning. AA-Briefcase added alongside GDPval-AA v2 in Agentic Knowledge Work.

Capability Evaluations v1.0 v1.1
Engineering Knowledge AA-Omniscience 35% 35%
Reasoning HLE, CritPt 35% 30%
Agentic Knowledge Work GDPval-AA v2, AA-Briefcase 25% 20%
Agentic Terminal Use Terminal-Bench 4.0 5% 15%

Economics

AA-Briefcase added alongside GDPval-AA v2 in Agentic Knowledge Work.

Capability Evaluations v1.0 v1.1
Economics Knowledge AA-Omniscience 35% 35%
Reasoning HLE 35% 35%
Agentic Knowledge Work GDPval-AA v2, AA-Briefcase 15% 25%
Long-Context Reasoning LCR 15% 5%

Full results and methodology #

- Finance & Accounting: [https://artificialanalysis.ai/models/capabilities/finance-and-accounting](https://artificialanalysis.ai/models/capabilities/finance-and-accounting)
- Strategy & Ops: [https://artificialanalysis.ai/models/capabilities/strategy-and-ops](https://artificialanalysis.ai/models/capabilities/strategy-and-ops)
- Legal: [https://artificialanalysis.ai/models/capabilities/legal](https://artificialanalysis.ai/models/capabilities/legal)
- Healthcare & Medical: [https://artificialanalysis.ai/models/capabilities/healthcare-and-medical](https://artificialanalysis.ai/models/capabilities/healthcare-and-medical)
- Engineering: [https://artificialanalysis.ai/models/capabilities/engineering](https://artificialanalysis.ai/models/capabilities/engineering)
- Economics: [https://artificialanalysis.ai/models/capabilities/economics](https://artificialanalysis.ai/models/capabilities/economics)

Explore all Capability Indices: [https://artificialanalysis.ai/models/capabilities](https://artificialanalysis.ai/models/capabilities)

Read the methodology: [https://artificialanalysis.ai/methodology/capability-indices](https://artificialanalysis.ai/methodology/capability-indices)

Read the latest

### Ant Group releases finance-focused Ling-3.0-flash-Fin

Ant Group has released their finance-focused flash model Ling-3.0-flash-Fin

September 16, 2026

Benchmarking GPT-6 Astra

GPT-6 Astra ties leadership with Claude Fable 5.1 in both of our flagship Indices, at lower cost. Astra equals Fable 5.1 in the Intelligence Index at ~40% of the cost, and in the Coding Agent Index at ~60% of the cost.

September 9, 2026

Announcing the Artificial Analysis Intelligence Index v4.3

We are upgrading Terminal-Bench to 4.0 and adding AutomationBench-AA, an agentic workflow automation benchmark with a private test set. This is a continuation of our rollout of Intelligence Index v5.

September 7, 2026

── more in #ai-research 4 stories · sorted by recency
── more on @artificial analysis 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/artificial-analysis-…] indexed:0 read:5min 2026-09-17 ·