# Artificial Analysis Updates Intelligence Index Twice In 2 Days, Fable 5.1 & GPT-6 Astra Now Tied For First Place

> Source: <https://officechai.com/ai/artificial-analysis-updates-intelligence-index-twice-in-2-days-fable-5-1-gpt-6-astra-now-tied-for-first-place/>
> Published: 2026-09-07 18:50:33+00:00

AI benchmarks seem to be changing as rapidly as AI models as we approach the singularly.

Just two days after its last [update](https://officechai.com/ai/artificial-analysis-rejigs-intelligence-index-with-v4-2-fable-5-1-tops-gpt-6-astra-placed-second/), Artificial Analysis has once again revised its widely-followed Intelligence Index, this time rolling out version 4.3. The benchmarking outfit says the update is part of its ongoing rollout of changes originally planned for a bigger Index v5, and this time the shakeup has produced a new joint leader: Claude Fable 5.1 (max, with fallback) and GPT-6 Astra (max) are now tied at the top of the leaderboard, both scoring 53 points.

### **What Changed In Artificial Analysis Intelligence Index v4.3**

The v4.3 update makes two changes to the underlying evaluation suite. Terminal-Bench, which tests whether an AI agent can complete complex, multi-step work through a terminal, has been upgraded from version 2.1 to the newly released version 4.0. The new version recalibrates compute and time allowances, sharpens task instructions and verification, and switches the evaluation harness from Terminus 2 to mini-SWE-agent, a minimal, model-agnostic harness. It comprises 66 tasks spanning software engineering, machine learning, science, operations, security, hardware, and media, each run three times with the average pass rate reported.

The second change swaps out τ³-Banking for AutomationBench-AA, a new agentic workflow benchmark built in collaboration with Zapier. It tests models on 657 simulated business workflows across apps like Gmail, Slack, Salesforce, and Jira, spanning categories such as Finance, HR, Marketing, Operations, Sales, and Support. Models are scored on the share of task objectives they complete, with any violation of a stated guardrail dropping that task’s score to zero. Because AutomationBench-AA uses a private, held-out test set, the weighting assigned to evaluations with private tasks or private answers rises from 40% to 45% of the overall Index, even though the four top-level category weights — Agents 30%, Coding 20%, General 30%, and Scientific Reasoning 20% — stay unchanged.

### **The New Top Of The Leaderboard**

With those changes in place, Claude Fable 5.1 (max, with fallback) and GPT-6 Astra (max) are now tied for first place with a score of 53 apiece. Anthropic’s Claude Opus 5 (max) sits just behind at 51, followed by Claude Fable 5 (with fallback) at 50 — meaning Anthropic still occupies three of the top four positions on the Index, even after losing sole possession of first place. Meta’s Muse Spark 1.3 (max) rounds out the top five with a score of 48, ahead of GPT-5.6 Sol (max) at 47.

On the individual evaluations behind the composite score, the two new tests reveal a clearer gap than the tied headline number suggests. On the upgraded Terminal-Bench 4.0, GPT-6 Astra (max) scores 59.1%, ahead of Claude Fable 5.1 (max, with fallback) at 52.0% and Claude Opus 5 (max) at 49.0%, putting Astra 19 percentage points clear of GPT-5.6 Sol (max) on the same test. On AutomationBench-AA, GPT-6 Astra (max) again leads with a score of 68.5%, ahead of Grok 4.6 (high) at 66.7% and GLM-5.3 (max) at 62.2%. Astra also completes every objective in a workflow without tripping a guardrail on 41.6% of tasks, compared with 32.1% for Fable 5.1 and 28.3% for Opus 5 — a reminder that fully following the rules on a workflow remains meaningfully harder than partially completing it.

Among open-weight models, Z AI’s GLM-5.3 and Moonshot’s Kimi K3 continue to lead the pack, both scoring 44, a full 9 points behind the joint leaders. GLM-5.3-Flash (42) is the next-strongest open-weights model, followed by Alibaba’s Qwen3.8 2.4T A95B (40) and DeepSeek V4 Pro 0813 (max) at 36.

On the cost-efficiency side, Artificial Analysis says four labs — OpenAI, Anthropic, Meta, and Z AI — now occupy the Intelligence-vs-Cost Pareto frontier. OpenAI dominates the lower and middle portions of that frontier, with all five reasoning tiers of GPT-6 Astra offering the cheapest cost-per-task at their respective intelligence levels, while Claude Fable 5.1 sits at the very top of the frontier as the most capable — and most expensive — option, alongside GLM-5.3-Flash and Xiaomi’s MiMo-V2.5-Pro at the cheaper end.

**How This Compares To v4.2 And v4.1**

The speed of these revisions is itself notable. Under [v4.2](https://officechai.com/ai/artificial-analysis-rejigs-intelligence-index-with-v4-2-fable-5-1-tops-gpt-6-astra-placed-second/), published just two days before v4.3, Claude Fable 5.1 held sole first place with a score of 57, a clear 2 points ahead of GPT-6 Astra’s 55. That version had already been a rebalancing exercise of its own, folding in early elements of the planned Index v5 after GPT-6 Astra’s launch failed to show the leap over GPT-5.6 Sol that some had expected.

Before that, under Index v4.1, Claude Fable 5.1 led with a score of 66, three points ahead of Claude Opus 5’s 63, with Meta’s Muse Spark 1.3 close behind at 62.

The biggest “loser” across the three versions is arguably Claude Fable 5.1 in relative terms — not because the model changed, but because each successive version of the Index has chipped away at its lead, first from 9 points over the nearest rival down to 2, and now to zero. OpenAI’s GPT-6 Astra, by contrast, has climbed from a tie for the score that [appeared “disappointing”](https://officechai.com/ai/openai-gpt-6-astra-scores-a-disappointing-61-on-artificial-analysis-intelligence-index-same-as-gpt-5-6-sol/) at launch, to now sharing the very top spot, largely on the strength of the harder agentic tests introduced in v4.2 and v4.3. Another loser is Google Gemini 3.8, which is no longer on the [Pareto frontier](https://officechai.com/ai/gemini-3-8-flash-is-on-pareto-frontier-of-intelligence-and-cost-on-artificial-analysis-intelligence-index/) like it was on v4.1. The “winners” of this particular reshuffle, then, are less about raw capability gains and more about which benchmarks got added: Terminal-Bench 4.0 and AutomationBench-AA both happen to favour OpenAI’s agentic and terminal-driven workflows, pulling GPT-6 Astra level with Anthropic’s flagship.

### A Word Of Caution On Fast-Moving Indices

All this raises a fair question: what does it mean for a benchmark to crown a new “leader” every couple of days? Artificial Analysis has now revised its flagship Intelligence Index three times in a matter of weeks, each time swapping out one or more underlying evaluations, reweighting private test sets, or recalibrating difficulty. Every individual change may well be justified on its own terms — reducing saturation, closing gaming loopholes, keeping pace with genuinely improving models — but the cumulative effect is a leaderboard whose “number one” position has become almost as unstable as the models it is meant to measure. If a benchmark’s ranking can flip from a 9-point lead for one company to a flat tie within two revisions, largely because of which new evaluations were added rather than any change in the models themselves, it becomes harder to treat any single snapshot as a durable signal of which AI system is actually “best.” Used well, these indices remain a genuinely useful cross-section of capability at a moment in time — but the moment they update this often, they risk becoming more a running commentary on Artificial Analysis’s own methodology than a stable measure of the state of the art.
