# Kimi K3: second only to Fable 5 on AA-Briefcase

> Source: <https://artificialanalysis.ai/articles/kimi-k3-agentic-knowledge-benchmark>
> Published: 2026-07-22 04:33:44+00:00

[All articles](/articles)

July 21, 2026

# Kimi K3: second only to Fable 5 on AA-Briefcase

**Kimi K3 is second only to Fable 5 on AA-Briefcase, our agentic knowledge work benchmark, but costs more than Opus 4.8 to run while averaging nearly an hour per task**

Last week Kimi (Moonshot AI) released Kimi K3, a 2.8T parameter model that scores 57 on the Artificial Analysis Intelligence Index, comparable to models such as Opus 4.8 and GPT-5.5. On AA-Briefcase, Kimi K3 scores an Elo of 1543, a +727 improvement over Kimi K2.6 and the second highest score recorded, behind only Claude Fable 5 (1574)

AA-Briefcase is our new proprietary benchmark for agentic knowledge work, testing models on a fully private dataset of realistic tasks across thousands of complex input files. Tasks require deliverables such as spreadsheets, presentations, and UI mock-ups, with performance combined into a single AA-Briefcase Elo based on correctness, analytical quality, and presentation quality
**Key results for Kimi K3 on AA-Briefcase:**

**➤ Second only to Fable 5:** Kimi K3 achieves an AA-Briefcase Elo of 1543, the second-highest score overall, ahead of GPT-5.6 Sol (max, 1501), Claude Sonnet 5 (max, 1388), and Claude Opus 4.8 (max, 1347). This is a +727 improvement over the previous-generation Kimi K2.6 (816) and puts Kimi K3 only behind Fable 5

**➤ Strong objective and analytical performance, with comparatively weaker presentation:** Kimi K3 achieves a rubric pass rate of 51%, second only to Claude Fable 5 (56%) and ahead of Claude Sonnet 5 (max, 42.3%) and GPT-5.6 Sol (max, 41.8%). It also records an analytical quality Elo of 1754, comparable to Claude Fable 5 (1744). Presentation quality is comparatively weaker, with a Presentation Elo of 1471, below GPT-5.6 Sol (max, 1660) and Claude Opus 4.8 (max, 1492)

**➤ ~10x increase in Cost per Task:** Kimi K3 averages a cost of $10.57 per task, placing it among the most expensive models to run on AA-Briefcase. This is driven by model token pricing, increased output tokens and relatively high turn use, averaging 83 turns per task, versus 67 for Claude Fable 5 and 50 for GPT-5.6 Sol (max). Kimi K3 is priced at $3/$15 per 1M input/output tokens, with a 90% discount for cached tokens

**➤ Averages nearly an hour per task:** Kimi K3 has an average Time per AA-Briefcase Task of 56.4 minutes. This is driven by a high number of turns, as well as higher output token use and lower speeds using the first party Kimi API. Kimi K3 uses 120k output tokens per task and 83 turns per task, up from 42k output tokens and 54 turns with Kimi K2.6. This average Time per Task is one of the highest recorded on AA-Briefcase, ~2.5x that of Claude Fable 5 and ~3.8x higher that of Grok 4.5 (high)

For full AA-Briefcase results, see [https://artificialanalysis.ai/evaluations/aa-briefcase](https://artificialanalysis.ai/evaluations/aa-briefcase)

#### Read the latest

### Gemini 3.6 Flash and Gemini 3.5 Flash-Lite: Halving Time per Task

Google has released Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. Both halve time per task relative to their predecessors and increase token efficiency, Gemini 3.5 Flash-Lite improves by 11 Intelligence Index points while Gemini 3.6 Flash does not improve in intelligence over 3.5 Flash

July 21, 2026

### Four frontier launches in eight days: six labs now field a model above 50 on the Artificial Analysis Intelligence Index

Yeah. Grok 4.5, GPT-5.6, Muse Spark 1.1, and Kimi K3 all launched within eight days. Six labs now have a model scoring above 50 on the Artificial Analysis Intelligence Index, up from two in early June - and the price of near-frontier intelligence has collapsed

July 17, 2026

### Kimi K3 achieves #3 in the Artificial Analysis Intelligence Index, comparable to Opus 4.8 and GPT-5.5

Benchmarks and Analysis of Kimi K3

July 17, 2026
