cd/entity/Artificial Analysis· home› entities› Artificial Analysis
grep -l @artificial analysis /news/*.json | wc -l → 460

Artificial Analysis

mentions 460 type Person page 19/23 feed RSS

// recent coverage 460 mentions

22:28
2026-07-16
tokenstead.ai
artificial-intelligence

Kimi K3

Moonshot AI released Kimi K3, a 2.8-trillion-parameter mixture-of-experts model with 16 active experts per token, achieving the top score of 1679 on the Artificial Analysis webdev arena ahead of Claud…

20:09
2026-07-16
artificialanalysis.ai
artificial-intelligence

Kimi K3 Intelligence, Performance and Price Analysis

Kimi K3, a reasoning model released on July 16, 2026, by Kimi, scores 57 on the Artificial Analysis Intelligence Index, well above the average of 30 among comparable models, but is slower than average…

19:29
2026-07-16
aws.amazon.com
artificial-intelligence

Introducing Grok on Amazon Bedrock

XAI's Grok 4.3 is now generally available on Amazon Bedrock, offering configurable reasoning effort, a 1 million token context window, and tool use for building agents. The model runs on Mantle, Amazo…

15:04
2026-07-16
platform.kimi.ai
artificial-intelligence

Introducing Kimi K3

Kimi released Kimi K3, its most capable model with 2.8 trillion parameters, built on Kimi Delta Attention and Attention Residuals, offering native visual understanding and a 1M-token context window. I…

14:36
2026-07-16
artificialanalysis.ai
artificial-intelligence

Inkling Benchmark Results

Thinking Machines has released Inkling, a 975B-parameter open weights model with 41B active parameters, debuting at 41 on the Artificial Analysis Intelligence Index and becoming the leading open weigh…

02:53
2026-07-15
artificialanalysis.ai
large-language-models

GPT-5.6 Sol, Terra, Luna compare on intelligence vs. cost

GPT-5.6 Sol and Luna outperform Terra at every point on the Intelligence vs Cost per Task chart, with Luna emerging as a particularly cost-efficient model, according to the Artificial Analysis Intelli…

07:54
2026-07-14
techstrong.ai
artificial-intelligence

You Can Keep the Benchmarks. I’ll Take the Test Drive

OpenAI's GPT-5.6 Sol and Anthropic's Claude Fable 5 represent a substantial advance in AI, handling complex, multi-step work with less human guidance, according to a hands-on evaluation by a senior ne…

06:09
2026-07-14
artificialanalysis.ai
artificial-intelligence

Harvey LAB-AA: evaluating AI agents on real-world legal work

Harvey LAB-AA, a new benchmark from Artificial Analysis evaluating AI agents on real-world legal work across 24 practice areas, shows Claude Fable 5 (max, with Opus 4.8 fallback) leading with a 14.2% …

00:00
2026-07-13
tomtunguz.com
artificial-intelligence

The AI Colander

AI models retain between high single digits and 40% of customers after five months, with the stickiest foundational cohorts near the top of that range, according to a study by OpenRouter and a16z. The…

20:01
2026-07-12
pub.towardsai.net
large-language-models

Grok 4.5 Uses 4.2x Fewer Tokens and Costs 17x Less Than Opus 4.8

SpaceXAI and Cursor released Grok 4.5 on July 8, 2026, which uses 4.2 times fewer tokens and costs 17 times less than Claude Opus 4.8 on SWE-Bench Pro, solving tasks for $0.096 versus $1.68. The model…

08:21
2026-07-12
neutralityproject.org
artificial-intelligence

Political Neutrality Benchmark of popular AI models

A new benchmark measuring the political neutrality of 18 AI models from 12 labs across four regions found that 97 out of 108 measured positions landed left of center, with only xAI's Grok models appro…

23:08
2026-07-11
byteiota.com
artificial-intelligence

Grok 4.5 Developer Guide: API, Benchmarks, and When to Use It

XAI released Grok 4.5 on July 8, 2026, claiming top performance on agentic tool use benchmarks and pricing at $2 per million input tokens, roughly 60% cheaper than Claude Opus 4.8. The model leads the…

19:01
2026-07-11
pub.towardsai.net
artificial-intelligence

Grok 4.5 Is xAI's Coding Comeback. The Price Is the Shock.

XAI released Grok 4.5, achieving competitive coding benchmark scores at significantly lower pricing than rivals, positioning it as a cost-effective option for coding-agent routing. The model scored 83…

← prev page 19 / 23 next →
// co-occurs with top 8 entities
// topics top 6 topics