Kimi K3
Moonshot AI released Kimi K3, a 2.8-trillion-parameter mixture-of-experts model with 16 active experts per token, achieving the top score of 1679 on the Artificial Analysis webdev arena ahead of Claud…
Moonshot AI released Kimi K3, a 2.8-trillion-parameter mixture-of-experts model with 16 active experts per token, achieving the top score of 1679 on the Artificial Analysis webdev arena ahead of Claud…
Anthropic has repeatedly extended the deadline for removing Claude Fable 5 from paid subscriptions, now set for July 19, in what critics call scarcity marketing. Meanwhile, Moonshot AI's open-weights …
Kimi K3, a reasoning model released on July 16, 2026, by Kimi, scores 57 on the Artificial Analysis Intelligence Index, well above the average of 30 among comparable models, but is slower than average…
XAI's Grok 4.3 is now generally available on Amazon Bedrock, offering configurable reasoning effort, a 1 million token context window, and tool use for building agents. The model runs on Mantle, Amazo…
Kimi K3 ranks third on Artificial Analysis's intelligence index, trailing the leader by only two points. The independent evaluation platform assesses leading AI models across intelligence, cost, speed…
Meta released Muse Spark 1.1 into public preview on July 9, offering a multimodal worker model priced at $1.25 per million input tokens and $4.25 per million output tokens, with 118.1 output tokens pe…
Kimi released Kimi K3, its most capable model with 2.8 trillion parameters, built on Kimi Delta Attention and Attention Residuals, offering native visual understanding and a 1M-token context window. I…
Thinking Machines has released Inkling, a 975B-parameter open weights model with 41B active parameters, debuting at 41 on the Artificial Analysis Intelligence Index and becoming the leading open weigh…
GPT-5.6 Sol and Luna outperform Terra at every point on the Intelligence vs Cost per Task chart, with Luna emerging as a particularly cost-efficient model, according to the Artificial Analysis Intelli…
GPT-5.6 Sol is the better choice over Claude Fable 5, according to a user who prefers Sol by a wide margin. Artificial Analysis gives Fable 5 a 60 on the Intelligence Index, one point ahead of Sol's 5…
OpenAI's GPT-5.6 Sol and Anthropic's Claude Fable 5 represent a substantial advance in AI, handling complex, multi-step work with less human guidance, according to a hands-on evaluation by a senior ne…
Harvey LAB-AA, a new benchmark from Artificial Analysis evaluating AI agents on real-world legal work across 24 practice areas, shows Claude Fable 5 (max, with Opus 4.8 fallback) leading with a 14.2% …
Meta's Muse Spark 1.1 scored 51 on the Artificial Analysis Intelligence Index, an 8-point gain over Muse Spark 1.0 in three months, tying with GLM-5.2, GPT-5.4, and GPT-5.6 Luna. The improvements are …
AI models retain between high single digits and 40% of customers after five months, with the stickiest foundational cohorts near the top of that range, according to a study by OpenRouter and a16z. The…
SpaceXAI and Cursor released Grok 4.5 on July 8, 2026, which uses 4.2 times fewer tokens and costs 17 times less than Claude Opus 4.8 on SWE-Bench Pro, solving tasks for $0.096 versus $1.68. The model…
Chinese developers are paying premium prices for OpenAI's GPT-5.6 via VPNs, despite it being blocked in mainland China, because the model's efficiency reduces total token usage and cost for complex ta…
A new benchmark measuring the political neutrality of 18 AI models from 12 labs across four regions found that 97 out of 108 measured positions landed left of center, with only xAI's Grok models appro…
XAI released Grok 4.5 on July 8, 2026, claiming top performance on agentic tool use benchmarks and pricing at $2 per million input tokens, roughly 60% cheaper than Claude Opus 4.8. The model leads the…
XAI released Grok 4.5, achieving competitive coding benchmark scores at significantly lower pricing than rivals, positioning it as a cost-effective option for coding-agent routing. The model scored 83…
OpenAI released GPT-5.6 with three tiers—Sol, Terra, and Luna—each offering different price-performance levels. The flagship Sol is 54% more token-efficient on AI coding tasks than previous models, an…