Models Are Getting Dumber on Purpose
Reasoning models are deliberately trading world knowledge for reasoning skill, with GLM-5.2 scoring 99.2% on AIME 2026 using about 40 billion active parameters per token, while factual recall remains …
Reasoning models are deliberately trading world knowledge for reasoning skill, with GLM-5.2 scoring 99.2% on AIME 2026 using about 40 billion active parameters per token, while factual recall remains …
SpaceXAI's Grok 4.6, released August 12, 2026, scores 61 on the Artificial Analysis Intelligence Index, tying OpenAI's GPT-5.6 Sol and beating Grok 4.5 by five points, while costing $2 per million inp…
Artificial Analysis has launched Optima, a platform that lets users build custom AI benchmarks from their own data and workflows, comparing models on quality, cost, and time per task. The tool address…
XAI released Grok 4.6 on August 12, 35 days after Grok 4.5, priced at $2 per million input tokens and $6 per million output tokens, making it 2.5x to 5x cheaper than GPT-5.6 Sol and Claude Opus 5 at c…
AI inference costs are becoming a major problem for many companies, with AI spend per employee per month rising sharply, according to the Ramp AI Index via a16z. To reduce costs, companies should audi…
Elon Musk's SpaceXAI (formerly xAI) released Grok 4.6, its latest flagship AI model for enterprise workflows, priced at $2 per million input tokens and $6 per million output tokens, less than half of …
ByteDance's Seedance 2.5 and MiniMax's H3, released on the same day two weeks ago, have propelled China to dominate the AI video race, with nine of the top 10 text-to-video systems on Artificial Analy…
Google DeepMind released Gemini 3.7 Flash, which scores 56 on the Artificial Analysis Intelligence Index with high reasoning, a 4-point improvement over Gemini 3.6 Flash, and achieves an average Time …
Qwen3.8 2.4T A95B, released on August 12, 2026, scores 58 on the Artificial Analysis Intelligence Index, well above the median of 27, but is priced at $2.00 per 1M input tokens and $6.00 per 1M output…
Google's Gemini 3.7 Flash, launched August 13, 2026 at $0.75 per million input tokens and $3.75 per million output, wins 9 of 19 benchmark rows in its own model card, while GPT-5.6 Terra wins 6 and Cl…
Z.ai released GLM-5.3 on August 14, 2026, claiming all benchmark gains, including a Terminal-Bench 3.0 score jump from 4.6 to 28.3, come from post-training alone on the identical GLM-5.2 base model. T…
State-of-the-art AI models are two-thirds smarter than last November, with two new models released every three days, yet 84% of tokens on OpenRouter are not state of the art, and the six most-used mod…
Investors in Anthropic, the developer of the Claude AI models, are targeting a $2 trillion valuation for its October IPO, according to the Financial Times, citing the company's projected revenue of $1…
Google's Gemini 3.7 Flash now sits on the Pareto frontier for intelligence versus speed, according to benchmarking firm Artificial Analysis, scoring 56 on the Artificial Analysis Intelligence Index wi…
Cerebras and OpenAI launched Ultrafast Mode, a new service tier in the OpenAI API powered by Cerebras that delivers GPT-5.6 Sol at up to 750 output tokens per second with no quality compromise. In Cer…
MiniMax-H3, a 33-billion-parameter open-weight video generation model from Chinese AI firm MiniMax, ranks first in the Video Edit Arena on arena.ai with 1390 points, 32 points ahead of Dreamina Seedan…
Ant Group's inclusionAI released Ling 3.0 Flash, an open-weights AI model that scores 38 points on the Artificial Analysis Intelligence Index, making it the smartest open model under 124 billion total…
Google announced that its Gemini app has surpassed 1 billion monthly users, with 63% of users interacting via voice and over 150 million images generated daily. Despite the milestone, independent test…
Google DeepMind CEO and cofounder Demis Hassabis is stepping back from the CEO role to become chair, with CTO Koray Kavukcuoglu taking over day-to-day control, while chief scientist Jeff Dean is also …
DeepSeek's V4 Pro 0813 (max) model scored 53 on the Artificial Analysis Intelligence Index, one point ahead of DeepSeek V4 Flash 0731, according to an independent evaluation by Artificial Analysis. Th…