The Plunging Price of Thought
The cost of achieving a given level of AI performance has fallen about 47% per quarter, or 13× per year, since 2023, according to an analysis by researcher David Roodman published with data and code o…
The cost of achieving a given level of AI performance has fallen about 47% per quarter, or 13× per year, since 2023, according to an analysis by researcher David Roodman published with data and code o…
Epoch AI's Benchmarking Hub held about 6,800 results covering roughly 1,200 model versions of 710 models as of 1 October 2026, spanning more than 80 benchmarks from GPQA Diamond and SWE-bench Verified…
The cost of reaching a fixed level of AI benchmark performance has fallen roughly 47% per quarter since 2023, according to Epoch AI, with the price of hitting a 75% score on the GPQA Diamond benchmark…
An Epoch AI report by Emberson and Roodman found that the cost of a given level of AI performance has fallen an average of about 47% per quarter over the past three years, a 13-fold drop every year. T…
Hamel Husain and Shreya Shankar published an AI Evals FAQ distilling the most common questions from teaching 700+ engineers and product managers about AI evaluation. The guide distinguishes model benc…
Researchers introduced QVAC Genesis III, a 191.43B-token STEM-focused synthetic corpus spanning 19 domains, built via a dual generation strategy that uses a weak edge-scale student model's failures as…
Artificial Analysis's Intelligence Index version 4.1.1 weights agentic workloads at 34% and general reasoning at 18%, a shift an analysis of 586 model evaluations argues distorts what the score measur…
A developer analysis of OpenRouter's LLM routing gateway warns that the abstraction between model and provider leaks in production, with the same model ID yielding significantly different performance,…
An operator who has sent more than 18 million messages through OpenRouter measured that identical DeepSeek weights score 90.2% on GPQA Diamond when routed to DeepSeek's own hosting but 75.3% on Digita…
DeepSeek released V4.1 Flash on September 10, 2026, cutting cache-hit costs to $0.003 per token during off-peak hours from the $0.022 charged for the outgoing V4-Pro, a 77-80% price reduction, and rai…
Qwen has released Qwen3.8-Flash-Next, an experimental open-weight causal language model with a vision encoder, designed for coding agents, long-horizon tool use, and multimodal computer tasks. The mod…
OpenAI's GPT-6 Astra, launched on September 3, scored 97.6% on FrontierMath Tier 4 (v2), up from an internal version's 17% and predecessor GPT-5.6 Sol's 83.0%. The model also achieved 99.9% on ARC-AGI…
Researchers propose Gradient-Aligned Reward (GAR), a dense reward method for reinforcement learning from verifiable rewards (RLVR) that uses cosine similarity between rollout and expert-anchor gradien…
A developer's benchmark of Qwen3.8 27B quantizations found that the 17 GB Q4_K_M 4-bit version matches the full BF16 model on Terminal-Bench 2.1, while 1-bit versions collapse to near-random performan…
Researchers at the University of Edinburgh and NVIDIA developed Dynamic Memory Sparsification (DMS), a technique that compresses the key-value cache of large language models by 8x while improving perf…
VIDRAFT, a Korean AI startup, has released Darwin Family, a training-free model merging framework that uses evolutionary algorithms to combine the parameters of two existing models. Their flagship mod…
Google DeepMind released DiffusionGemma, an experimental open-weight model that converts the Gemma 4 architecture into a discrete text-diffusion system using less than 10% of Gemma 4's original traini…
Hyperbolic Sparse Autoencoders (HyperSAE) outperform standard Euclidean Sparse Autoencoders (FlatSAEs) in reconstructing Google Gemma-2-2B activations, reducing reconstruction mean squared error (MSE)…
Anthropic's Claude Sonnet 5 and Claude Opus 5, released in mid-2026, offer distinct strengths and pricing trade-offs. Sonnet 5 excels at high-volume coding and content tasks with a 72.7% SWE-bench Ver…
AGI Ranker, an independent AI leaderboard, released v2.0.0 of its AGI Score, causing every model's score to drop by 6 to 15 points due to a correction in how human-parity ceilings are applied. The rev…