cd/entity/GPQA Diamond· home› entities› GPQA Diamond
grep -l @gpqa diamond /news/*.json | wc -l → 26

GPQA Diamond

mentions 26 type Person page 1/2 feed RSS

// recent coverage 26 mentions

23:16
2026-10-06
epoch.ai
artificial-intelligence

The Plunging Price of Thought

The cost of achieving a given level of AI performance has fallen about 47% per quarter, or 13× per year, since 2023, according to an analysis by researcher David Roodman published with data and code o…

11:36
2026-10-02
mostlyright.md
large-language-models

A table of 1,200 model benchmarks since 2019, updated daily

Epoch AI's Benchmarking Hub held about 6,800 results covering roughly 1,200 model versions of 710 models as of 1 October 2026, spanning more than 80 benchmarks from GPQA Diamond and SWE-bench Verified…

21:03
2026-10-01
firethering.com
artificial-intelligence

AI Is Getting Cheaper, But AI Bills Could Still Go Up

The cost of reaching a fixed level of AI benchmark performance has fallen roughly 47% per quarter since 2023, according to Epoch AI, with the price of hitting a 75% score on the GPQA Diamond benchmark…

11:17
2026-09-23
marginalrevolution.com
artificial-intelligence

The Price of Intelligence is Falling Rapidly

An Epoch AI report by Emberson and Roodman found that the cost of a given level of AI performance has fallen an average of about 47% per quarter over the past three years, a 13-fold drop every year. T…

07:00
2026-09-18
hamel.dev
ai-research

AI Evals: Everything You Need to Know

Hamel Husain and Shreya Shankar published an AI Evals FAQ distilling the most common questions from teaching 700+ engineers and product managers about AI evaluation. The guide distinguishes model benc…

18:23
2026-09-16
itsmonkey.business
artificial-intelligence

Artificial Analysis: What Is the Intelligence Index Measuring?

Artificial Analysis's Intelligence Index version 4.1.1 weights agentic workloads at 34% and general reasoning at 18%, a shift an analysis of 586 model evaluations argues distorts what the score measur…

09:23
2026-09-10
forkast.news
large-language-models

DeepSeek’s New Architecture Slashes Agentic Costs by 80%

DeepSeek released V4.1 Flash on September 10, 2026, cutting cache-hit costs to $0.003 per token during off-peak hours from the $0.022 charged for the outgoing V4-Pro, a 77-80% price reduction, and rai…

20:52
2026-08-25
cryptobriefing.com
artificial-intelligence

Researchers shrink AI model while enhancing its intelligence

Researchers at the University of Edinburgh and NVIDIA developed Dynamic Memory Sparsification (DMS), a technique that compresses the key-value cache of large language models by 8x while improving perf…

22:07
2026-07-31
dev.to
large-language-models

Claude Sonnet 5 vs Opus 5: A Real-World Comparison (2026)

Anthropic's Claude Sonnet 5 and Claude Opus 5, released in mid-2026, offer distinct strengths and pricing trade-offs. Sonnet 5 excels at high-volume coding and content tasks with a 72.7% SWE-bench Ver…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics