cd/entity/GPQA· home entities GPQA
grep -l @gpqa /news/*.json | wc -l → 10

GPQA

mentions 10 type Organization feed RSS

// recent coverage 10 mentions

21:39
2026-09-01
huggingface.co
artificial-intelligence

BenchMIRT: What are LLM benchmarks actually measuring?

The Allen Institute for AI introduced BenchMIRT, a new method for auditing LLM benchmarks at the level of individual prompts, which uses multidimensional Item Response Theory to separate the underlyin…

17:21
2026-08-12
dev.to
machine-learning

Benchmarks for Scientific Reasoning: What a Score Establishes

Multigrid AI explains how graduate-level science benchmarks like GPQA are constructed, emphasizing the use of a non-expert baseline to make scores interpretable. The process involves domain experts wr…

20:44
2026-07-30
lesswrong.com
ai-research

Hint-based CoT faithfulness evals still mostly work on Claude

Redwood Research finds that hint-based chain-of-thought faithfulness evaluations still work on Claude models, contradicting Anthropic system card claims that recent models no longer use hints. The rep…

21:47
2026-07-25
promptcube3.com
artificial-intelligence

Model Benchmarks: The New Arms Race

Model benchmarks are creating a fragmented AI landscape where 'the best model' depends on which specific test is valued most, according to a news analysis. GPT-5.6 leads on GPQA while Opus 5 dominates…

04:00
2026-07-14
arxiv.org
large-language-models

Depth-Entropy Guided Sampling for Training-Free LLM Reasoning

Researchers introduce Depth-Entropy Guided Sampling (DEGS), a training-free test-time method that exploits layer-wise entropy collapse in transformer forward passes to improve LLM reasoning. Across th…

15:03
2026-06-17
dev.to
large-language-models

Claude 3.5 Sonnet Isn't Just an Upgrade. It's a New Baseline.

Anthropic released Claude 3.5 Sonnet, a new AI model that outperforms the previous top-tier Claude 3 Opus in intelligence, speed, and cost. The model achieves a 64% solve rate on internal agentic codi…

11:24
2026-06-04
huggingface.co
artificial-intelligence

Task-Seeded Synthetic Q&A Generation for Nemotron Pretraining

NVIDIA researchers developed a task-seeded synthetic Q&A generation workflow for Nemotron-family pretraining that uses public task training splits as capability seeds to generate new task-aligned exam…

// co-occurs with top 8 entities
// topics top 6 topics