cd/entity/LiveCodeBench· home entities LiveCodeBench
grep -l @livecodebench /news/*.json | wc -l → 11

LiveCodeBench

mentions 11 type Organization feed RSS

// recent coverage 11 mentions

07:35
2026-08-04
leaddev.com
artificial-intelligence

Your AI-coding agents might need an org chart

A controlled experiment by researchers testing Claude Opus 4.7 and Codex GPT-5.5 on 116 Python tasks found that pairing a weaker reviewer with a stronger writer can degrade performance: Claude's 91.4%…

04:00
2026-08-04
machinebrief.com
large-language-models

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

A new arXiv preprint (2608.00355v1) finds that apparent progress toward harder tasks in large language models is mostly a ceiling effect, but a smaller hard-task effect persists after controlling for …

15:29
2026-07-09
lesswrong.com
ai-safety

Debate with Self-Play Best-of-N Optimization

Researchers at an undisclosed lab introduced a best-of-N (BoN) optimization method as a proxy for self-play training in debate protocols, aiming to improve scalable oversight for AI systems. Their exp…

15:17
2026-06-30
byteiota.com
large-language-models

Gemini 2.5 Pro Deep Think: What the Benchmarks Mean

Google's Gemini 2.5 Pro with Deep Think reasoning mode topped coding and reasoning benchmarks this week, scoring 82.4% on GPQA Diamond and 94.1% on HumanEval+, but the mode multiplies token costs by r…

// co-occurs with top 8 entities
// topics top 6 topics