cd/entity/LiveCodeBench· home› entities› LiveCodeBench
grep -l @livecodebench /news/*.json | wc -l → 21

LiveCodeBench

mentions 21 type Organization page 1/2 feed RSS

// recent coverage 21 mentions

13:42
2026-09-18
blog.jetbrains.com
ai-tools

Making Local AI Smarter and Faster

JetBrains released Qwen3.8-3.6-27B-blend, a merged 27B local coding model that completed 37 of 100 internal coding tasks versus 34 for Qwen3.6 with reasoning disabled and 39 for Qwen3.8, while generat…

00:00
2026-08-25
mindstudio.ai
artificial-intelligence

Escha-W2: 2-Bit Quantization That Shrinks a 27B Model to 10GB

Escha Labs Inc. released Escha-W2, a 2-bit quantized build of Qwen3.8-27B that compresses the 27-billion-parameter model to 10.15GB, enabling 128k context on a single 24GB GPU while matching FP8 quali…

07:35
2026-08-04
leaddev.com
artificial-intelligence

Your AI-coding agents might need an org chart

A controlled experiment by researchers testing Claude Opus 4.7 and Codex GPT-5.5 on 116 Python tasks found that pairing a weaker reviewer with a stronger writer can degrade performance: Claude's 91.4%…

04:00
2026-08-04
machinebrief.com
large-language-models

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

A new arXiv preprint (2608.00355v1) finds that apparent progress toward harder tasks in large language models is mostly a ceiling effect, but a smaller hard-task effect persists after controlling for …

15:29
2026-07-09
lesswrong.com
ai-safety

Debate with Self-Play Best-of-N Optimization

Researchers at an undisclosed lab introduced a best-of-N (BoN) optimization method as a proxy for self-play training in debate protocols, aiming to improve scalable oversight for AI systems. Their exp…

15:17
2026-06-30
byteiota.com
large-language-models

Gemini 2.5 Pro Deep Think: What the Benchmarks Mean

Google's Gemini 2.5 Pro with Deep Think reasoning mode topped coding and reasoning benchmarks this week, scoring 82.4% on GPQA Diamond and 94.1% on HumanEval+, but the mode multiplies token costs by r…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics