cd/entity/SlopCodeBench· home entities SlopCodeBench
grep -l @slopcodebench /news/*.json | wc -l → 7

SlopCodeBench

mentions 7 type Organization feed RSS

// recent coverage 7 mentions

00:00
2026-09-10
coderabbit.ai
ai-agents

A software factory needs a review gate it can trust

A 2026 National Bureau of Economic Research working paper, "Writing Code vs. Shipping Code," studied more than 100,000 GitHub developers and found large gains in coding activity across successive gene…

14:10
2026-08-17
blog.sshh.io
artificial-intelligence

How I use AI in 2026 (Coding, Writing, Learning, Assistant-ing)

In a personal update to his 2025 post, sshh.io blogger details his 2026 AI workflow, reporting that 95%+ of code in his side projects is written in a single mega build run using vanilla Codex and Clau…

15:08
2026-08-04
github.com
artificial-intelligence

Benchmarking Fable, Sol, and Kimi K3 on SlopCodeBench

In a benchmark run on Thursday, Fable and Sol tied at 33.3% (10/30 strict checkpoint passes) on SlopCodeBench, with Expo and Kimi K3 following at 26.7% and 23.3%, respectively. The evaluation, conduct…

01:08
2026-07-28
sourcefeed.dev
artificial-intelligence

Opus 5 Quadruples a SlopCodeBench Subset — by Tripling the Code

Claude Opus 5 passed 24% of 17 checkpoints on a subset of SlopCodeBench, quadrupling the 6% pass rates of Opus 4.8 and Sonnet 5, but achieved this by producing roughly 29,000 source lines versus about…

01:06
2026-07-28
scbench.ai
ai-agents

SlopCodeBench

SlopCodeBench, a community benchmark measuring code erosion as agents iteratively extend their own solutions across checkpoints, has released its top 10 models. The leading model is 01GPT 5.5/Codex wi…

22:37
2026-07-27
github.com
artificial-intelligence

Benchmarking Opus 5 on SlopCodeBench

Opus 5 achieved a 24% strict pass rate on a subset of the SlopCodeBench coding benchmark from UW Madison, only marginally higher than Opus 4.6's 17% in the original paper, and wrote five times more fu…

00:00
2026-06-18
jasonrobert.dev
artificial-intelligence

Skill Rot Is Real

Developers are reporting 'skill rot' as AI coding assistants automate debugging and problem-solving, reducing hands-on practice. A METR study found developers believed AI made them 20% faster but were…

// co-occurs with top 8 entities
// topics top 6 topics