cd /news/ai-agents/slopcodebench · home topics ai-agents article
[ARTICLE · art-76205] src=scbench.ai ↗ pub= topic=ai-agents verified=true sentiment=· neutral

SlopCodeBench

SlopCodeBench, a community benchmark measuring code erosion as agents iteratively extend their own solutions across checkpoints, has released its top 10 models. The leading model is 01GPT 5.5/Codex with a score of 28.1% and a code erosion metric of 0.494, followed by GPT 5.3-Codex/Codex at 26.0% and GPT 5.4/Codex at 25.5%. The benchmark evaluates coding agents through repeated requirement changes and extensions, tracking both correctness and code quality.

read1 min views1 publishedJul 28, 2026

A community benchmark measuring code erosion as agents iteratively extend their own solutions across checkpoints.

top 10 models

full leaderboard →

  • 01GPT 5.5/Codex28.1%0.4940.269
  • 02GPT 5.3-Codex/Codex26.0%0.6440.336
  • 03GPT 5.4/Codex25.5%0.2780.193
  • 04GPT 5.2-Codex/Codex21.9%0.7280.398
  • 05Opus 4.6/Claude Code20.9%0.7370.318
  • 06Opus 4.7/Claude Code20.9%0.7590.357
  • 07KIKimi K2.6/Kimi CLI18.9%0.7640.399
  • 08Opus 4.5/Claude Code17.3%0.6910.297
  • 09Sonnet 4.6/Claude Code16.8%0.7410.316
  • 10Composer 2/Cursor CLI16.3%0.7160.353

browse all →

Overview #

SlopCodeBench evaluates coding agents the way real software actually gets built: through repeated requirement changes and extensions. Each problem is a sequence of checkpoints — the agent implements an initial version, then extends its own solution as new requirements arrive. Evaluation is black-box: only a CLI or API contract is given, with no prescribed architecture, function signatures, or module boundaries, so early design decisions compound across the run. Beyond correctness, we measure code erosion — verbosity, dead branches, and redundant structure — to surface the agents that stay clean under sustained change instead of patching their way into slop.

Supported by

Special thanks to Snorkel AI for supporting this work through the Open Benchmarks Grant.

── more in #ai-agents 4 stories · sorted by recency
── more on @slopcodebench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/slopcodebench] indexed:0 read:1min 2026-07-28 ·