{"slug": "slopcodebench", "title": "SlopCodeBench", "summary": "SlopCodeBench, a community benchmark measuring code erosion as agents iteratively extend their own solutions across checkpoints, has released its top 10 models. The leading model is 01GPT 5.5/Codex with a score of 28.1% and a code erosion metric of 0.494, followed by GPT 5.3-Codex/Codex at 26.0% and GPT 5.4/Codex at 25.5%. The benchmark evaluates coding agents through repeated requirement changes and extensions, tracking both correctness and code quality.", "body_md": "# SlopCodeBench\n\nA community benchmark measuring code erosion as agents iteratively extend their own solutions across checkpoints.\n\n### top 10 models\n\n[full leaderboard →](/leaderboard)\n\n- 01GPT 5.5/Codex28.1%0.4940.269\n- 02GPT 5.3-Codex/Codex26.0%0.6440.336\n- 03GPT 5.4/Codex25.5%0.2780.193\n- 04GPT 5.2-Codex/Codex21.9%0.7280.398\n- 05Opus 4.6/Claude Code20.9%0.7370.318\n- 06Opus 4.7/Claude Code20.9%0.7590.357\n- 07KIKimi K2.6/Kimi CLI18.9%0.7640.399\n- 08Opus 4.5/Claude Code17.3%0.6910.297\n- 09Sonnet 4.6/Claude Code16.8%0.7410.316\n- 10Composer 2/Cursor CLI16.3%0.7160.353\n\n## Featured Problems\n\n[browse all →](/problems)\n\n## Overview\n\nSlopCodeBench evaluates coding agents the way real software actually gets built: through repeated requirement changes and extensions. Each problem is a sequence of checkpoints — the agent implements an initial version, then extends its own solution as new requirements arrive. Evaluation is black-box: only a CLI or API contract is given, with no prescribed architecture, function signatures, or module boundaries, so early design decisions compound across the run. Beyond correctness, we measure code erosion — verbosity, dead branches, and redundant structure — to surface the agents that stay clean under sustained change instead of patching their way into slop.\n\nSupported by\n\nSpecial thanks to [Snorkel AI](https://snorkel.ai) for supporting this work through the Open Benchmarks Grant.", "url": "https://wpnews.pro/news/slopcodebench", "canonical_source": "https://www.scbench.ai", "published_at": "2026-07-28 01:06:13+00:00", "updated_at": "2026-07-28 01:22:20.859614+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "ai-research", "developer-tools", "machine-learning"], "entities": ["SlopCodeBench", "Snorkel AI", "01GPT 5.5/Codex", "GPT 5.3-Codex/Codex", "GPT 5.4/Codex", "GPT 5.2-Codex/Codex", "Opus 4.6/Claude Code", "Opus 4.7/Claude Code"], "alternates": {"html": "https://wpnews.pro/news/slopcodebench", "markdown": "https://wpnews.pro/news/slopcodebench.md", "text": "https://wpnews.pro/news/slopcodebench.txt", "jsonld": "https://wpnews.pro/news/slopcodebench.jsonld"}}