cd /news/artificial-intelligence/grok-4-6-medium-outperforms-high-eff… · home topics artificial-intelligence article
[ARTICLE · art-95264] src=deepswe.datacurve.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Grok 4.6 /medium outperforms /high effort on DeepSWE

DeepSWE, a new long-horizon software engineering benchmark, reports that Grok 4.6 /medium outperforms /high effort on its leaderboard, which measures frontier coding agents on original tasks across 91 repositories and 5 languages. The benchmark is contamination-free, with tasks written from scratch, and requires 5.5x more code and ~2x more output tokens than SWE-bench Pro, aiming to separate top models that cluster within narrow score bands on existing benchmarks.

read2 min views1 publishedAug 13, 2026
Grok 4.6 /medium outperforms /high effort on DeepSWE
Image: source

Measuring frontier coding agents on original, long-horizon engineering tasks

Get notified when new models drop

Leaderboard #

All models run on mini-swe-agent for consistency. Read why. Today's leading public coding benchmarks are starting to saturate at the frontier: top models cluster within a narrow score band where adjacent configurations often overlap on confidence intervals. DeepSWE is a long-horizon software engineering benchmark built to separate them. It delivers four advances over existing public benchmarks:

Contamination free: Tasks are written from scratch, not adapted from existing commits or PRs, so no model has seen the solution during pretraining.High diversity: Tasks span a broad pool of 91 repositories across 5 languages.** Real-world complexity**: Prompts are ~half the length of SWE-bench Pro's, yet solutions require 5.5x more code and ~2x more output tokens.** Reliable verification**: Verifiers are hand-written to test software behavior rather than implementation details.

The result is a benchmark that reflects how today's frontier coding agents actually perform in software engineering work.

Task Examples #

capricorn86/happy-domtypescript

Abort pending body reads on shutdown

Ensure interrupted request and response body reads, formData parsing, and discarded timers abort cleanly during shutdown.

prometheus/prometheusgo

Fix PromQL label sorting across typed and untyped values

PromQL label sorting must order mixed typed and untyped label values with stable typed comparison rules.

c4spar/cliffytypescript

Add config file parsing to Cliffy commands

Add command-level config file , parsing, merging, and precedence handling.

yjs/yjsjavascript

Add deterministic map conflict detection to Y.Map writes

Add strict, deterministic conflict detection for Y.Map key writes with collect and error policies.

wasmi-labs/wasmirust

Add trap coredump generation to wasmi

Generate opt-in Wasm coredumps on traps and attach the bytes to errors.

beevik/etreego

Add XML diff, patch, and merge operations to etree

Add recursive XML diffing, patch generation and application, reverse patching, three-way merge, and diff summaries.

All 113 tasks

Sign up for leaderboard updates #

New frontier models are added to the DeepSWE leaderboard as they're released.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @deepswe 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/grok-4-6-medium-outp…] indexed:0 read:2min 2026-08-13 ·