{"slug": "grok-4-6-medium-outperforms-high-effort-on-deepswe", "title": "Grok 4.6 /medium outperforms /high effort on DeepSWE", "summary": "DeepSWE, a new long-horizon software engineering benchmark, reports that Grok 4.6 /medium outperforms /high effort on its leaderboard, which measures frontier coding agents on original tasks across 91 repositories and 5 languages. The benchmark is contamination-free, with tasks written from scratch, and requires 5.5x more code and ~2x more output tokens than SWE-bench Pro, aiming to separate top models that cluster within narrow score bands on existing benchmarks.", "body_md": "Measuring frontier coding agents on original, long-horizon engineering tasks\n\n[Get notified when new models drop](#updates)\n\n## Leaderboard\n\nAll models run on [mini-swe-agent](https://github.com/SWE-agent/mini-swe-agent) for consistency. [Read why](/blog/deepswe#evaluation-harness).\n\nToday's leading public coding benchmarks are starting to saturate at the frontier: top models cluster within a narrow score band where adjacent configurations often overlap on confidence intervals. DeepSWE is a long-horizon software engineering benchmark built to separate them. It delivers four advances over existing public benchmarks:\n\n**Contamination free**: Tasks are written from scratch, not adapted from existing commits or PRs, so no model has seen the solution during pretraining.**High diversity**: Tasks span a broad pool of 91 repositories across 5 languages.** Real-world complexity**: Prompts are ~half the length of SWE-bench Pro's, yet solutions require 5.5x more code and ~2x more output tokens.** Reliable verification**: Verifiers are hand-written to test software behavior rather than implementation details.\n\nThe result is a benchmark that reflects how today's frontier coding agents actually perform in software engineering work.\n\n## Task Examples\n\n[capricorn86/happy-domtypescript](/data/tasks/happy-dom-abort-pending-body-reads)\n\n### Abort pending body reads on shutdown\n\nEnsure interrupted request and response body reads, formData parsing, and discarded timers abort cleanly during shutdown.\n\n[prometheus/prometheusgo](/data/tasks/prometheus-typed-label-sorting)\n\n### Fix PromQL label sorting across typed and untyped values\n\nPromQL label sorting must order mixed typed and untyped label values with stable typed comparison rules.\n\n[c4spar/cliffytypescript](/data/tasks/cliffy-config-file-parsing)\n\n### Add config file parsing to Cliffy commands\n\nAdd command-level config file loading, parsing, merging, and precedence handling.\n\n[yjs/yjsjavascript](/data/tasks/yjs-map-conflict-detection)\n\n### Add deterministic map conflict detection to Y.Map writes\n\nAdd strict, deterministic conflict detection for Y.Map key writes with collect and error policies.\n\n[wasmi-labs/wasmirust](/data/tasks/wasmi-trap-coredumps)\n\n### Add trap coredump generation to wasmi\n\nGenerate opt-in Wasm coredumps on traps and attach the bytes to errors.\n\n[beevik/etreego](/data/tasks/etree-xml-diff-patch)\n\n### Add XML diff, patch, and merge operations to etree\n\nAdd recursive XML diffing, patch generation and application, reverse patching, three-way merge, and diff summaries.\n\n[All 113 tasks](/data/tasks)\n\n## Sign up for leaderboard updates\n\nNew frontier models are added to the DeepSWE leaderboard as they're released.", "url": "https://wpnews.pro/news/grok-4-6-medium-outperforms-high-effort-on-deepswe", "canonical_source": "https://deepswe.datacurve.ai/#grok4.6", "published_at": "2026-08-13 12:56:50+00:00", "updated_at": "2026-08-13 13:13:32.584503+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "ai-tools"], "entities": ["DeepSWE", "Grok 4.6", "mini-swe-agent", "SWE-bench Pro", "capricorn86/happy-dom", "prometheus/prometheus", "c4spar/cliffy", "yjs/yjs"], "alternates": {"html": "https://wpnews.pro/news/grok-4-6-medium-outperforms-high-effort-on-deepswe", "markdown": "https://wpnews.pro/news/grok-4-6-medium-outperforms-high-effort-on-deepswe.md", "text": "https://wpnews.pro/news/grok-4-6-medium-outperforms-high-effort-on-deepswe.txt", "jsonld": "https://wpnews.pro/news/grok-4-6-medium-outperforms-high-effort-on-deepswe.jsonld"}}