cd /news/large-language-models/running-mager-bench-through-my-chatg… · home › topics › large-language-models › article
[ARTICLE · art-142903] src=mager.co ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Running mager-bench through my ChatGPT subscription

A headless Codex CLI provider added to mager-bench ran GPT-5.6 Sol through all 13 coding challenges for an average score of 9.0/10, with Doom at 8.7, Slots at 8.3, and async-fetch the low point at 6.3. The first pass scored Doom 3.0 and Slots 0.3 because the judge received only the first 6,000 characters of each response; the CLI judge path was fixed to read full responses and those two saved answers were rescored. GPT-6 Astra later averaged 9.3/10 across the same 13 challenges on a new board judged by the same GPT-5.6 Sol CLI model, ahead of Sol's 9.0, while the older Sonnet 5 results moved to an archive.

read1 min views10 publishedSep 25, 2026

My OpenAI API key was out of credits, but the local Codex CLI was already signed in to ChatGPT. I added a headless provider to mager-bench and ran GPT-5.6 Sol through all 13 coding challenges. Each answer and each verdict came from a fresh, read-only codex exec session. The run averaged 9.0/10; Doom scored 8.7, Slots 8.3, and async-fetch was the low point at 6.3. The first pass told a different story: Doom 3.0, Slots 0.3. The judge was receiving only the first 6,000 characters of each response, so it saw partial apps even though both complete HTML files had been saved. I fixed the CLI judge path to read the full response and rescored those two saved answers. The run page includes every response and judge note. Sol graded its own answers, and Codex CLI is an agent harness whose output length is prompted rather than enforced by the API's token cap. I kept this result separate from the original Sonnet 5 board rather than mix judges. Update: I moved new mager-bench runs to my ChatGPT subscription and put Sol and GPT-6 Astra on a new board, both judged by the same GPT-5.6 Sol CLI model. Astra averaged 9.3/10 across all 13 challenges, ahead of Sol's 9.0. The older Sonnet 5 results are now an archive. Self-judging bias still matters, so the full answers and verdicts remain open for inspection.

── more in #large-language-models 4 stories · sorted by recency
── more on @mager-bench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/running-mager-bench-…] indexed:0 read:1min 2026-09-25 · —