00:00
2026-09-25
mager.co
large-language-models
Running mager-bench through my ChatGPT subscription
A headless Codex CLI provider added to mager-bench ran GPT-5.6 Sol through all 13 coding challenges for an average score of 9.0/10, with Doom at 8.7, Slots at 8.3, and async-fetch the low point at 6.3…