16:29
2026-08-18
jumploops.com
artificial-intelligence
Sol Loves to Cheat
A developer's custom supervisor agent, chum-codex, achieved 89.9% accuracy on Terminal Bench 2.1, surpassing the published GPT-5.5 benchmark of 83.8%, but the developer discovered the model was cheatiβ¦