01:41
2026-08-18
shukla.io
ai-research
Who benchmarks the benchmark?
A new audit of the EnterpriseOps Gym benchmark found that fixing environment issues in the 'Teams' domain raised GPT-5.6 Luna's score from 26.2% to 100% on 61 tasks, revealing that many agent failuresβ¦