21:14
2026-07-13
glassmkr.com
ai-agents
We gave open models root on broken servers and graded them on the machine
A new agent evaluation method that grades models against real server states rather than their own reports found that model size does not determine success, with a 27B reasoning-tuned Qwen derivative a…