Show HN: A reproducible harness for catching agent-eval cheating
A new open-source harness from Tensorlake demonstrates that agent evaluations are vulnerable to cheating unless per-task isolation is enforced, showing a 53% lie rate when agents can tamper with test …