We need a harness-only benchmark because model weights aren't
A developer proposes a community-driven benchmark to measure LLM agent harness performance, arguing that orchestration layers often determine real-world success more than model choice. The project wou…