Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x FrontierHarness Eval, a new benchmark from Runta, tested nine AI coding harnesses on the same model and found median cost per successful task varies by 17x, with Claude Code v2.1.237 passing 19 tasks but reaching $18.34 per task, while OpenCode v1.18.19's cost rises to $3.24 per task when failed attempts are included. The evaluation, run on Runta agent runtimes with identical golden checkpoints, highlights that cache hit rate is not cost and quality and cost can diverge. Median cost per successful task Median cost per task Median cache hit rate per successful task Median time per successful task Beyond the numbers - 01 OpenCode: failures excluded. It only covers 15 passes. Count failed attempts and the number becomes $3.24 per task. - 02 Cache hit rate is not cost. A cached 300-turn failure can still burn more than a short cache miss. - 03 Quality and cost can diverge. Claude Code passes 19 tasks, but reaches $18.34 in cost per task. Run your harness on Runta. If you want to test your own harness on Runta, we’ll give you $100 in credits to get started. Tested harnesses Codex v0.148.0 DeepSeek Harness v0.1.0-rc.8 Claude Code v2.1.237 Pi v0.84.2 Oh My Pi v17.4.0 Kimi Code v0.37.2 Exo Harness v0.1.0 OpenCode v1.18.19 Hermes v0.20.4 - FrontierHarness v1.0 focuses on software engineering contexts and terminal-based tasks. It may not generalize to other areas of knowledge work. - Evaluated on Runta agent runtimes. All harnesses and the task environment are prepared once as a golden checkpoint. Every run is a fresh restore with identical vCPU, memory, disk size, disk contents, and memory state.