Benchmarks show scores. Dashboards show usage. What ships? OpenCode 2.0 introduces Pragmatikos, a new scoring system that ranks AI model pairings by real-world shipping outcomes rather than benchmark scores. Based on 1,210 developer sessions from 2 contributors, the system evaluates planner-builder pairs across six weighted axes including ship rate, cost per ship, and precision, with scores expected to change as more data is contributed. For OpenCode 2.0 Benchmarks test single model. Developers often pair them. A benchmark is an exam: a clean task, a hidden answer key, one model, nobody steering. Real work is the job: a messy repo, your tools, your steering, and increasingly one model planning while another builds. Pragmatikos scores the job, not the exam : every planner → builder pairing, on real sessions, by what actually shipped. Rankings below are pooled from real developer sessions. claude-fable → grok · 62 / 100 the rankings Models and pairings, ranked by what ships Did the work land in a commit, how many nudges did it take, how clean was it, what did it cost. Scored on real sessions, not lab tasks. Scores are expected to change as more data is contributed. 1,210 Cycles