{"slug": "benchmarks-show-scores-dashboards-show-usage-what-ships", "title": "Benchmarks show scores. Dashboards show usage. What ships?", "summary": "OpenCode 2.0 introduces Pragmatikos, a new scoring system that ranks AI model pairings by real-world shipping outcomes rather than benchmark scores. Based on 1,210 developer sessions from 2 contributors, the system evaluates planner-builder pairs across six weighted axes including ship rate, cost per ship, and precision, with scores expected to change as more data is contributed.", "body_md": "For OpenCode 2.0\n\n# Benchmarks test single model. Developers often pair them.\n\nA benchmark is an exam: a clean task, a hidden answer key, one model, nobody steering. Real work is the job: a messy repo, your tools, your steering, and increasingly one model planning while another builds. **Pragmatikos scores the job, not the exam**: every planner → builder pairing, on real sessions, by what actually shipped.\n\nRankings below are pooled from real developer sessions.\n\nclaude-fable → grok · 62 / 100\n\nthe rankings\n\n## Models and pairings, ranked by what ships\n\nDid the work land in a commit, how many nudges did it take, how clean was it, what did it cost. Scored on real sessions, not lab tasks. **Scores are expected to change as more data is contributed.**\n\n1,210 Cycles<sup>*</sup> · 2 contributors\n\njoin in\n\nThis is a first answer, not the final one. It's built from the sessions developers have shared so far. The more histories in the pool, the harder the ranking gets to argue with.\n\nwhy it exists\n\n## Three questions. Three instruments.\n\nThe first two are useful and stay useful. Developers choosing a model and a workflow are missing the third answer.\n\nquestion 01 · the lab\n\n### Benchmarks ask: how capable is this model?\n\nSWE-bench, Terminal-bench, Aider, LMArena. Controlled tasks, hidden tests, preference votes — the best way to compare models on equal footing and know each model’s ceiling.\n\n- signal\n- pass rate · votes\n- blind spot\n- One model, a clean task, nobody steering.\n\nquestion 02 · the crowd\n\n### Usage rankings ask: what is the market using?\n\nOpenCode Data, OpenRouter Ranking. Tokens, users, retention, dollars per session — the best way to see adoption and where the mix is shifting.\n\n- signal\n- tokens · users · $/session\n- blind spot\n- Popular is not productive — and still one model at a time.\n\nquestion 03 · the recorder\n\n### Pragmatikos asks: what ships in real work?\n\nTurns, phases, edits and errors from real sessions, correlated with what shipped. The one instrument that scores the setups developers actually run.\n\n- signal\n- session record × shipped outcomes\n- blind spot\n- Anything it never saw. Observational, not a benchmark.\n\nCapable. Popular. Effective.\n\nOnly one of them has a row for the setup you actually run.\n\nthe blind spot both miss\n\n## Same builder. Different planner. Opposite results.\n\nThe build model is identical. One planner got it to a commit in a few turns; the other burned dozens and shipped nothing. A model ranking can't see this. A pairing ranking can.\n\nclaude-fable-5.1 → grok-4.6\n\n0 turns\n\ngrok-4.6 → grok-4.6\n\n0 turns\n\nReal sessions. Unshipped work counts against the model.\n\nhow the score works\n\n## Six tiers, ten axes, one weighted mean\n\nOutcome\n\nweight 30\n\nShip rate and one-shot rate (prompt plus approval).\n\nCost of a ship\n\nweight 20\n\nTurns, hours and dollars per shipped cycle.\n\nPrecision\n\nweight 15\n\nTool errors and human interventions.\n\nDiscipline\n\nweight 15\n\nShare of cycles verified after the last edit.\n\nEfficiency\n\nweight 10\n\nEdits per human turn.\n\nLatency\n\nweight 10\n\nMedian seconds per assistant step.\n\nPool-relative\n\nEvery axis is a log distance to the pooled average — log odds for rates — on a fixed ×4 span, so ×2 and ÷2 sit the same distance from the line.\n\nEvidence-weighted\n\nSmall groups shrink toward the pool with k = 10 pseudo-cycles. A two-cycle fluke nearly vanishes.\n\nFailure lowers the rate\n\nUnshipped cycles stay in the ship-rate denominator, so few-turn dead ends can never look good. Turns, hours and dollars are priced per shipped cycle.\n\ntwo records, correlated\n\n## The session record comes first\n\nEvery turn, plan and build phase, edit, tool error and dollar is read from the agent's own record — that is where every axis is measured. Commits only decide which cycles count as shipped: a window from the previous commit, file overlap decides credit, active-but-absent sessions advised only.\n\nwhere it falls short\n\n## A hypothesis to try, not a verdict\n\nPragmatikos reads sessions after the fact. Nobody assigned pairings to developers or tasks, so the ranking says which setups shipped for the people who ran them — not which setup would ship for you. Here is where that bites, and what we intend to do about it.\n\nObservational, not controlled\n\nThe developers who run one pairing are not the developers who run another, and they are not working the same tasks. Skill, repo and task difficulty can move a score as much as the models do. “Ranked by what ships” is a correlation. Read it as one.\n\nSmall, and uneven\n\n1,210 cycles from 2 contributors, across 7 family pairings — 72% of them from one pairing. Groups near the ten-cycle floor are shrunk toward the pool, which guards against flukes but does not stand in for evidence. Two scores a point or two apart are a tie.\n\nSelf-selected\n\nEveryone in the pool installed a plugin and left sharing on. The ranking says nothing about developers who did not.\n\nShipped is a floor, not a grade\n\nA commit means the work landed. It does not mean it survived review, was never reverted, or was any good.\n\nMore histories\n\nEvery new contributor puts the same pairings in front of different developers, repos and tasks. That is what lets a pairing’s effect separate from the developer’s. The scale and the shrinkage were built for a bigger pool; the pool is what is missing.\n\n[Share your sessions →](#contribute)\n\nIf that is not enough\n\nThe pool has one thing going for it: everyone in it is a developer running OpenCode on real work — the population this ranking is for. That does not fix the task and skill mix, but a different approach can start from it. If more histories do not settle the picture, the method changes, and this page will say what changed.\n\ncontribute\n\n## Add your sessions in two minutes\n\nFor OpenCode v2 — other coding agents coming soon.\n\nInstall the ocInsights plugin for OpenCode: add \"@pfoundation/ocinsights\" to the \"plugin\" list in my global opencode.json config. Then tell me to restart OpenCode, and afterwards verify the plugin loaded and the insights deck answers at http://127.0.0.1:4173/. Also report whether insight contribution is on, without changing that setting.\n\n1. 1Paste the prompt into your OpenCode agent and let it edit your config.\n2. 2Restart OpenCode when it tells you to, then open the deck it verifies.\n3. 3Preview with Contribute in the deck header — sharing is on by default.\n\nTwenty-five fields per cycle — day, models, effort, harness, turns, edits, cost, shipping. No paths, prompts, session ids or projects. Preview in the deck's Contribute panel before anything leaves your machine.", "url": "https://wpnews.pro/news/benchmarks-show-scores-dashboards-show-usage-what-ships", "canonical_source": "https://pragmatikos.ai/", "published_at": "2026-09-07 16:16:46+00:00", "updated_at": "2026-09-07 16:26:42.745029+00:00", "lang": "en", "topics": ["ai-tools", "ai-research", "developer-tools"], "entities": ["OpenCode", "Pragmatikos", "claude-fable", "grok"], "alternates": {"html": "https://wpnews.pro/news/benchmarks-show-scores-dashboards-show-usage-what-ships", "markdown": "https://wpnews.pro/news/benchmarks-show-scores-dashboards-show-usage-what-ships.md", "text": "https://wpnews.pro/news/benchmarks-show-scores-dashboards-show-usage-what-ships.txt", "jsonld": "https://wpnews.pro/news/benchmarks-show-scores-dashboards-show-usage-what-ships.jsonld"}}