{"slug": "show-hn-evals-coach-a-claude-plugin-to-help-pms-write-good-evals", "title": "Show HN: Evals Coach – a Claude plugin to help PMs write good evals", "summary": "Evals Coach, a Claude plugin by justshipai, guides product managers through designing AI evals step-by-step, producing importable test cases, judge prompts, and calibration guides. In a six-task blinded comparison using GPT-5.6 Sol at medium reasoning, the eval-design method behind Evals Coach outperformed a no-method baseline on rubric score and won tasks, with fewer critical failures.", "body_md": "Evals Coach is a Claude plugin that guides you through designing evals,\n\none decision at a time.\n\nStep 3, after reviewing twelve real outputs. Your notes become named failure modes, each one traceable back to the outputs it came from, and ranked by how often it actually happened.\n\nYou bring product judgement. It brings the eval-design method: criteria, cases, graders and thresholds.\n\nStart from a PRD, a feature idea, or a sentence describing what you’re building. No codebase required.\n\nThe output isn’t tied to any eval platform, so it imports into whatever your team already runs.\n\nDescribe your feature; Evals Coach drafts each step, you edit it, and the outputs assemble as you go.\n\nA sentence or two on what it does and the worst thing it could get wrong. If it's already live, paste a few real outputs too, and every later step gets grounded in what actually happened.\n\nIt drafts the one decision your eval should inform. You adjust the wording until it fits.\n\nTurn “helpful” into observable must and must-not behaviours you can actually score.\n\nIt drafts a small representative set: normal, edge, adversarial and critical. If you pasted real outputs, half the cases are grounded in them. You edit any of it or add your own.\n\nFor each criterion it picks deterministic, code, LLM-judge or human, and writes the judge prompt for you.\n\nThe one threshold that decides ship or hold, stated plainly instead of left implied.\n\nEverything assembles into an eval plan, a `test-cases.csv`, judge prompts and a calibration guide, ready for your team.\n\nNot advice about evals. Ready-to-import artifacts in formats your team already uses.\n\nThe decision, the criteria, the cases and the release gate, written up so anyone on the team can follow it.\n\nA validated, importable test set you can drop straight into whatever eval tool your team already uses.\n\nReady-to-run prompts for each LLM-graded criterion, matched to the good and bad you defined.\n\nHow to spot-check each judge against 20 hand-labelled cases before it's allowed to gate a release. So you never ship on an uncalibrated score.\n\nMost AI features are shipped on vibes because turning fuzzy product intent into something measurable is genuinely hard. These are the traps Evals Coach helps you catch.\n\n“Helpful”, “accurate”, “on-brand”. Until they become observable must / must-not behaviours, no grader can score them and no one can agree whether a run passed.\n\nA green dashboard measuring the wrong thing is worse than none. Evals Coach hunts for what could score well while real users still get hurt.\n\nThe same production incident keeps recurring because nothing turned it into a test case. Feed in the failures; get back regression cases, without inventing evidence.\n\nAn uncalibrated judge quietly gating your release is a coin flip in a lab coat. Get a path to check it agrees with human labels before it holds the gate.\n\nIt’s a Claude plugin. Install it once, in the app or in Claude Code, then ask it to build you an eval.\n\nDesktop, web and Cowork. No terminal needed.\n\n`justshipai/evals-coach` and confirm\nThen ask it to build you an eval, in your own words.\n\nPrefer the command line? Two commands in your terminal.\n\nEither way, Evals Coach builds you a private page that asks Claude from inside itself. Drafting runs on the Claude account of whoever opens that page, and they consent on first use.\n\nUpdates aren't automatic, and a page you've already made won't change on its own. It's a snapshot from the moment it was published. To pick up the latest version, update the plugin and then ask for a new page.\n\nRestart Claude to apply it. [See what changed →](/changelog)\n\nA six-task blinded comparison using GPT-5.6 Sol at medium reasoning: the eval-design method behind Evals Coach versus a no-method baseline.\n\nRubric score, method vs baseline (%)\n\nTasks won, out of six\n\nCritical failures vs baseline\n\nEvals Coach encodes this same method. The run also exposed a weakness (early outputs ran long) that we’ve since tightened, adding PM usability to the rubric. Because the rubric was developed alongside the method, with one run per task and no independent replication yet, treat this as promising early evidence rather than a universal claim. [Read the full method, scores and limitations →](https://github.com/justshipai/evals-coach/blob/main/evals/results/2026-08-24-gpt-5.6-sol-medium/report.md)\n\nInstall once (it’s free!) and move from vague “vibe checks” into a precise quality bar for your product.", "url": "https://wpnews.pro/news/show-hn-evals-coach-a-claude-plugin-to-help-pms-write-good-evals", "canonical_source": "https://evalscoach.com", "published_at": "2026-09-07 12:12:45+00:00", "updated_at": "2026-09-07 12:27:13.793703+00:00", "lang": "en", "topics": ["ai-tools", "ai-products", "ai-agents"], "entities": ["Evals Coach", "Claude", "justshipai", "GPT-5.6 Sol"], "alternates": {"html": "https://wpnews.pro/news/show-hn-evals-coach-a-claude-plugin-to-help-pms-write-good-evals", "markdown": "https://wpnews.pro/news/show-hn-evals-coach-a-claude-plugin-to-help-pms-write-good-evals.md", "text": "https://wpnews.pro/news/show-hn-evals-coach-a-claude-plugin-to-help-pms-write-good-evals.txt", "jsonld": "https://wpnews.pro/news/show-hn-evals-coach-a-claude-plugin-to-help-pms-write-good-evals.jsonld"}}