# Show HN: Evals Coach – a Claude plugin to help PMs write good evals

> Source: <https://evalscoach.com>
> Published: 2026-09-07 12:12:45+00:00

Evals Coach is a Claude plugin that guides you through designing evals,

one decision at a time.

Step 3, after reviewing twelve real outputs. Your notes become named failure modes, each one traceable back to the outputs it came from, and ranked by how often it actually happened.

You bring product judgement. It brings the eval-design method: criteria, cases, graders and thresholds.

Start from a PRD, a feature idea, or a sentence describing what you’re building. No codebase required.

The output isn’t tied to any eval platform, so it imports into whatever your team already runs.

Describe your feature; Evals Coach drafts each step, you edit it, and the outputs assemble as you go.

A sentence or two on what it does and the worst thing it could get wrong. If it's already live, paste a few real outputs too, and every later step gets grounded in what actually happened.

It drafts the one decision your eval should inform. You adjust the wording until it fits.

Turn “helpful” into observable must and must-not behaviours you can actually score.

It drafts a small representative set: normal, edge, adversarial and critical. If you pasted real outputs, half the cases are grounded in them. You edit any of it or add your own.

For each criterion it picks deterministic, code, LLM-judge or human, and writes the judge prompt for you.

The one threshold that decides ship or hold, stated plainly instead of left implied.

Everything assembles into an eval plan, a `test-cases.csv`, judge prompts and a calibration guide, ready for your team.

Not advice about evals. Ready-to-import artifacts in formats your team already uses.

The decision, the criteria, the cases and the release gate, written up so anyone on the team can follow it.

A validated, importable test set you can drop straight into whatever eval tool your team already uses.

Ready-to-run prompts for each LLM-graded criterion, matched to the good and bad you defined.

How to spot-check each judge against 20 hand-labelled cases before it's allowed to gate a release. So you never ship on an uncalibrated score.

Most AI features are shipped on vibes because turning fuzzy product intent into something measurable is genuinely hard. These are the traps Evals Coach helps you catch.

“Helpful”, “accurate”, “on-brand”. Until they become observable must / must-not behaviours, no grader can score them and no one can agree whether a run passed.

A green dashboard measuring the wrong thing is worse than none. Evals Coach hunts for what could score well while real users still get hurt.

The same production incident keeps recurring because nothing turned it into a test case. Feed in the failures; get back regression cases, without inventing evidence.

An uncalibrated judge quietly gating your release is a coin flip in a lab coat. Get a path to check it agrees with human labels before it holds the gate.

It’s a Claude plugin. Install it once, in the app or in Claude Code, then ask it to build you an eval.

Desktop, web and Cowork. No terminal needed.

`justshipai/evals-coach` and confirm
Then ask it to build you an eval, in your own words.

Prefer the command line? Two commands in your terminal.

Either way, Evals Coach builds you a private page that asks Claude from inside itself. Drafting runs on the Claude account of whoever opens that page, and they consent on first use.

Updates aren't automatic, and a page you've already made won't change on its own. It's a snapshot from the moment it was published. To pick up the latest version, update the plugin and then ask for a new page.

Restart Claude to apply it. [See what changed →](/changelog)

A six-task blinded comparison using GPT-5.6 Sol at medium reasoning: the eval-design method behind Evals Coach versus a no-method baseline.

Rubric score, method vs baseline (%)

Tasks won, out of six

Critical failures vs baseline

Evals Coach encodes this same method. The run also exposed a weakness (early outputs ran long) that we’ve since tightened, adding PM usability to the rubric. Because the rubric was developed alongside the method, with one run per task and no independent replication yet, treat this as promising early evidence rather than a universal claim. [Read the full method, scores and limitations →](https://github.com/justshipai/evals-coach/blob/main/evals/results/2026-08-24-gpt-5.6-sol-medium/report.md)

Install once (it’s free!) and move from vague “vibe checks” into a precise quality bar for your product.
