# How to Use /goal and Verification Checks in Claude Code

> Source: <https://www.mindstudio.ai/blog/claude-code-goal-command-verification/>
> Published: 2026-10-11 00:00:00+00:00

# How to Use /goal and Verification Checks in Claude Code

A practical guide to Claude Code's /goal command and verification checks, letting Claude run tasks autonomously until criteria pass.

## What does /goal do in Claude Code?

The `/goal` command tells Claude Code to keep working on a task, turn after turn, until it meets criteria you define, instead of stopping after one response and waiting for you to type “continue.” After each turn, a separate small model reads the goal and the latest output and decides whether the condition has been met. If not, Claude takes another turn automatically. It’s the mechanism that lets you hand off a task and walk away instead of babysitting the chat.

## TL;DR

- **Verification checks** (a script, a pass/fail test, a rubric document) let Claude confirm its own work instead of you manually reviewing every output.
- **/goal automates the follow-up loop** , removing the need to keep typing “continue” until Claude satisfies your criteria.
- A **judge model** (not Claude itself) evaluates whether the goal is met after each turn, and it only sees the chat and Claude’s output, not the actual files or commands.
- Vague goals like “the report is finished” risk **premature completion** because the judge has no independent way to check reality.
- Good goals, according to Anthropic’s own guidance, specify **one end state, how to prove it, what must not change, and a turn limit** .
- Combining a written **definition of done** with an executable check script is what makes autonomous loops reliable rather than risky.
- This workflow only works well if your **Claude.md and skills are lean** , since bloated instructions get followed less and cost more tokens per turn.

- ✕a coding agent
- ✕no-code
- ✕vibe coding
- ✕a faster Cursor

The one that tells the coding agents what to build.

## Why does verification matter more than planning now?

For most of the past year, the standard Claude Code advice was to start every session in plan mode, laying out steps before letting the model touch any code. That made sense when older models tended to jump straight into execution without thinking things through. Newer models plan as part of normal execution, so the old plan-first ritual has become less necessary. Boris Cherny, who built Claude Code, has said as much publicly: plan mode solved a model limitation that largely no longer exists.

What replaced it is a shift from controlling the process to verifying the result. Instead of reviewing a plan before work starts, the more effective pattern is to describe the task, the guardrails around it, and what “done” actually looks like, then let Claude work it out and check the outcome afterward. That’s a meaningful change in how you spend your attention: less time reviewing intentions, more time reviewing evidence.

Anthropic’s own guidance now treats this as a first-tier recommendation. Giving Claude a way to verify its own work can substantially improve the quality of the final result, because the model can catch and fix its own mistakes before you ever see them, rather than you catching them after the fact.

## How do you build a verification check?

A verification check is anything that can return a clear pass or fail against your definition of done. The most reliable version is an actual script. If your task is a weekly client report, for example, the definition of done might be “every figure matches the CRM export and every section of the template is filled in.” You can turn that directly into a small script, something like `check_report`, that validates exactly those two conditions.

The workflow then looks like this:

1. Claude completes the task.
2. It runs the check script against its own output.
3. If something fails (a figure doesn’t match the export, say), Claude fixes the issue and reruns the check.
4. Once the check passes, Claude shows you the output as evidence.

You’re no longer running the check yourself or manually comparing numbers. You’re reading the evidence that the check already passed.

Checks don’t have to be code. A brand voice document that Claude compares a draft against works the same way conceptually. But the more deterministic the check, the more reliably it catches errors. A script that fails loudly on a mismatched number is more trustworthy than a subjective read-through.

One important detail: you don’t need to tell Claude to “double check its work” inside your prompt. Current models already attempt this without being asked, so that instruction mostly just adds tokens without improving results. What actually moves the needle is giving Claude a concrete, testable target to check against.

## How does the /goal command actually work?

## One coffee. One working app.

You bring the idea. Remy manages the project.

Once you have a verification check, `/goal` is what keeps Claude iterating against it without you manually prompting each retry. After every turn, a separate, smaller model (Haiku-class) reads your stated goal plus Claude’s latest output and judges whether the condition has been satisfied. If it hasn’t, Claude takes another turn to try again. This repeats until the goal is met or a turn limit is reached.

The catch is that this judge model doesn’t run commands or read files on its own. It only sees what’s in the chat and what Claude reports back. That means if your goal is loosely worded, like “the report is finished,” the judge has no independent way to verify that claim and may decide the condition is met when it actually isn’t.

This is why the goal statement itself has to be explicit and tied directly to your verification check. Anthropic’s guidance describes a good goal as having four parts:

- **One end state** (a single, unambiguous thing that must be true)
- **A way to prove it** (reference the check script or specific evidence)
- **What must not change** (your guardrails, like “don’t touch the template”)
- **A turn limit** (to stop the loop from running indefinitely if something goes wrong)

An example goal built around the report scenario above might read something like: “The check_report script passes and you show me the output. The template remains unchanged. Stop after 15 turns.” That single instruction gives Claude a concrete loop to run inside, with a hard stop if it can’t get there.

## Is /goal worth using over just prompting manually?

For tasks you’d otherwise supervise turn by turn, yes. The value of `/goal` is time, not capability. It doesn’t make Claude smarter or more accurate. It removes the manual labor of repeatedly typing “continue” or re-explaining what still needs fixing. Your job shifts to writing a precise goal once and then validating the final evidence, rather than steering every intermediate step.

The tradeoff is that a poorly specified goal can backfire. Because the judge model can’t independently inspect files or run commands, it’s trusting the honesty and clarity of what Claude reports. A vague or overly broad goal increases the risk of the loop ending early on a false positive. The fix isn’t to avoid `/goal`, it’s to write goals as specifically as you’d write the verification check itself.

## How does this fit with Claude.md and skills?

Autonomous loops only work as well as the instructions underneath them. Anthropic has pushed significant cleanup into Claude Code’s own default instructions, on the reasoning that long instruction files cost more tokens and get followed less reliably as they grow. The general guidance is to keep a project’s `Claude.md` short, under roughly 200 lines, and to treat it like an index rather than a manual.

Practically, that means `Claude.md` should hold only things that are always true: tools and commands Claude can’t guess, anything you do differently from Claude’s default approach, recurring mistakes worth flagging, and hard boundaries. Everything else should live in separate reference files that Claude opens only when needed, referenced by file path rather than loaded wholesale into every session.

Skills work on a similar principle. Rather than writing a skill that documents an entire process up front, the more effective approach is to run the task without a skill first, note exactly where Claude got it wrong, and build the skill around those specific corrections. Many effective skills end up being short, just a handful of lines addressing particular gotchas, because the model already handles most of the task correctly out of the box.

## Frequently Asked Questions

### What is the difference between a verification check and just asking Claude to double-check its work?

A verification check is a concrete, testable condition (a script, a rubric, a specific comparison) that returns pass or fail. Asking Claude to “double check” is a vague instruction that current models largely do automatically anyway. A defined check gives both Claude and the judge model something objective to evaluate against.

### Can /goal run indefinitely if the check never passes?

No, as long as you set a turn limit in the goal statement. Anthropic’s guidance recommends always including one specifically so a stuck or impossible task doesn’t loop forever.

### Does the judge model that evaluates /goal actually run my code or check my files?

No. It only reads the chat history and Claude’s reported output, not the underlying files or commands. That’s why the goal needs to be specific and tied to verifiable evidence rather than a vague description of completion.

### Should I still use plan mode in Claude Code?

It’s no longer considered essential advice the way it was earlier on. Newer models tend to plan as part of execution, so many practitioners now favor describing the task, guardrails, and definition of done up front, then reviewing the result, rather than reviewing a plan before any work starts.

### How short should Claude.md actually be?

General guidance points to keeping it under roughly 200 lines, functioning as an index that points to other reference files rather than containing every instruction directly.
