Why your Claude Code routine reported success and did nothing Anthropic's Claude Code routines documentation states that a green status in the run list only means the session started and exited without an infrastructure error, not that the prompt's task succeeded, according to an analysis of the three failure modes that report success. The failures include unreadable sources that render as empty results, health checks that never ran — on the default cloud environment network access is Trusted, so requests outside Anthropic's allowlist fail with a 403 and the header x-deny-reason: host_not_allowed — and runs that fire late, which Anthropic recommends avoiding by scheduling at 9:07 rather than 9:00. The piece prescribes an eight-point pre-scheduling pass covering self-containment, named sources, verification, idempotency, honest handling of blind sources, least privilege, loud failures, and time awareness. There is one sentence in Anthropic’s routines documentation that explains more confusion than anything else on the page: a green status in the run list means the session started and exited without an infrastructure error — it does not mean the task in your prompt succeeded . That is the whole problem with unattended automation in one line. A routine runs with nobody watching. The run list is the only thing most people ever look at. And the run list is reporting on the wrong thing: it tells you the machine worked, not that the job got done. This article covers the three ways a routine fails while reporting success, and the eight-point pass that catches all three before you schedule anything. Three failures that look exactly like success 1. The source that could not be read Your mail connector times out. The routine writes “no urgent messages” . You read that over coffee as good news. It is not good news. It is no news — and the two are indistinguishable in the output. This is the most expensive failure mode in unattended work, because nothing about it looks wrong. A calendar connector that fails renders an empty schedule, and an empty schedule reads as a free day. You find out on Thursday, in the meeting you missed on Tuesday. 2. The check that never ran A health-check routine that cannot reach your endpoint, and says nothing about it, reads exactly like a healthy system. Silence is ambiguous, and ambiguity always defaults to “fine” in the reader’s head. Worth knowing the specific signature here: on the default cloud environment, network access is Trusted , which allows only Anthropic’s default allowlist. A request to a host outside that list fails with a 403 and the header x-deny-reason: host not allowed . That is a real, named, catchable error — but only if your prompt is written to catch it rather than to carry on. 3. The run that fired at the wrong time Schedule a routine exactly on the hour and it can start several minutes late; Anthropic’s own recommendation is to pick 9:07 rather than 9:00. Local Desktop tasks are worse. A 7am brief missed because the laptop was shut can run as a catch-up when the machine wakes at 11pm — and still write “today” in the present tense, about a day that is nearly over. None of these three produce an error. All three produce green. The eight-point pass Run any routine prompt — a template you accepted, or one you wrote yourself — through these eight checks before you schedule it. Each one closes a specific hole. | | Check | The question it answers | |---|---|---| | 1 | Self-contained | If a stranger ran this with no context, would it work? | | 2 | Sources named | Can every number in the output be traced to something read this run? | | 3 | Verified | Did the routine re-check its own work before reporting? | | 4 | Idempotent | If this runs twice, do I get one result or two? | | 5 | Honest when blind | Does an unreadable source look different from an empty one? | | 6 | Least privilege | Which connectors are attached, and which of them can write? | | 7 | Fails loudly | When it breaks, do I find out — and does it tell me what to do? | | 8 | Time-aware | If this fires nine hours late, is the output still true? | Check 1 — Self-contained A routine session has no memory of your conversations. It cannot ask a follow-up question. Anything the prompt assumes you will supply is simply missing. Write it as instructions to a competent stranger who has never met you: name the files, name the windows, name the thresholds. Check 2 — Sources named Every figure in the output should be traceable to something the routine actually read during this run. The failure this prevents is subtle: a plausible number that came from nowhere. Require the prompt to list its sources and counts, and a fabricated figure has nowhere to hide. Check 3 — Verified Add a step where the routine re-opens its own output and checks it against what it read. Counts match the items listed. Times appear as read, not rounded. Conclusions trace to evidence. Make that step able to fail the run — a verification section that always passes is decoration. Check 4 — Idempotent Routines fire more than once. A GitHub trigger fires on every push to an open pull request, and Claude Code does not reuse sessions across events — five pushes means five independent runs. Key your state on something stable the head commit SHA, not the PR number so a repeat run writes nothing instead of writing a duplicate. Check 5 — Honest when blind This is the one almost every prompt fails, and the one that costs the most. The rule is short: A source that could not be read must never render as a source with nothing in it. The fix is a few lines, and it is the single highest-value edit you can make to any routine prompt: For each source, record exactly one of two outcomes: read: