How I review an AI agent's diff (the diff comes last) A developer outlined a review workflow for AI-agent-generated code diffs that checks the agent's claims against evidence the model did not write, rather than reading the diff first. Applied to a Spring Boot 2.7-to-3.5 migration performed by Claude Code, the process — a read-only reviewer subagent plus a curl-based old-versus-new response diff — caught a public API change (an added details field and reformatted timestamp) that 47 passing tests had missed. The developer recommends requiring the build's own test-count line as proof of a green build and having a human explicitly decide every choice the agent made. When an AI agent hands you a diff, reading it line by line is the last step, not the first. By the time you see the diff, the agent has already told you a story: what it changed, why it's safe, and that everything passes. A senior review checks that story against evidence. Here's the order I use, walked through on one real step: a legacy demo service moving from Spring Boot 2.7 to 3.5, done by Claude Code and replayed from its transcripts. Here's how that step ended, in the agent's own words: The independent review confirms no hidden regressions. Green build. All 47 tests passing. Those are claims. The question for each one is the same: can I check it against output the model didn't write? Scroll back to the final build. It ran in quiet mode, so a success prints almost nothing. Claude tried to read the exit code, and that command needed an approval nobody was there to give. So it concluded success from the absence of an error. It was right this time. The build really was green. But silence is weak evidence, and the same reasoning in another run could hide a failure. The fix is cheap: ask for the line the build itself prints, with the test count, pasted into the report. Something like: Tests run: 47, Failures: 0, Errors: 0, Skipped: 0 If the agent can't show you that line, you don't have a green build. You have a sentence about one. Not the changes, the choices. There was one, flagged honestly at the end of the summary. A hook fired before a human could answer, so Claude defaulted to "option B": a choice about what the public API returns for unknown URLs. Maybe it's the right choice. That's not the point. A decision about your public API was made by the model, so it needs an explicit yes or no from a person before the step is done. Make a list of every choice like that in the session, and answer each one. We use a reviewer subagent: one markdown file in .claude/agents/ , its own context, and no edit or write tools. It sees the diff, not the reasoning behind it. Its checklist is ordered by damage: public API first, then security, the frozen tests, configuration, dependencies. Every finding gets a severity, and the report ends with one verdict: ship, or fix first. The report earned its keep. It flagged the model's decision as an unreviewed change to the error body. It spotted deprecated annotations in two test classes that a hook had kept Claude from touching. And a security rule that was probably dead. Then the verdict: ship. Now read how it knows . The build: green, because no output means success. The public API: the error tests pass unmodified, so no drift. The reviewer trusted the same tests the author trusted. A second opinion built on the same evidence isn't independent. It's a filter. Useful, but read its reasons, not just its verdict. This is the check that found the problem, and it's almost embarrassingly simple. Start the old version and the new one. Send both the same request. Compare. old version on :8080, new one on :8081 — adapt the paths and users to your API for path in /api/parcels/does-not-exist /api/parcels/42 /api/stats; do curl -s -u clerk:clerk "http://localhost:8080$path" | jq -S . old.json curl -s -u clerk:clerk "http://localhost:8081$path" | jq -S . new.json diff -u old.json new.json || echo "^^^ $path changed" done An unknown URL on 2.7, then on the new build: a details field that was never there, and a timestamp in a different format. Forty-seven green tests, a ship verdict, and the public API had changed. Why did the tests miss it? The test for that URL looked at four fields and ignored the rest of the body. A new field couldn't turn it red. Green meant those four fields were fine, not that the body was. Build your own list of requests your clients depend on. Put at least one request that should fail in it, and one that tests who's allowed. Every difference gets a human decision: bring the old behavior back, or keep the new one deliberately, with a note that says why. When a check like this finds a gap, fix the contract before the code. A person tightened the test exactly five fields, the old timestamp format and ran it against the old version first. It passed there, so it described what clients really got. Then Claude got both real responses as evidence, a note of what the human had already changed, and one instruction with a clear finish line: restore the default error body exactly as on 2.7, and leave the tests alone. It took its own handler out and changed just the message through Spring Boot's error attributes. Same request again: five fields, old timestamp style, matches 2.7. One more habit from that exchange. Claude's last line separated its edits from the human's. In a diff with two authors, require that every time. You need to know whose work you're approving. Claude's plan included calling the running app by hand. Launching it in the background needed an approval, and in a headless run there was no one to give it. So it dropped that check and wrote that the automated tests already covered it. Mostly. But not one test in the suite ever requested the OpenAPI description or the Swagger page. So those went on the human's list: Everything the agent couldn't verify is your checklist. Look for the words "should", "already covered" and "assumed" in its report. Only now. Read it for scope, and for the next change, not just this one. In a later step Java 25 , Claude declared Lombok for the compiler plugin with a hard-coded version, and left a note: a human should bump it by hand if Lombok is ever upgraded. The review changed one line, to use the version property the Spring Boot parent already manages. Now the next framework upgrade moves the annotation processor and the library together. A note that asks for manual work later is a small bug with no date on it yet. None of this is specific to Claude Code. Copilot, Cursor and Codex all document reviewer agents of their own, and the checks don't care which agent wrote the diff. The video version, with the transcripts and both responses on screen: Personal project, views my own. The sessions and transcripts are real; this write-up was drafted with AI from the video and its transcripts.