Claude Code made 6 mistakes upgrading a legacy Spring Boot service. None reached a commit. A developer ran Claude Code through a full upgrade of a legacy Java 11 / Spring Boot 2.7 service to Java 25 / Spring Boot 4.1, a roughly two-hour, $24 session that ended with 47 tests passing. Along the way the agent made six mistakes — including a summary claiming 56 tests passed, two false claims in its upgrade plan, and an autonomous product decision forced by a Stop hook — none of which reached a commit because each was caught by build output, plan review, or hook fixes. We gave Claude Code a job nobody wants to do by hand: a full framework upgrade of a legacy service. Java 11 and Spring Boot 2.7 at the start, with Springfox for the API docs, the old WebSecurityConfigurerAdapter that Spring Security 6 removed, and a TODO from 2019. Java 25 and Spring Boot 4.1 at the end. It took about two hours of Claude time and roughly $24 at API prices. It ended green, 47 tests passing. The result isn't the interesting part. On the way, Claude made six mistakes, and not one of them made it into a commit. Here they are, straight from the session transcripts, with whatever caught each one. The service is a demo, built on purpose to look like the legacy code a lot of us keep alive a parcel tracking REST API called ShipTrack . The sessions are real. After the move to Java 21, Claude modernized the code and wrote a tidy summary. Everything in it was accurate except one number: it said 56 tests passed. The build had run 47. Small, but it's the cleanest example of the whole problem. A model's summary describes the work. It doesn't prove it. The evidence is in the build output. That's why the definition of done in our CLAUDE.md doesn't say "tests pass". It names a command: Definition of done - ./mvnw verify passes. - No new deprecation warnings in the compiler output. - UPGRADE NOTES.md says what a reviewer should double-check. Caught by: reading the build's own output, not the summary. For the jump to Spring Boot 3.5, Claude didn't touch code at first. It went into plan mode, explored with subagents, and wrote a plan: every file to change, the order of operations, even the limits our hooks put on the work. A good plan. It also contained two false statements. It said Spring Boot 3 no longer had a path-matching property it still does . And it said server.max-http-header-size kept its name in 3.5. It didn't: since 3.0 it's server.max-http-request-header-size . Checking took one command. Spring Boot jars ship their own configuration metadata, including every deprecated key and its replacement: unzip -p path/to/spring-boot-autoconfigure-3.5.x.jar META-INF/spring-configuration-metadata.json \ | jq '.properties | select .name | test "header-size" ' A long plan with two wrong claims is normal. That's exactly why you review a plan like a pull request, before any code changes. A wrong assumption costs one sentence to fix in a plan. After the edits, it costs an afternoon. Caught by: plan review plus the configuration metadata. This one is my favorite, because it was partly our fault. After the jump to 3.5, two tests failed. Since Spring Framework 6.1, unknown URLs go through the static resource handler, and the 404 body now says "No static resource". A real change to the public API. Claude did the right thing: it laid out two options and asked us to choose. Then it tried to stop and wait for an answer. Our Stop hook said no. Its job is to keep Claude working until the build is green, and the build was red. So Claude picked option B itself, got the build green, and flagged the choice for review. The reasoning was sensible. But a product decision that belonged to a human was made by the model, and our guardrail pushed it there. The fix went into the hook: when the only failing tests are characterization tests, a red build means behavior changed, and that's a question for a person. Now the hook lets Claude stop and ask. For reference, the wiring is just two entries in .claude/settings.json : { "hooks": { "PreToolUse": { "matcher": "Edit|Write", "hooks": { "type": "command", "command": ".claude/hooks/protect-characterization-tests.sh" } } , "Stop": { "hooks": { "type": "command", "command": ".claude/hooks/require-green-build.sh" } } } } The first one blocks edits to the tests that freeze today's behavior exit code 2 blocks the tool call . The second one refuses to end the turn on a red build. If you copy the idea, give the Stop hook a way out for "this needs a human", or it will make decisions for you. Caught by: reading the transcript. The reviewer subagent flagged it too. Now the build was green, all 47 tests passed, and the reviewer subagent said "ship it". Before committing we did one boring thing: ran the old version and the new one side by side, sent the same unknown URL to both, and compared the real responses. Same status, same message. But the new body had an extra details field, and the timestamp was formatted differently. Two changes to the public API, with every test green. Why? The test for unknown URLs checked four fields. It never checked that there was nothing else. A test only protects what it asserts. The fix started with the contract, not the code. A human made that test stricter exact number of fields, exact timestamp format and ran it against the old version first, to prove it described what clients really got. Then Claude got the evidence and one instruction: go back to Spring Boot's default error body and leave the tests alone. It removed its own handler and overrode just the message, inside Spring Boot's error attributes. Five fields, old timestamp format. Then we committed. Caught by: a human comparing real responses. On to Spring Boot 4.1. Two characterization tests red again. In one, the error body for malformed JSON had lost its message field. Claude investigated and gave a clear diagnosis: Spring Boot 4 no longer puts the message in the response, so we'd need extra code to bring it back. Plausible. Wrong. Spring Boot silently ignores properties it doesn't recognize. One more look at the configuration metadata showed the key had been renamed in 4.0: server.error.include-message became spring.web.error.include-message . The old key was ignored, so the message disappeared. No code needed. The plan had also skipped spring-boot-properties-migrator , which our upgrade instructions explicitly ask for. It exists to report exactly this kind of rename. Caught by: the configuration metadata, again. Last step, Java 25. Should be a one-line change. Claude switched the version, ran the build: green. Then it checked for warnings with something like mvn compile | grep -i warning . Nothing. Looks done. Except that first build wasn't clean, so classes compiled earlier were still sitting in target/ . And the compile had actually failed. The error lines just didn't contain the word "warning", and the pipe swallowed the exit code. Before finishing, Claude ran one more full clean build, and the truth came out: cannot find symbol . The Lombok getters were missing. Since JDK 23, javac no longer runs annotation processors found on the class path unless you ask it to, so Lombok has to be declared in the compiler plugin. Claude recognized it and said its earlier check was flawed. Output you filtered is not a result. True for humans, true for models. Our Stop hook now runs a clean build whenever the pom.xml changes. Caught by: a clean build. Look at the list of catches: the build output, a plan review, the transcript and the reviewer, a human comparing responses, the configuration metadata, a clean build. No single guardrail caught everything. Each one caught something the others missed. And Claude did almost all of the work, and did it well. It found a real production bug before the upgrade even started, refused to work around a guardrail, and checked APIs in the actual jars. The human part was small. But it was the part that mattered: building the guardrails and making the decisions. Every one of those guardrails is something you can add to your own repo this week. The full walkthrough, with the transcripts on screen: Personal project, views my own. The sessions, numbers and transcripts are real; this write-up was drafted with AI from the video and its transcripts.