{"slug": "claude-sonnet-5-5-for-code-review-more-catches-than-sonnet-5-in-half-the-time", "title": "Claude Sonnet 5.5 for code review: More catches than Sonnet 5, in half the time", "summary": "Anthropic released Claude Sonnet 5.5 at the same price as Sonnet 5 ($2 per million input tokens, $10 per million output tokens), claiming 30% faster output and up to 30% lower cost per task. In CodeRabbit's code-review evaluation on 13 hardest known-bug cases, Sonnet 5.5 caught 6 of 13 issues versus 4 for Sonnet 5, at 41.2% actionable precision (40.0% for Sonnet 5), in about half the wall-clock time and roughly 60% less per review at list prices. CodeRabbit notes the 13-case set is small and a second run on 44 open-source pull requests confirmed the speed and comment-volume findings.", "body_md": "Anthropic has released [Claude Sonnet 5.5](https://www.anthropic.com/claude-sonnet-5-5), the second model in the Claude 5.5 family, at the same price as Sonnet 5 and with a promise of 30% faster output and up to 30% lower cost per task. We ran it through CodeRabbit’s review pipeline the same way we tested [Opus 5.5](https://www.coderabbit.ai/blog/opus-5-5-model-review) earlier this month. The question for most teams is not whether Sonnet 5.5 beats Opus. It is whether it fixes the one thing that held [Sonnet 5](https://www.coderabbit.ai/blog/claude-sonnet-5-review) back as a reviewer: clean comments, but too many bugs left uncaught.\n\nOn our 13 hardest known-bug cases, it does. Sonnet 5.5 caught **6 of 13 issues** through actionable comments, against **4 for Sonnet 5**, at **41.2% actionable precision** (40.0% for Sonnet 5), with 17 reported comments to Sonnet 5’s 15. It did this in **about half the wall-clock time**, and at list prices the Claude model calls cost **about 60% less per review**, twice the saving Anthropic advertises. Four of its catches were bugs Sonnet 5 waved through; it missed two that Sonnet 5 caught, so this is a “different misses” story as much as a “more catches” one.\n\nThirteen cases is a small set, and we say so throughout. A second, larger run on 44 open-source pull requests confirms the speed and comment-volume side of the story. Together they make this the first Sonnet release we would consider for the main review pass rather than only for comment quality.\n\n*How the Sonnet line has evolved in our code-review evaluations, from Sonnet 4 to Sonnet 5.5. Qualitative direction only.*\n\nWe have put every Sonnet on the bench since Sonnet 4, so this release fits a pattern. [Sonnet 4.5](https://www.coderabbit.ai/blog/claude-sonnet-45-better-performance-but-a-paradox) added reasoning depth, with a hedging habit. Sonnet 4.6 became the highest-coverage Sonnet we had measured, catching about 63% of known issues, but at 29% precision it commented on everything. [Sonnet 5](https://www.coderabbit.ai/blog/claude-sonnet-5-review) traded that coverage for precision: cleaner comments, but coverage fell to about 50% and nitpicks multiplied. Sonnet 5.5 is the first release in the line to move coverage back up without giving the precision back.\n\n## What’s new in Claude Sonnet 5.5\n\nAnthropic positions Sonnet 5.5 as the faster, lower-cost complement to Opus 5.5: strongest at well-scoped everyday work such as fixing bugs and producing documents, slides, and spreadsheets, while Opus 5.5 stays the model for open-ended work that needs sustained judgment. Five things from the [launch announcement](https://www.anthropic.com/claude-sonnet-5-5) matter for teams building review and coding workflows:\n\n- **Same price, fewer tokens.** Sonnet 5.5 costs $2 per million input tokens, $10 per million output tokens, $0.20 per million cache reads, and $2.50 per million cache writes, unchanged from Sonnet 5 and half of Opus 5.5’s $4 and $20. Anthropic says it typically needs far fewer tokens for the same work, up to 30% less per task, and generates output 30%+ faster. Our review runs show a larger gap than that; see the cost section below.\n- **Large gains on agentic coding benchmarks.** On Anthropic’s launch table (below), Sonnet 5.5 jumps from 10.3% to 70.6% on Terminal-Bench 4.0, ahead of Opus 5.5, and lands within about three points of Opus 5.5 on most other rows; AA-Briefcase is the widest gap, at 11 points. Our code-review results follow that shape, except that on our hardest cases the gap to Opus 5.5 is wider.\n- **Thinking is on by default, and effort is the dial.** Adaptive thinking is the default; effort (low, medium, high, xhigh, max) controls cost and depth, with Medium the default in the Claude apps and High on the Claude Platform. Thinking can still be turned off at the lower effort levels, and teams that ran Sonnet 5 with thinking off need to move to the new`between_tools` setting when they migrate. Our thinking-off run below uses that path.\n- **API behavior carried over from Opus 5.5.** Forced tool use is retired in favor of structured outputs, and tool definitions can be added or changed mid-conversation without invalidating the prompt cache or earlier thinking blocks. Anthropic’s guidance also notes the model follows instructions literally, so phrasing like “minimize tool calls” is obeyed to the letter, and that at low effort it can report a code change as done without running a check unless the prompt asks for one.\n- **Safeguards and deployment.** Sonnet 5.5’s cyber capabilities are comparable to Opus 5’s, so it is the first Sonnet to ship with Opus-style cyber safeguards: routine bug fixing is unaffected, but higher-risk security requests fall back to Sonnet 5. Biology safeguards are unchanged from Sonnet 5.\n\n| Anthropic’s launch benchmarks | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | \n| Terminal-Bench 4.0 (agentic coding) | 70.6% | 10.3% | 66.4% | \n| FrontierCode 1.1 Main (agentic coding) | 52.1% at Xhigh · 46.2% at Max | 42.4% | 54.4% | \n| GDPval-AA v2.1 (knowledge work) | 1844 | 1449 | 1846 | \n| AA-Briefcase v1.1 (knowledge work) | 1811 | 1359 | 1822 | \n| Humanity’s Last Exam, with tools | 64.5% | 54.9% | 67.7% | \n| OSWorld 2.1, partial (computer use) | 80.1% | 57.0% | 81.8% | \n| Chartography, no tools (chart recognition) | 61.6% | 15.6% | 64.4% | \n\nFor reviewers, the practical change is what we measured below: Sonnet 5.5 spends far fewer tokens per review call than Sonnet 5 and still finds more.\n\n## What we tested\n\nWe used two benchmarks. **Signal** is the same 13 harder cases we used for the Opus 5.5 evaluation: each is a real pull request from an open-source project (Elasticsearch, Puma, vLLM, Cilium, axios, and Next.js) with one verified issue the reviewer should catch. Nine are difficulty 3, three are difficulty 4, and one is difficulty 5. **OSS August** is our broader set, 85 known issues across 44 open-source pull requests, weighted toward logic errors, API misuse, race conditions, null references, and security issues.\n\n| Configuration | What it is | \n| Sonnet 5.5, thinking on | Low, medium, and high effort for the trivial, junior, and senior review cohorts respectively, with adaptive thinking left on. Run on both benchmarks. | \n| Sonnet 5.5, thinking off | Same effort ladder with thinking disabled, which the model allows at these three effort levels. Signal only. | \n| Sonnet 5 | Same effort ladder, run the same day as the Sonnet 5.5 configurations. Both benchmarks. | \n\nAll runs replayed the same recorded file summaries, walkthroughs, and layer grouping from a frozen cassette, so they reviewed identical inputs; the review model and its thinking setting are what changed between them. On Signal, an independent judge scored every comment with three votes against the known issue, and only majority-PASS comments count.\n\n**Known issues caught** is the share of the 13 cases where at least one regular actionable comment passed (outside-diff and nitpick comments excluded). **Actionable precision** is passing actionable comments divided by all actionable comments. **Reported comments** is the post-pipeline count after verification, deduplication, and filtering; someone still has to read each one.\n\n## TL;DR\n\n| Signal · 13 patterns | Known issues caught | Caught incl. outside-diff | Actionable precision | Reported comments | Mean time per review | \n| **Sonnet 5.5, thinking on** | **6/13 · 46.2%** | **8/13** | **41.2%** | 17 | **5:27** | \n| Sonnet 5.5, thinking off | 5/13 · 38.5% | 7/13 | 38.5% | 13 | 5:58 | \n| Sonnet 5 | 4/13 · 30.8% | 4/13 | 40.0% | 15 | 9:55 | \n| Opus 5.5 Standard (from our Opus 5.5 post) | 8/13 · 61.5% | 10/13 | 66.7% | 21 | — | \n| Opus 5.5 Max (from our Opus 5.5 post) | 10/13 · 76.9% | 10/13 | 52.0% | 25 | — | \n\n*Figure 1. Known issues caught through actionable comments on the 13 Signal cases. Squares include findings outside the changed lines. Opus 5.5 figures are from our September evaluation of the same cases.*\n\nSonnet 5.5 with thinking on catches half again as many known issues as Sonnet 5 on this set, at the same precision, with two more comments. Sonnet 5 remains the precise-but-quiet reviewer we described in June, and on this harder set that quietness cost it: it caught the fewest issues of any configuration.\n\n[Opus 5.5](https://www.coderabbit.ai/blog/opus-5-5-model-review) caught 8 of 13 (Standard) and 10 of 13 (Max) in the same cases. Sonnet 5.5 does not close that gap, and we come back to it below.\n\n## What Sonnet 5.5 means for your code reviews\n\n### Same precision, more of it\n\nSonnet 5.5 and Sonnet 5 landed at almost the same actionable precision, 41.2% versus 40.0%, so the extra catches did not come from posting more speculative comments. Sonnet 5.5 posted 17 reported comments to Sonnet 5’s 15, with fewer nitpicks and no comments labeled critical, where Sonnet 5 labeled two critical and one of those was wrong. The two models write comments of similar quality; Sonnet 5.5 simply lands more of them on the bug.\n\n### On 44 real pull requests, it is quieter and twice as fast\n\nThe Signal set is small, so we also ran Sonnet 5.5 and Sonnet 5 on our Open-Source benchmark: 85 known issues across 44 pull requests.\n\n| OSS August · 44 PRs | Reported comments | Critical / major / minor | Nitpicks | Mean time per review | Median | Total for 44 reviews | \n| **Sonnet 5.5, thinking on** | **111** | 4 / 50 / 57 | **9** | **6:33** | **5:44** | **4 h 49 min** | \n| Sonnet 5 | 146 | 14 / 87 / 45 | 30 | 13:31 | 13:49 | 9 h 55 min | \n\n*Figure 2. OSS August benchmark, 44 pull requests: time per review and post-pipeline comment volume. Judge scoring pending.*\n\nSonnet 5.5 posted 24% fewer comments than Sonnet 5 with a milder severity mix, and a third of the nitpicks. Sonnet 5 labeled 14 comments critical to Sonnet 5.5’s four, and added 30 nitpicks, the same nitpick-heavy pattern we reported in June. The judged Signal results say Sonnet 5’s extra comments did not translate into more catches; the OSS judge run will tell us whether that holds at scale, and whether Sonnet 5.5’s lower volume costs it any coverage. Until then, read the volume numbers as workload, not quality.\n\nThe latency result needs no such caveat. Across 44 reviews, Sonnet 5.5 averaged 6:33 per review against 13:31 for Sonnet 5, and 5:44 against 13:49 on the median review. Sonnet 5 also produced four generations that ran longer than ten minutes; Sonnet 5.5 produced none.\n\n**Want to see what this looks like on your own code?** [Try CodeRabbit on your next PR](https://coderabbit.link/NcmblIe). It is free to start and takes about two minutes to connect a repo.\n\n## Thinking on or off?\n\nWe ran Sonnet 5.5 twice on Signal: once with adaptive thinking on across the effort ladder, once with it turned off. Unlike Opus 5.5, which rejects explicit thinking toggles, Sonnet 5.5 allows thinking off at low, medium, and high effort (the setting Anthropic’s migration guide now calls `between_tools`), so this is a configuration teams can actually deploy.\n\n| Sonnet 5.5 · Signal | Known issues caught | Actionable precision | Reported comments | Nitpicks | Mean time per review | \n| Thinking on | 6/13 | 41.2% | 17 | 2 | 5:27 | \n| Thinking off | 5/13 | 38.5% | 13 | 3 | 5:58 | \n\nThinking on won by one case and 2.7 points of precision, while producing four more comments. The two configurations disagreed in both directions: thinking-on caught the vLLM config-context bug and a streaming tool-call serialization case that thinking-off missed, while thinking-off caught the Elasticsearch terms-enum case through a regular comment that thinking-on only reached outside the diff.\n\nThe cost of thinking was smaller than we expected. On the core review calls, thinking-on produced about 5,800 output tokens per call versus 2,900 with thinking off, on identical input, and mean latency per review was 31 seconds *shorter*, which is within noise but shows the extra tokens did not slow reviews down. The default should be thinking on.\n\n## What a review actually costs\n\nSonnet 5.5 and Sonnet 5 share a price list, so the cost difference between them is entirely a token difference. We priced the Claude model calls in each run at Anthropic’s published rates ($2 input, $10 output, $0.20 cache read, $2.50 cache write per million tokens). The shared smaller models that handle summaries and verification are the same on both sides and are left out.\n\n| Claude model calls at list price | Signal (13 reviews) | Per review | OSS August (44 reviews) | Per review | \n| **Sonnet 5.5, thinking on** | **$6.16** | **$0.47** | **$20.32** | **$0.46** | \n| Sonnet 5.5, thinking off | $5.37 | $0.41 | — | — | \n| Sonnet 5 | $15.06 | $1.16 | $50.95 | $1.16 | \n\nOn both benchmarks, Sonnet 5.5’s Claude calls cost about 40% of Sonnet 5’s for the same reviews, a saving of roughly 60% per review. Anthropic’s launch figure is “up to 30% less per task”; a code-review workload, where Sonnet 5 read the same files repeatedly and wrote long deliberations, sits well beyond that. Thinking on added about 15% to the bill over thinking off and bought one more catch and a little more precision. The difference comes from token usage.\n\n*Figure 3. Cost of the Claude model calls per review at Anthropic’s list prices, Signal and OSS August.*\n\n*Figure 4. Average input and output tokens per core review call on Signal (22 calls per run).*\n\n| Avg per core review call | Signal: input | Signal: output | Signal: thinking words | OSS: input | OSS: output | OSS: thinking words | \n| Sonnet 5.5, thinking on | 110.7k | 5.8k | 464 | 87.3k | 5.7k | 523 | \n| Sonnet 5.5, thinking off | 110.7k | 2.9k | 0 | — | — | — | \n| Sonnet 5 | 247.5k | 21.6k | 2,771 | 191.5k | 23.8k | 3,143 | \n\nSignal has 22 core review calls per run, OSS August 84. Input is gross prompt size (uncached tokens plus cache reads and writes), standardized across providers.\n\nThe two benchmarks agree. Per review call, Sonnet 5 reads more than twice as much as Sonnet 5.5, writes about four times as much, and thinks roughly six times as many words, on both sets. On Signal it also caught fewer bugs. Whatever Anthropic changed between the two releases, the new model reaches its conclusions with far less deliberation.\n\nAcross the whole pipeline, including the smaller models that handle summaries and verification, Sonnet 5 consumed 27% more total tokens than Sonnet 5.5 on Signal and 49% more on the 44 OSS reviews, and more than twice the output tokens on both.\n\nLatency is where Sonnet 5.5 separates most clearly from its predecessor, and the larger run makes the point more firmly than the small one. On Signal, mean time per full review was 5:27 against 9:55 for Sonnet 5; on the 44 OSS reviews, 6:33 against 13:31. Over both sets, Sonnet 5 needed just over twelve hours of wall-clock time for 57 reviews; Sonnet 5.5 needed six.\n\n*Figure 5. Mean and median wall-clock time per full review on Signal.*\n\n## Sonnet 5.5 against the Sonnet line and Opus 5.5\n\nOur test sets, judges, and pipeline versions change between evaluations, so the numbers below are two different kinds of comparison. The Signal rows are a same-set comparison: identical 13 cases and identical recorded inputs, run within days of each other. The earlier Sonnet rows are historical context from our previous posts and should be read as direction, not as a leaderboard.\n\n| Model | Evaluation | Known issues caught | Actionable precision | Comments and noise | \n| [Sonnet 4.5](https://www.coderabbit.ai/blog/claude-sonnet-45-better-performance-but-a-paradox) | Oct 2025, 25 hard PRs | closed much of the gap to the flagship of the time | 35% comment-level | hedged in a third of its comments | \n| Sonnet 4.6 | Jun 2026, our standard benchmark of the time ( [Sonnet 5 review](https://www.coderabbit.ai/blog/claude-sonnet-5-review) ) | about 63% | about 29% | the noisiest Sonnet we measured | \n| [Sonnet 5](https://www.coderabbit.ai/blog/claude-sonnet-5-review) | Jun 2026, same set as 4.6 | about 50–51% | 38–40% | nitpick-heavy | \n| Sonnet 5 | Signal, this post | 4/13 · 30.8% | 40.0% | 15 comments, 3 nitpicks, 2 critical | \n| **Sonnet 5.5, thinking on** | **Signal, this post** | **6/13 · 46.2%** | **41.2%** | **17 comments, 2 nitpicks, 0 critical** | \n| [Opus 5.5 Standard](https://www.coderabbit.ai/blog/opus-5-5-model-review) | Signal, Sep 2026 | 8/13 · 61.5% | 66.7% | 21 comments | \n| [Opus 5.5 Max](https://www.coderabbit.ai/blog/opus-5-5-model-review) | Signal, Sep 2026 | 10/13 · 76.9% | 52.0% | 25 comments | \n| [Opus 5.5 Standard](https://www.coderabbit.ai/blog/opus-5-5-model-review) | OSS August, 80 cases, Sep 2026 | 63.8% | 38.6% | 127 comments | \n\nThree things stand out. First, within the Sonnet line, 5.5 is the first release to improve coverage without losing precision: 4.6 had coverage but not precision, and 5 had precision but not coverage. Second, on the same 13 Signal cases, Opus 5.5 caught 8 to 10 issues where Sonnet 5.5 caught 6, at higher precision. On the hardest cases the flagship is clearly the stronger reviewer, and Sonnet 5.5 does not close that gap. That is a wider margin than Anthropic’s launch benchmarks show, where Sonnet 5.5 lands within about three points of Opus 5.5 on most rows, but it fits Anthropic’s own framing that Opus 5.5 stays clearly stronger on complex, open-ended work that needs sustained judgment. Hard review cases are exactly that kind of work. Third, Sonnet 5.5 gets its result cheaply: it finishes reviews in about half the time of Sonnet 5, at half of Opus 5.5’s list price and 40% of Sonnet 5’s actual cost per review.\n\nThose are different roles. If missing a bug on a high-risk change is the expensive failure, Opus 5.5 is worth its extra tokens. If you are choosing the model for the review pass that runs on every pull request, fast and with comments developers will actually read, Sonnet 5.5 is the first Sonnet we would put on that list.\n\n### How it builds: a side-by-side with Opus 5.5\n\nWe also gave Sonnet 5.5 and Opus 5.5 the same long prompt in side-by-side Claude Code sessions: build “Brick Studio,” a brick-model designer with an Alpine Chalet demo model, and record a 45-second showcase video. Sonnet 5.5 finished first, in 29 minutes 27 seconds, against 44 minutes 50 seconds for Opus 5.5, so it was about 1.5× faster. The results were close to identical, with Opus 5.5 a little higher in fidelity. This is one run, not a benchmark. The sessions shared a working folder, so Sonnet 5.5 moved into an isolated subfolder partway through, and neither model received the reference screenshot. It does match the review results above: Opus 5.5 is the stronger model on the hardest work, and Sonnet 5.5 gets close in much less time.\n\n## How to evaluate Sonnet 5.5 yourself\n\nThirteen cases is enough to see direction, not to settle a decision. Before switching:\n\n- Run it on pull requests from your own repositories with known outcomes, and compare which bugs it catches and misses against the model you run today. Sonnet 5.5 and Sonnet 5 disagreed on six of our 13 cases; our two Sonnet 5.5 configurations disagreed with each other on three.\n- Test thinking on and off. In our runs thinking on caught more, was more precise, and was not slower; it cost four more comments and about twice the output tokens per review call.\n- Measure the complete review, not the model call. Verification agents, retries, and summary models all add tokens and time. Track input, output, cache-read, and cache-write tokens separately.\n- Read the comments that pass and the ones that do not. Several of Sonnet 5.5’s misses were plausible findings on the right file, which is harder to triage than noise.\n\n## Our verdict\n\nSonnet 5.5 fixes the specific weakness that kept Sonnet 5 off our main review path. On our hardest cases it catches half again as many known issues as Sonnet 5 at the same precision, and on 44 broader pull requests it posts a quarter fewer comments, a third of the nitpicks, and finishes in less than half the time, at about 40% of the cost.\n\nIt is not an Opus 5.5 replacement, four of our 13 hard cases defeated every Sonnet configuration we ran, and the coverage numbers for the larger set are still to come. But for the review that has to run on every pull request, fast and with comments developers will actually read, Sonnet 5.5 is the most convincing Sonnet we have tested. As our VP of AI, David Loker, put it in [Anthropic’s launch announcement](https://www.anthropic.com/claude-sonnet-5-5), Sonnet 5.5 “shows better judgment than Sonnet 5 across different levels of complexity” while spending far fewer output tokens, and Sonnet 5’s habit of reaching for web search too often is gone. We are moving simple and moderate reviews over now, and more in the coming weeks.\n\n[**Get started with CodeRabbit**](https://coderabbit.link/NcmblIe): connect your repo, get your first review in minutes. Free to try, no credit card required.\n\n*Methodology notes.* Signal is a 13-case subset of harder known-bug patterns drawn from real open-source pull requests; OSS August is 85 known issues across 44 open-source pull requests, for which only comment counts, latency, token usage, and cost are reported here because judge scoring was pending. A case counts as caught when at least one regular actionable comment is accepted by a majority of three judge votes. Precision measures the share of actionable comments that pass the judge for the target issue, not developer acceptance. All runs replayed identical upstream summaries from a recorded cassette. Individual judge calls matter at this sample size: on the Puma case, a finding from Sonnet 5 was accepted while a near-identical finding from Sonnet 5.5 was not, and one of Sonnet 5.5’s seven passing comments passed on a two-to-one vote; excluding it gives 35.3% precision. Token figures are averaged across generation calls with input standardized as uncached tokens plus cache reads and writes. Dollar figures are recomputed from each run’s token counts at Anthropic’s published list prices for Sonnet 5.5, which are unchanged from Sonnet 5, and cover the Claude model calls only; the evaluation harness itself used placeholder rates for the pre-release model.", "url": "https://wpnews.pro/news/claude-sonnet-5-5-for-code-review-more-catches-than-sonnet-5-in-half-the-time", "canonical_source": "https://coderabbit.ai/blog/sonnet-5-5-model-review", "published_at": "2026-09-28 00:00:00+00:00", "updated_at": "2026-09-28 21:47:32.857690+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-tools", "developer-tools", "ai-products"], "entities": ["Anthropic", "Claude Sonnet 5.5", "Claude Sonnet 5", "Claude Opus 5.5", "CodeRabbit", "Terminal-Bench 4.0", "AA-Briefcase"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/claude-sonnet-5-5-for-code-review-more-catches-than-sonnet-5-in-half-the-time", "markdown": "https://wpnews.pro/news/claude-sonnet-5-5-for-code-review-more-catches-than-sonnet-5-in-half-the-time.md", "text": "https://wpnews.pro/news/claude-sonnet-5-5-for-code-review-more-catches-than-sonnet-5-in-half-the-time.txt", "jsonld": "https://wpnews.pro/news/claude-sonnet-5-5-for-code-review-more-catches-than-sonnet-5-in-half-the-time.jsonld"}}