Anthropic SDK max_retries Stacked on My Retries: 36 Calls per Job A developer running the interview-practice tool Preterview traced a retry storm to the Anthropic Python SDK's default max_retries=2 stacking with a tenacity retry decorator and RQ job retries, multiplying one failing call into 36 HTTP attempts and up to 42 requests per job. During a burst of 529 overloaded_error responses, the piled-up background retries fired simultaneously on recovery, pushing the shared API key past its input-tokens-per-minute limit and causing 429s on live interview traffic. The fix is to set max_retries explicitly per client or call and to count actual HTTP requests with an httpx event hook rather than application-level function calls. My application logs said the report worker called Claude 1,302 times in 22 minutes. An httpx hook I added the next morning said it was 3,904 HTTP requests. Both were right, which was the problem. That gap came from the Anthropic SDK max retries default sitting underneath two other retry layers I had written myself. Each layer looked reasonable on its own. Together they multiplied: 3 × 4 × 3 = 36 attempts for one failing call. During a short stretch of 529 overloaded error responses, that multiplication turned a blip into a retry storm. When the API recovered, the storm hit my own rate limit and took down live traffic. messages.create in tenacity, and your job queue also retries, the attempts The official Python SDK anthropic retries a failed request twice by default, so one client.messages.create can become 3 HTTP requests. It retries connection errors, 408 Request Timeout, 409 Conflict, 429 Rate Limit and any 5xx, which includes the 529 overloaded error. It uses exponential backoff with jitter and respects the retry-after header when the server sends one. This is a good default. It's also invisible. Nothing in your code says "retry," so you forget it exists, and then you write your own retry on top. You can change it per client or per call: python import anthropic client = anthropic.Anthropic max retries=0 whole client client.with options max retries=5 .messages.create ... one call The system is Preterview, an interview practice tool I built and run full disclosure: I built it . A user does a realistic voice interview, and when the session ends, a background job sends the transcript to Claude to score it and write a report. You can see it at preterview.com/en https://preterview.com/en . The live interview turns and the background report jobs shared one API key, which matters later. Here is roughly what the report path looked like: python import anthropic from tenacity import retry, stop after attempt, wait fixed from rq import Retry client = anthropic.Anthropic max retries=2, I never typed it @retry stop=stop after attempt 4 , wait=wait fixed 2 def call claude kw : return client.messages.create kw def build report session id : rubric = call claude ... step 1: score against rubric evidence = call claude ... step 2: pull quotes per criterion report = call claude ... step 3: write the report save session id, report queue.enqueue build report, session id, retry=Retry max=2 Three layers, three people's worth of good intentions: | Layer | Attempts | Who added it | |---|---|---| | Anthropic SDK | 3 | The SDK, silently | | tenacity | 4 | Me, "just in case" | | RQ job retry | 3 | Me, months earlier | | Product | 36 | Nobody | And there was a fourth multiplier hiding in build report . An RQ retry restarts the whole function. If step 3 failed, the retry re-ran steps 1 and 2, which had already succeeded. My worst job that evening made 42 requests and still failed. Background retries piled up during the outage and all fired at once when the API recovered, which pushed my input tokens per minute past my rate limit. The live interview turns, which used the same key, started getting 429s. The timeline, from my logs: wait fixed 2 in tenacity has no jitter, so workers that failed together retried together. The numbers for the report worker in that window: call claude according to my app logs Nothing here was a bug in the SDK. The SDK did exactly what it documents. I just never did the multiplication. Pass your own httpx client with an event hook. Your application logs count function calls; this counts what actually goes over the wire, including the SDK's internal retries. python import httpx, anthropic def on request request : metrics.incr "anthropic.http requests", tags={"path": request.url.path} client = anthropic.Anthropic http client=anthropic.DefaultHttpxClient event hooks={"request": on request } Put the ratio of HTTP requests to app-level calls on a dashboard. Mine normally sits around 1.01. During the incident it hit 3.0, which is the SDK's ceiling and a clear sign every call was exhausting its retries. Decide which layer owns which timescale, then write the product of all layers in a comment next to the code. Here is what I changed. 1. The SDK owns fast HTTP retries. Nobody else does. I deleted tenacity around API calls entirely. The SDK already has jittered exponential backoff and reads retry-after , which my wait fixed 2 ignored. 2. Separate clients per traffic class. Live voice turns: latency beats persistence. One retry, then filler. live = anthropic.Anthropic max retries=1, timeout=20.0 Background reports: patient, but bounded. batch = anthropic.Anthropic max retries=3 3. Outer retries are slow and checkpointed. The job retry now waits minutes, not seconds, and each step's result is stored so a retry resumes instead of restarting. python def step session id, name, kw : if hit := store.get session id, name : return hit if breaker.is open : raise CoolingDown with input token bucket.reserve estimate tokens kw : out = batch.messages.create kw store.put session id, name, out return out Retry product: 4 SDK x 3 RQ = 12 attempts max, and the RQ attempts are 3 and 10 minutes apart. queue.enqueue build report, session id, retry=Retry max=2, interval= 180, 600 4. A circuit breaker on 529s. If more than 30% of the last 20 background calls got 529, workers stop pulling report jobs for 60 seconds. Reports wait in the queue instead of hammering a struggling API. 5. Background traffic gets a token budget. A simple token bucket caps the report worker at half of my input-tokens-per-minute limit. The recovery burst can still happen, but it can't eat the half that live turns need. A week later there was another overload stretch, about 15 minutes. The report worker sent 1.3× its normal request volume instead of 9.3×, the HTTP-to-call ratio peaked at 2.4, and live turns saw zero 429s. The slowest report arrived 11 minutes late, which is worse than 40 seconds but much better than a failed report with a 42-request bill attached. The honest caveat: I have two incidents, not a controlled experiment. The second outage was shorter and may have been milder. What I can say for certain is the arithmetic: the old setup could reach 36 attempts per failing call by design, and the new one is capped at 12, spread over 13 minutes. @retry , Retry , backoff , and max retries in any code path that touches an LLM client. Multiply what you find. max retries=0 on that client. Pick one. By itself, very little: the Anthropic SDK retries each request 2 times by default with jittered backoff, which is a sensible setting. The cost appears when you stack it under your own retry decorator and a job-queue retry, because retry layers multiply. My 3 × 4 × 3 setup allowed 36 attempts per failing call, produced 3,904 requests where 420 were needed, and the recovery burst caused 429s on 18% of live turns. Let the SDK own fast HTTP retries, make outer retries slow and checkpointed, write the retry product in a comment, and budget background traffic so it can't starve the requests your users are waiting on. Written by the developer behind Preterview https://preterview.com/en , an interview prep platform.