{"slug": "anthropic-sdk-max-retries-stacked-on-my-retries-36-calls-per-job", "title": "Anthropic SDK max_retries Stacked on My Retries: 36 Calls per Job", "summary": "A developer running the interview-practice tool Preterview traced a retry storm to the Anthropic Python SDK's default max_retries=2 stacking with a tenacity retry decorator and RQ job retries, multiplying one failing call into 36 HTTP attempts and up to 42 requests per job. During a burst of 529 overloaded_error responses, the piled-up background retries fired simultaneously on recovery, pushing the shared API key past its input-tokens-per-minute limit and causing 429s on live interview traffic. The fix is to set max_retries explicitly per client or call and to count actual HTTP requests with an httpx event hook rather than application-level function calls.", "body_md": "My application logs said the report worker called Claude 1,302 times in 22 minutes. An httpx hook I added the next morning said it was 3,904 HTTP requests. Both were right, which was the problem.\n\nThat gap came from the **Anthropic SDK max_retries** default sitting underneath two other retry layers I had written myself. Each layer looked reasonable on its own. Together they multiplied: 3 × 4 × 3 = 36 attempts for one failing call. During a short stretch of `529 overloaded_error` responses, that multiplication turned a blip into a retry storm. When the API recovered, the storm hit my own rate limit and took down live traffic.\n\n`messages.create()` in tenacity, and your job queue also retries, the attempts The official Python SDK (`anthropic`) retries a failed request twice by default, so one `client.messages.create()` can become 3 HTTP requests. It retries connection errors, 408 Request Timeout, 409 Conflict, 429 Rate Limit and any 5xx, which includes the 529 overloaded error. It uses exponential backoff with jitter and respects the `retry-after` header when the server sends one.\n\nThis is a good default. It's also invisible. Nothing in your code says \"retry,\" so you forget it exists, and then you write your own retry on top.\n\nYou can change it per client or per call:\n\n``` python\nimport anthropic\n\nclient = anthropic.Anthropic(max_retries=0)          # whole client\nclient.with_options(max_retries=5).messages.create(...)  # one call\n```\n\nThe system is Preterview, an interview practice tool I built and run (full disclosure: I built it). A user does a realistic voice interview, and when the session ends, a background job sends the transcript to Claude to score it and write a report. You can see it at [preterview.com/en](https://preterview.com/en). The live interview turns and the background report jobs shared one API key, which matters later.\n\nHere is roughly what the report path looked like:\n\n``` python\nimport anthropic\nfrom tenacity import retry, stop_after_attempt, wait_fixed\nfrom rq import Retry\n\nclient = anthropic.Anthropic()  # max_retries=2, I never typed it\n\n@retry(stop=stop_after_attempt(4), wait=wait_fixed(2))\ndef call_claude(**kw):\n    return client.messages.create(**kw)\n\ndef build_report(session_id):\n    rubric = call_claude(...)    # step 1: score against rubric\n    evidence = call_claude(...)  # step 2: pull quotes per criterion\n    report = call_claude(...)    # step 3: write the report\n    save(session_id, report)\n\nqueue.enqueue(build_report, session_id, retry=Retry(max=2))\n```\n\nThree layers, three people's worth of good intentions:\n\n| Layer | Attempts | Who added it | \n|---|---|---|\n| Anthropic SDK | 3 | The SDK, silently | \n| tenacity | 4 | Me, \"just in case\" | \n| RQ job retry | 3 | Me, months earlier | \n| **Product** | **36** | Nobody | \n\nAnd there was a fourth multiplier hiding in `build_report`. An RQ retry restarts the whole function. If step 3 failed, the retry re-ran steps 1 and 2, which had already succeeded. My worst job that evening made **42 requests** and still failed.\n\nBackground retries piled up during the outage and all fired at once when the API recovered, which pushed my input tokens per minute past my rate limit. The live interview turns, which used the same key, started getting 429s.\n\nThe timeline, from my logs:\n\n`wait_fixed(2)` in tenacity has no jitter, so workers that failed together retried together.\nThe numbers for the report worker in that window:\n\n`call_claude()` according to my app logs\nNothing here was a bug in the SDK. The SDK did exactly what it documents. I just never did the multiplication.\n\nPass your own httpx client with an event hook. Your application logs count function calls; this counts what actually goes over the wire, including the SDK's internal retries.\n\n``` python\nimport httpx, anthropic\n\ndef on_request(request):\n    metrics.incr(\"anthropic.http_requests\", tags={\"path\": request.url.path})\n\nclient = anthropic.Anthropic(\n    http_client=anthropic.DefaultHttpxClient(\n        event_hooks={\"request\": [on_request]}\n    )\n)\n```\n\nPut the ratio of HTTP requests to app-level calls on a dashboard. Mine normally sits around 1.01. During the incident it hit 3.0, which is the SDK's ceiling and a clear sign every call was exhausting its retries.\n\nDecide which layer owns which timescale, then write the product of all layers in a comment next to the code. Here is what I changed.\n\n**1. The SDK owns fast HTTP retries. Nobody else does.** I deleted tenacity around API calls entirely. The SDK already has jittered exponential backoff and reads `retry-after`, which my `wait_fixed(2)` ignored.\n\n**2. Separate clients per traffic class.**\n\n```\n# Live voice turns: latency beats persistence. One retry, then filler.\nlive = anthropic.Anthropic(max_retries=1, timeout=20.0)\n\n# Background reports: patient, but bounded.\nbatch = anthropic.Anthropic(max_retries=3)\n```\n\n**3. Outer retries are slow and checkpointed.** The job retry now waits minutes, not seconds, and each step's result is stored so a retry resumes instead of restarting.\n\n``` python\ndef step(session_id, name, **kw):\n    if (hit := store.get(session_id, name)):\n        return hit\n    if breaker.is_open():\n        raise CoolingDown()\n    with input_token_bucket.reserve(estimate_tokens(kw)):\n        out = batch.messages.create(**kw)\n    store.put(session_id, name, out)\n    return out\n\n# Retry product: 4 (SDK) x 3 (RQ) = 12 attempts max,\n# and the RQ attempts are 3 and 10 minutes apart.\nqueue.enqueue(build_report, session_id,\n              retry=Retry(max=2, interval=[180, 600]))\n```\n\n**4. A circuit breaker on 529s.** If more than 30% of the last 20 background calls got 529, workers stop pulling report jobs for 60 seconds. Reports wait in the queue instead of hammering a struggling API.\n\n**5. Background traffic gets a token budget.** A simple token bucket caps the report worker at half of my input-tokens-per-minute limit. The recovery burst can still happen, but it can't eat the half that live turns need.\n\nA week later there was another overload stretch, about 15 minutes. The report worker sent **1.3×** its normal request volume instead of 9.3×, the HTTP-to-call ratio peaked at 2.4, and live turns saw **zero** 429s. The slowest report arrived 11 minutes late, which is worse than 40 seconds but much better than a failed report with a 42-request bill attached.\n\nThe honest caveat: I have two incidents, not a controlled experiment. The second outage was shorter and may have been milder. What I can say for certain is the arithmetic: the old setup could reach 36 attempts per failing call by design, and the new one is capped at 12, spread over 13 minutes.\n\n`@retry`, `Retry(`, `backoff`, and `max_retries` in any code path that touches an LLM client. Multiply what you find.`max_retries=0` on that client. Pick one.\nBy itself, very little: the Anthropic SDK retries each request 2 times by default with jittered backoff, which is a sensible setting. The cost appears when you stack it under your own retry decorator and a job-queue retry, because retry layers multiply. My 3 × 4 × 3 setup allowed 36 attempts per failing call, produced 3,904 requests where 420 were needed, and the recovery burst caused 429s on 18% of live turns. Let the SDK own fast HTTP retries, make outer retries slow and checkpointed, write the retry product in a comment, and budget background traffic so it can't starve the requests your users are waiting on.\n\n*Written by the developer behind [Preterview](https://preterview.com/en), an interview prep platform.*", "url": "https://wpnews.pro/news/anthropic-sdk-max-retries-stacked-on-my-retries-36-calls-per-job", "canonical_source": "https://dev.to/ji_ai/anthropic-sdk-maxretries-stacked-on-my-retries-36-calls-per-job-283l", "published_at": "2026-10-06 16:40:44+00:00", "updated_at": "2026-10-06 16:48:45.114111+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "developer-tools", "mlops"], "entities": ["Anthropic", "Preterview", "Claude", "httpx", "tenacity", "RQ"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/anthropic-sdk-max-retries-stacked-on-my-retries-36-calls-per-job", "markdown": "https://wpnews.pro/news/anthropic-sdk-max-retries-stacked-on-my-retries-36-calls-per-job.md", "text": "https://wpnews.pro/news/anthropic-sdk-max-retries-stacked-on-my-retries-36-calls-per-job.txt", "jsonld": "https://wpnews.pro/news/anthropic-sdk-max-retries-stacked-on-my-retries-36-calls-per-job.jsonld"}}