# Anthropic SDK max_retries Stacked on My Retries: 36 Calls per Job

> Source: <https://dev.to/ji_ai/anthropic-sdk-maxretries-stacked-on-my-retries-36-calls-per-job-283l>
> Published: 2026-10-06 16:40:44+00:00

My application logs said the report worker called Claude 1,302 times in 22 minutes. An httpx hook I added the next morning said it was 3,904 HTTP requests. Both were right, which was the problem.

That gap came from the **Anthropic SDK max_retries** default sitting underneath two other retry layers I had written myself. Each layer looked reasonable on its own. Together they multiplied: 3 × 4 × 3 = 36 attempts for one failing call. During a short stretch of `529 overloaded_error` responses, that multiplication turned a blip into a retry storm. When the API recovered, the storm hit my own rate limit and took down live traffic.

`messages.create()` in tenacity, and your job queue also retries, the attempts The official Python SDK (`anthropic`) retries a failed request twice by default, so one `client.messages.create()` can become 3 HTTP requests. It retries connection errors, 408 Request Timeout, 409 Conflict, 429 Rate Limit and any 5xx, which includes the 529 overloaded error. It uses exponential backoff with jitter and respects the `retry-after` header when the server sends one.

This is a good default. It's also invisible. Nothing in your code says "retry," so you forget it exists, and then you write your own retry on top.

You can change it per client or per call:

``` python
import anthropic

client = anthropic.Anthropic(max_retries=0)          # whole client
client.with_options(max_retries=5).messages.create(...)  # one call
```

The system is Preterview, an interview practice tool I built and run (full disclosure: I built it). A user does a realistic voice interview, and when the session ends, a background job sends the transcript to Claude to score it and write a report. You can see it at [preterview.com/en](https://preterview.com/en). The live interview turns and the background report jobs shared one API key, which matters later.

Here is roughly what the report path looked like:

``` python
import anthropic
from tenacity import retry, stop_after_attempt, wait_fixed
from rq import Retry

client = anthropic.Anthropic()  # max_retries=2, I never typed it

@retry(stop=stop_after_attempt(4), wait=wait_fixed(2))
def call_claude(**kw):
    return client.messages.create(**kw)

def build_report(session_id):
    rubric = call_claude(...)    # step 1: score against rubric
    evidence = call_claude(...)  # step 2: pull quotes per criterion
    report = call_claude(...)    # step 3: write the report
    save(session_id, report)

queue.enqueue(build_report, session_id, retry=Retry(max=2))
```

Three layers, three people's worth of good intentions:

| Layer | Attempts | Who added it | 
|---|---|---|
| Anthropic SDK | 3 | The SDK, silently | 
| tenacity | 4 | Me, "just in case" | 
| RQ job retry | 3 | Me, months earlier | 
| **Product** | **36** | Nobody | 

And there was a fourth multiplier hiding in `build_report`. An RQ retry restarts the whole function. If step 3 failed, the retry re-ran steps 1 and 2, which had already succeeded. My worst job that evening made **42 requests** and still failed.

Background retries piled up during the outage and all fired at once when the API recovered, which pushed my input tokens per minute past my rate limit. The live interview turns, which used the same key, started getting 429s.

The timeline, from my logs:

`wait_fixed(2)` in tenacity has no jitter, so workers that failed together retried together.
The numbers for the report worker in that window:

`call_claude()` according to my app logs
Nothing here was a bug in the SDK. The SDK did exactly what it documents. I just never did the multiplication.

Pass your own httpx client with an event hook. Your application logs count function calls; this counts what actually goes over the wire, including the SDK's internal retries.

``` python
import httpx, anthropic

def on_request(request):
    metrics.incr("anthropic.http_requests", tags={"path": request.url.path})

client = anthropic.Anthropic(
    http_client=anthropic.DefaultHttpxClient(
        event_hooks={"request": [on_request]}
    )
)
```

Put the ratio of HTTP requests to app-level calls on a dashboard. Mine normally sits around 1.01. During the incident it hit 3.0, which is the SDK's ceiling and a clear sign every call was exhausting its retries.

Decide which layer owns which timescale, then write the product of all layers in a comment next to the code. Here is what I changed.

**1. The SDK owns fast HTTP retries. Nobody else does.** I deleted tenacity around API calls entirely. The SDK already has jittered exponential backoff and reads `retry-after`, which my `wait_fixed(2)` ignored.

**2. Separate clients per traffic class.**

```
# Live voice turns: latency beats persistence. One retry, then filler.
live = anthropic.Anthropic(max_retries=1, timeout=20.0)

# Background reports: patient, but bounded.
batch = anthropic.Anthropic(max_retries=3)
```

**3. Outer retries are slow and checkpointed.** The job retry now waits minutes, not seconds, and each step's result is stored so a retry resumes instead of restarting.

``` python
def step(session_id, name, **kw):
    if (hit := store.get(session_id, name)):
        return hit
    if breaker.is_open():
        raise CoolingDown()
    with input_token_bucket.reserve(estimate_tokens(kw)):
        out = batch.messages.create(**kw)
    store.put(session_id, name, out)
    return out

# Retry product: 4 (SDK) x 3 (RQ) = 12 attempts max,
# and the RQ attempts are 3 and 10 minutes apart.
queue.enqueue(build_report, session_id,
              retry=Retry(max=2, interval=[180, 600]))
```

**4. A circuit breaker on 529s.** If more than 30% of the last 20 background calls got 529, workers stop pulling report jobs for 60 seconds. Reports wait in the queue instead of hammering a struggling API.

**5. Background traffic gets a token budget.** A simple token bucket caps the report worker at half of my input-tokens-per-minute limit. The recovery burst can still happen, but it can't eat the half that live turns need.

A week later there was another overload stretch, about 15 minutes. The report worker sent **1.3×** its normal request volume instead of 9.3×, the HTTP-to-call ratio peaked at 2.4, and live turns saw **zero** 429s. The slowest report arrived 11 minutes late, which is worse than 40 seconds but much better than a failed report with a 42-request bill attached.

The honest caveat: I have two incidents, not a controlled experiment. The second outage was shorter and may have been milder. What I can say for certain is the arithmetic: the old setup could reach 36 attempts per failing call by design, and the new one is capped at 12, spread over 13 minutes.

`@retry`, `Retry(`, `backoff`, and `max_retries` in any code path that touches an LLM client. Multiply what you find.`max_retries=0` on that client. Pick one.
By itself, very little: the Anthropic SDK retries each request 2 times by default with jittered backoff, which is a sensible setting. The cost appears when you stack it under your own retry decorator and a job-queue retry, because retry layers multiply. My 3 × 4 × 3 setup allowed 36 attempts per failing call, produced 3,904 requests where 420 were needed, and the recovery burst caused 429s on 18% of live turns. Let the SDK own fast HTTP retries, make outer retries slow and checkpointed, write the retry product in a comment, and budget background traffic so it can't starve the requests your users are waiting on.

*Written by the developer behind [Preterview](https://preterview.com/en), an interview prep platform.*
