# Show HN: One line so LLM workers share a 429 wait and keep jobs an outage stops

> Source: <https://github.com/baldurhq/baldur>
> Published: 2026-10-02 16:40:24+00:00

**English** | [한국어](https://github.com/baldurhq/baldur/blob/main/README.ko.md)

**Early access — feedback wanted.** Trying Baldur on a real service? If anything gets in your way — installing, the docs, behavior you didn't expect — tell us in [Discussions](https://github.com/baldurhq/baldur/discussions) or [open an issue](https://github.com/baldurhq/baldur/issues/new/choose).

**An API you depend on goes down for an hour. What happens to your app?**

Requests hang until they time out, every worker fills up, and the jobs that failed in that hour are gone. Whether it's OpenAI, your payment provider, or your email service — Baldur fixes all three with one decorator, for Python services that don't have anyone on call.

``` python
import baldur
from openai import OpenAI

llm = baldur.llm.wrap(OpenAI(), timeout=60.0)

@baldur.protected("summarize", replay=True)
def summarize(doc_id: str) -> str:
    response = llm.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": load_document(doc_id)}],
    )
    return response.choices[0].message.content
```

No Redis, no Docker, no config to start: those two lines run in-memory until
you go multi-process. Not calling an LLM? The decorator alone does the same for
[any dependency](#the-same-decorator-any-dependency).

When the provider rate-limits you, dies, or just gets slow — mid-traffic:

- **Your workers back off together.** A 429 or an "overloaded" answer installs
one wait every worker shares, at least as long as the provider asked: the
first worker is refused and the rest wait, instead of each one discovering
the limit alone. A request the provider rejected (400, 422) is never retried.
- **Your app keeps answering.** A hung request fails at the 60-second bound,
the circuit breaker opens, and calls fail fast — so a slow provider doesn't
take every worker down with it. Give the wrap`fallbacks=[...]` and a call
moves to the next endpoint instead (OpenAI SDK, Anthropic SDK, google-genai).
- **Failed jobs are kept, not lost.** Every call that failed for good is
captured with its arguments and listed in the built-in console at`http://127.0.0.1:9090/` . In a plain Python process, call`baldur.init()` once at startup to start the console; the Django, FastAPI, and Flask
integrations do that for you.
- **They come back.**`replay=True` lets Baldur re-run a parked job from its
stored arguments: from the console with a click, or automatically once the
provider recovers and the job's breaker closes — with a Celery worker running
the replay.

See it against the real `openai` SDK and a local fake provider — a rate limit,
then an outage, then every parked job replayed:

*A real run. Playback shortens each pause to 3 seconds; the times on screen
are the run's own. Run it yourself:*

```
pip install "baldur-framework[celery]" openai
python -m baldur.scripts.demo_llm_outage
```

**If Baldur breaks, does my call break?** No. A fault in Baldur's own
bookkeeping, such as a failed dead-letter write, is logged: your call still
returns its own result or raises its own error. If Redis is unreachable, calls
keep running on each process's own state, and the workers stop sharing the 429
wait and the breaker until it is back. One exception, by design: a call with
`idempotency_key=` is refused with `IdempotencyUnavailableError` (after a few
seconds' wait) rather than run without the shared record that keeps it from
running twice.

**Can a retry or a replay bill me twice?** Yes, so plan for it. A replay runs
the whole job again: a model call that had already succeeded inside the job is
made, and billed, a second time. A request that timed out may have finished on
the provider's side, so retrying it can bill twice too. Use `replay=True` on
jobs that are safe to run twice, or store each step's result yourself and skip
the steps that already finished.

Django, FastAPI, Flask, and Celery adapters included.

**Already using your SDK's retries?** Around a decorated call, keep them.
Baldur doesn't replace retry — it adds what retry can't: a breaker so one
incident doesn't cost every request its retries, one wall-clock bound on what
the caller waits, a fallback, and the capture-and-replay no retry library gives
you. A `baldur.llm.wrap` client is the one exception: it turns the SDK's own
retries off on its copy of the client, because there one coordinated retry
replaces them.

The Python package is `baldur` (you `import baldur`); the PyPI distribution is
`baldur-framework`.

```
pip install baldur-framework                 # framework-agnostic core
pip install baldur-framework[django]         # Django integration
pip install baldur-framework[django-api]     # Baldur's Django REST API (baldur.api.django.urls)
pip install baldur-framework[fastapi]        # FastAPI integration
pip install baldur-framework[flask]          # Flask integration
pip install baldur-framework[celery]         # Celery task protection
pip install baldur-framework[redis]          # Redis-backed shared state
pip install baldur-framework[prometheus]     # Prometheus metrics
```

A payment gateway, your database, an email provider — the call site never changes:

``` python
@baldur.protected("charge-customer", dlq=True)
def charge(order_id: str, amount_cents: int) -> dict:
    # Circuit breaker by default; dlq=True parks the call if it fails for
    # good, with its arguments, and replays it once the gateway recovers.
    return payment_gateway.charge(order_id, amount_cents)
```

When the gateway dies, the breaker opens and your service answers fast instead
of stacking up timeouts; the charges that failed on the way out wait in the
dead-letter queue and come back when it closes. (Replay is for work that failed
on the way out — never for a business rejection, and never for a checkout the
customer already walked away from:
[where that line sits](https://baldur.sh/concepts/foundations/dlq-replay/).)

*The payment demo: the gateway goes unreachable mid-traffic, seven charges are
captured with their arguments, and all seven are replayed on recovery. Zero
lost. A real run, with the breaker states and DLQ tallies read live from the
framework. The decorator itself is `pip install baldur-framework` and nothing
else; the demo adds the `celery` extra for its in-process stand-in worker —
still one process, no Redis, no broker. Run it yourself:*

```
pip install "baldur-framework[celery]"
python -m baldur.scripts.demo_self_healing
```

Need more than the default? Compose the pipeline declaratively:

```
@baldur.protected(
    "summarize",
    timeout=30.0,                            # one bound on what the caller waits
    fallback=lambda: last_good_summary(),    # graceful answer while OPEN
    idempotency_key="doc_id",                # a redelivered job pays once
)
def summarize(doc_id: str) -> str:
    return llm_api.summarize(doc_id)
```

**Notice what isn't there: `retry=`.** Your SDK almost certainly retries
already — `anthropic` and `openai` default to two attempts with backoff, boto3
has an adaptive mode — and it retries better than a generic wrapper can,
because it knows which status codes are worth another attempt and honours
`retry-after`. Keep it. What no SDK gives you is the rest: a breaker, so a
provider incident doesn't mean *every* request pays its retries before failing;
one wall-clock bound on what your caller waits, retries included (an SDK's own
worst case is `timeout × (max_retries + 1)` — 30 minutes at `anthropic`'s
defaults); a fallback; and a dedup key that survives a job redelivery the SDK
never sees. `retry=True` is there for the calls that don't retry themselves.

Sync and async callables are both supported — the decorator auto-detects coroutine functions.

| Capability | What it gives you | 
|---|---|
| [Circuit breaker](https://baldur.sh/concepts/oss/circuit-breaker/) | Stops cascading failure; bounded half-open probes on recovery | 
| [Retry with backoff](https://baldur.sh/concepts/oss/retry/) | Exponential backoff with jitter and bounded attempts | 
| [Fallback & composition](https://baldur.sh/concepts/foundations/composition/) | One ordered pipeline for all resilience patterns | 
| [Idempotency](https://baldur.sh/concepts/oss/idempotency/) | Concurrent duplicate calls execute the side effect exactly once | 
| [Bulkhead isolation](https://baldur.sh/concepts/foundations/bulkhead/) | Each dependency gets a fixed slice of concurrency, so one slow dependency can't drain every worker | 
| [Dead-letter queue + replay](https://baldur.sh/concepts/foundations/dlq-replay/) | A call that fails for good is captured with its context and replayed once the dependency recovers | 
| [Health checks](https://baldur.sh/concepts/oss/health-check/) | Liveness/readiness that reflect real dependency state | 
| [Graceful shutdown](https://baldur.sh/concepts/oss/graceful-shutdown/) | Drain in-flight work cleanly on restart and deploy | 
| [Metrics](https://baldur.sh/concepts/oss/metrics/) | Prometheus and OpenTelemetry, emitted by default | 
| [System control](https://baldur.sh/concepts/oss/system-control/) | Instant kill switch and dry-run mode for Baldur's automation — no redeploy | 
| [Web console](https://baldur.sh/concepts/foundations/web-console/) | Built-in operations console: live breaker state, controls, recovery | 
| [Precomputed cache](https://baldur.sh/concepts/oss/precomputed-cache/) | Health/status endpoints answer from a warm cache, so constant probing stays cheap | 

The read path heals the same way. Here a Django app under live HTTP traffic (recorded from a demo harness driving it) loses its network path to Redis for 21 seconds — every request keeps returning 200 off the in-memory cache tier, and the Redis tier resyncs itself on recovery:

Full documentation lives at **[https://baldur.sh](https://baldur.sh)**.

- [What is Baldur?](https://baldur.sh/what-is-baldur/) — the problem it solves and how
- Getting started: [Django](https://baldur.sh/getting-started/django/) ·[FastAPI](https://baldur.sh/getting-started/fastapi/) ·[Flask](https://baldur.sh/getting-started/flask/) ·[Celery](https://baldur.sh/getting-started/celery/)
- [Concept guides](https://baldur.sh) — one page per capability, linked
throughout this README
- [API reference](https://baldur.sh/reference/)
- [Troubleshooting](https://baldur.sh/troubleshooting/)
- [Compatibility](https://baldur.sh/compatibility/)

Building with an AI coding assistant (Claude Code, Cursor, Copilot, Codex)? Run
`baldur init-ai` in your repo to drop an `AGENTS.md` (read by Cursor, Copilot,
and Codex) plus a `CLAUDE.md` that imports it for Claude Code — together they
teach the assistant to reach for `@baldur.protected("name")` instead of
hand-rolling a circuit breaker. See
[Using Baldur with AI assistants](https://baldur.sh/getting-started/ai-assistants/).

| Component | Minimum | Tested in CI | 
|---|---|---|
| Python | 3.11 | 3.11 · 3.12 · 3.13 | 
| Django | 4.2 | 4.2 LTS · 5.2 LTS · 6.0 | 
| FastAPI | 0.100 | latest ≥ floor (smoke) | 
| Flask | 2.3 | latest ≥ floor (smoke) | 
| Celery | 5.3 | 5.4 | 
| Redis server | — | 7.x | 

See [Compatibility](https://baldur.sh/compatibility/) for the full matrix, the
Python × Django test grid, and the version support policy.

Baldur PRO adds the fleet-level machinery on top of the same API — nothing in
the core gets relicensed or replaced:
[DLQ at scale](https://baldur.sh/concepts/foundations/dlq-replay/) (batch replay from the
console, success-rate-driven pacing, and archive/purge
retention), a hash-chained [audit trail](https://baldur.sh/concepts/pro/audit/),
[unified notifications](https://baldur.sh/concepts/pro/unified-notification/),
[emergency mode](https://baldur.sh/concepts/pro/emergency-mode/),
[bulkhead thread-pool isolation](https://baldur.sh/concepts/foundations/bulkhead/),
[adaptive throttling](https://baldur.sh/concepts/pro/throttle/),
[canary recovery](https://baldur.sh/concepts/pro/canary-recovery/),
[governance gates](https://baldur.sh/concepts/pro/governance/), and a
[meta-watchdog](https://baldur.sh/concepts/pro/meta-watchdog/) that watches Baldur itself.
See the full [OSS vs PRO capability matrix](https://baldur.sh/concepts/oss-vs-pro/) and
[pricing](https://baldur.sh/pricing/).

Baldur is in early access: the API is stable and the core is tested under
sustained load with Sentinel failover, but the project is young — minor
releases may still ship breaking changes, always with a changelog entry. It is
looking for a small number of teams already running a Python service in
production to work with directly. If that is you, the details and how to reach
me are in [Discussions](https://github.com/baldurhq/baldur/discussions).

How the project got here — including why it was nearly shelved in September
2026: [retrospective (Korean)](https://github.com/baldurhq/baldur/blob/main/POSTMORTEM.ko.md).

Baldur is released under the Apache License 2.0 — see [LICENSE](https://github.com/baldurhq/baldur/blob/main/LICENSE) and
[NOTICE](https://github.com/baldurhq/baldur/blob/main/NOTICE).

Contributions are welcome under the Apache License 2.0. Pull requests are
accepted through a sign-off-based [DCO](https://developercertificate.org/) flow —
see [CONTRIBUTING.md](https://github.com/baldurhq/baldur/blob/main/CONTRIBUTING.md) for the full model.

- **Ideas, or showing what you built** →[Discussions](https://github.com/baldurhq/baldur/discussions) .
- **Bugs / feature requests / docs** → open an issue or a pull request.
- **Security** → see[SECURITY.md](https://github.com/baldurhq/baldur/blob/main/SECURITY.md) (no public issues for vulnerabilities).
- **Usage questions / commercial** →`support@baldur.sh` .
