Show HN: One line so LLM workers share a 429 wait and keep jobs an outage stops Baldur, an early-access Python framework from baldurhq, released a decorator that makes LLM workers share a single backoff wait on a 429 or "overloaded" response and parks failed jobs for replay instead of losing them. The framework wraps the OpenAI, Anthropic, and google-genai SDKs, runs in-memory with no Redis or Docker, opens a circuit breaker at a 60-second request bound, and lists dead-lettered calls with their arguments in a built-in console at http://127.0.0.1:9090/. Baldur warns that replay=True re-runs an entire job, so a model call that already succeeded inside it is billed a second time, and a timed-out request may also bill twice on retry. English | 한국어 https://github.com/baldurhq/baldur/blob/main/README.ko.md Early access — feedback wanted. Trying Baldur on a real service? If anything gets in your way — installing, the docs, behavior you didn't expect — tell us in Discussions https://github.com/baldurhq/baldur/discussions or open an issue https://github.com/baldurhq/baldur/issues/new/choose . An API you depend on goes down for an hour. What happens to your app? Requests hang until they time out, every worker fills up, and the jobs that failed in that hour are gone. Whether it's OpenAI, your payment provider, or your email service — Baldur fixes all three with one decorator, for Python services that don't have anyone on call. python import baldur from openai import OpenAI llm = baldur.llm.wrap OpenAI , timeout=60.0 @baldur.protected "summarize", replay=True def summarize doc id: str - str: response = llm.chat.completions.create model="gpt-4o-mini", messages= {"role": "user", "content": load document doc id } , return response.choices 0 .message.content No Redis, no Docker, no config to start: those two lines run in-memory until you go multi-process. Not calling an LLM? The decorator alone does the same for any dependency the-same-decorator-any-dependency . When the provider rate-limits you, dies, or just gets slow — mid-traffic: - Your workers back off together. A 429 or an "overloaded" answer installs one wait every worker shares, at least as long as the provider asked: the first worker is refused and the rest wait, instead of each one discovering the limit alone. A request the provider rejected 400, 422 is never retried. - Your app keeps answering. A hung request fails at the 60-second bound, the circuit breaker opens, and calls fail fast — so a slow provider doesn't take every worker down with it. Give the wrap fallbacks= ... and a call moves to the next endpoint instead OpenAI SDK, Anthropic SDK, google-genai . - Failed jobs are kept, not lost. Every call that failed for good is captured with its arguments and listed in the built-in console at http://127.0.0.1:9090/ . In a plain Python process, call baldur.init once at startup to start the console; the Django, FastAPI, and Flask integrations do that for you. - They come back. replay=True lets Baldur re-run a parked job from its stored arguments: from the console with a click, or automatically once the provider recovers and the job's breaker closes — with a Celery worker running the replay. See it against the real openai SDK and a local fake provider — a rate limit, then an outage, then every parked job replayed: A real run. Playback shortens each pause to 3 seconds; the times on screen are the run's own. Run it yourself: pip install "baldur-framework celery " openai python -m baldur.scripts.demo llm outage If Baldur breaks, does my call break? No. A fault in Baldur's own bookkeeping, such as a failed dead-letter write, is logged: your call still returns its own result or raises its own error. If Redis is unreachable, calls keep running on each process's own state, and the workers stop sharing the 429 wait and the breaker until it is back. One exception, by design: a call with idempotency key= is refused with IdempotencyUnavailableError after a few seconds' wait rather than run without the shared record that keeps it from running twice. Can a retry or a replay bill me twice? Yes, so plan for it. A replay runs the whole job again: a model call that had already succeeded inside the job is made, and billed, a second time. A request that timed out may have finished on the provider's side, so retrying it can bill twice too. Use replay=True on jobs that are safe to run twice, or store each step's result yourself and skip the steps that already finished. Django, FastAPI, Flask, and Celery adapters included. Already using your SDK's retries? Around a decorated call, keep them. Baldur doesn't replace retry — it adds what retry can't: a breaker so one incident doesn't cost every request its retries, one wall-clock bound on what the caller waits, a fallback, and the capture-and-replay no retry library gives you. A baldur.llm.wrap client is the one exception: it turns the SDK's own retries off on its copy of the client, because there one coordinated retry replaces them. The Python package is baldur you import baldur ; the PyPI distribution is baldur-framework . pip install baldur-framework framework-agnostic core pip install baldur-framework django Django integration pip install baldur-framework django-api Baldur's Django REST API baldur.api.django.urls pip install baldur-framework fastapi FastAPI integration pip install baldur-framework flask Flask integration pip install baldur-framework celery Celery task protection pip install baldur-framework redis Redis-backed shared state pip install baldur-framework prometheus Prometheus metrics A payment gateway, your database, an email provider — the call site never changes: python @baldur.protected "charge-customer", dlq=True def charge order id: str, amount cents: int - dict: Circuit breaker by default; dlq=True parks the call if it fails for good, with its arguments, and replays it once the gateway recovers. return payment gateway.charge order id, amount cents When the gateway dies, the breaker opens and your service answers fast instead of stacking up timeouts; the charges that failed on the way out wait in the dead-letter queue and come back when it closes. Replay is for work that failed on the way out — never for a business rejection, and never for a checkout the customer already walked away from: where that line sits https://baldur.sh/concepts/foundations/dlq-replay/ . The payment demo: the gateway goes unreachable mid-traffic, seven charges are captured with their arguments, and all seven are replayed on recovery. Zero lost. A real run, with the breaker states and DLQ tallies read live from the framework. The decorator itself is pip install baldur-framework and nothing else; the demo adds the celery extra for its in-process stand-in worker — still one process, no Redis, no broker. Run it yourself: pip install "baldur-framework celery " python -m baldur.scripts.demo self healing Need more than the default? Compose the pipeline declaratively: @baldur.protected "summarize", timeout=30.0, one bound on what the caller waits fallback=lambda: last good summary , graceful answer while OPEN idempotency key="doc id", a redelivered job pays once def summarize doc id: str - str: return llm api.summarize doc id Notice what isn't there: retry= . Your SDK almost certainly retries already — anthropic and openai default to two attempts with backoff, boto3 has an adaptive mode — and it retries better than a generic wrapper can, because it knows which status codes are worth another attempt and honours retry-after . Keep it. What no SDK gives you is the rest: a breaker, so a provider incident doesn't mean every request pays its retries before failing; one wall-clock bound on what your caller waits, retries included an SDK's own worst case is timeout × max retries + 1 — 30 minutes at anthropic 's defaults ; a fallback; and a dedup key that survives a job redelivery the SDK never sees. retry=True is there for the calls that don't retry themselves. Sync and async callables are both supported — the decorator auto-detects coroutine functions. | Capability | What it gives you | |---|---| | Circuit breaker https://baldur.sh/concepts/oss/circuit-breaker/ | Stops cascading failure; bounded half-open probes on recovery | | Retry with backoff https://baldur.sh/concepts/oss/retry/ | Exponential backoff with jitter and bounded attempts | | Fallback & composition https://baldur.sh/concepts/foundations/composition/ | One ordered pipeline for all resilience patterns | | Idempotency https://baldur.sh/concepts/oss/idempotency/ | Concurrent duplicate calls execute the side effect exactly once | | Bulkhead isolation https://baldur.sh/concepts/foundations/bulkhead/ | Each dependency gets a fixed slice of concurrency, so one slow dependency can't drain every worker | | Dead-letter queue + replay https://baldur.sh/concepts/foundations/dlq-replay/ | A call that fails for good is captured with its context and replayed once the dependency recovers | | Health checks https://baldur.sh/concepts/oss/health-check/ | Liveness/readiness that reflect real dependency state | | Graceful shutdown https://baldur.sh/concepts/oss/graceful-shutdown/ | Drain in-flight work cleanly on restart and deploy | | Metrics https://baldur.sh/concepts/oss/metrics/ | Prometheus and OpenTelemetry, emitted by default | | System control https://baldur.sh/concepts/oss/system-control/ | Instant kill switch and dry-run mode for Baldur's automation — no redeploy | | Web console https://baldur.sh/concepts/foundations/web-console/ | Built-in operations console: live breaker state, controls, recovery | | Precomputed cache https://baldur.sh/concepts/oss/precomputed-cache/ | Health/status endpoints answer from a warm cache, so constant probing stays cheap | The read path heals the same way. Here a Django app under live HTTP traffic recorded from a demo harness driving it loses its network path to Redis for 21 seconds — every request keeps returning 200 off the in-memory cache tier, and the Redis tier resyncs itself on recovery: Full documentation lives at https://baldur.sh https://baldur.sh . - What is Baldur? https://baldur.sh/what-is-baldur/ — the problem it solves and how - Getting started: Django https://baldur.sh/getting-started/django/ · FastAPI https://baldur.sh/getting-started/fastapi/ · Flask https://baldur.sh/getting-started/flask/ · Celery https://baldur.sh/getting-started/celery/ - Concept guides https://baldur.sh — one page per capability, linked throughout this README - API reference https://baldur.sh/reference/ - Troubleshooting https://baldur.sh/troubleshooting/ - Compatibility https://baldur.sh/compatibility/ Building with an AI coding assistant Claude Code, Cursor, Copilot, Codex ? Run baldur init-ai in your repo to drop an AGENTS.md read by Cursor, Copilot, and Codex plus a CLAUDE.md that imports it for Claude Code — together they teach the assistant to reach for @baldur.protected "name" instead of hand-rolling a circuit breaker. See Using Baldur with AI assistants https://baldur.sh/getting-started/ai-assistants/ . | Component | Minimum | Tested in CI | |---|---|---| | Python | 3.11 | 3.11 · 3.12 · 3.13 | | Django | 4.2 | 4.2 LTS · 5.2 LTS · 6.0 | | FastAPI | 0.100 | latest ≥ floor smoke | | Flask | 2.3 | latest ≥ floor smoke | | Celery | 5.3 | 5.4 | | Redis server | — | 7.x | See Compatibility https://baldur.sh/compatibility/ for the full matrix, the Python × Django test grid, and the version support policy. Baldur PRO adds the fleet-level machinery on top of the same API — nothing in the core gets relicensed or replaced: DLQ at scale https://baldur.sh/concepts/foundations/dlq-replay/ batch replay from the console, success-rate-driven pacing, and archive/purge retention , a hash-chained audit trail https://baldur.sh/concepts/pro/audit/ , unified notifications https://baldur.sh/concepts/pro/unified-notification/ , emergency mode https://baldur.sh/concepts/pro/emergency-mode/ , bulkhead thread-pool isolation https://baldur.sh/concepts/foundations/bulkhead/ , adaptive throttling https://baldur.sh/concepts/pro/throttle/ , canary recovery https://baldur.sh/concepts/pro/canary-recovery/ , governance gates https://baldur.sh/concepts/pro/governance/ , and a meta-watchdog https://baldur.sh/concepts/pro/meta-watchdog/ that watches Baldur itself. See the full OSS vs PRO capability matrix https://baldur.sh/concepts/oss-vs-pro/ and pricing https://baldur.sh/pricing/ . Baldur is in early access: the API is stable and the core is tested under sustained load with Sentinel failover, but the project is young — minor releases may still ship breaking changes, always with a changelog entry. It is looking for a small number of teams already running a Python service in production to work with directly. If that is you, the details and how to reach me are in Discussions https://github.com/baldurhq/baldur/discussions . How the project got here — including why it was nearly shelved in September 2026: retrospective Korean https://github.com/baldurhq/baldur/blob/main/POSTMORTEM.ko.md . Baldur is released under the Apache License 2.0 — see LICENSE https://github.com/baldurhq/baldur/blob/main/LICENSE and NOTICE https://github.com/baldurhq/baldur/blob/main/NOTICE . Contributions are welcome under the Apache License 2.0. Pull requests are accepted through a sign-off-based DCO https://developercertificate.org/ flow — see CONTRIBUTING.md https://github.com/baldurhq/baldur/blob/main/CONTRIBUTING.md for the full model. - Ideas, or showing what you built → Discussions https://github.com/baldurhq/baldur/discussions . - Bugs / feature requests / docs → open an issue or a pull request. - Security → see SECURITY.md https://github.com/baldurhq/baldur/blob/main/SECURITY.md no public issues for vulnerabilities . - Usage questions / commercial → support@baldur.sh .