{"slug": "show-hn-one-line-so-llm-workers-share-a-429-wait-and-keep-jobs-an-outage-stops", "title": "Show HN: One line so LLM workers share a 429 wait and keep jobs an outage stops", "summary": "Baldur, an early-access Python framework from baldurhq, released a decorator that makes LLM workers share a single backoff wait on a 429 or \"overloaded\" response and parks failed jobs for replay instead of losing them. The framework wraps the OpenAI, Anthropic, and google-genai SDKs, runs in-memory with no Redis or Docker, opens a circuit breaker at a 60-second request bound, and lists dead-lettered calls with their arguments in a built-in console at http://127.0.0.1:9090/. Baldur warns that replay=True re-runs an entire job, so a model call that already succeeded inside it is billed a second time, and a timed-out request may also bill twice on retry.", "body_md": "**English** | [한국어](https://github.com/baldurhq/baldur/blob/main/README.ko.md)\n\n**Early access — feedback wanted.** Trying Baldur on a real service? If anything gets in your way — installing, the docs, behavior you didn't expect — tell us in [Discussions](https://github.com/baldurhq/baldur/discussions) or [open an issue](https://github.com/baldurhq/baldur/issues/new/choose).\n\n**An API you depend on goes down for an hour. What happens to your app?**\n\nRequests hang until they time out, every worker fills up, and the jobs that failed in that hour are gone. Whether it's OpenAI, your payment provider, or your email service — Baldur fixes all three with one decorator, for Python services that don't have anyone on call.\n\n``` python\nimport baldur\nfrom openai import OpenAI\n\nllm = baldur.llm.wrap(OpenAI(), timeout=60.0)\n\n@baldur.protected(\"summarize\", replay=True)\ndef summarize(doc_id: str) -> str:\n    response = llm.chat.completions.create(\n        model=\"gpt-4o-mini\",\n        messages=[{\"role\": \"user\", \"content\": load_document(doc_id)}],\n    )\n    return response.choices[0].message.content\n```\n\nNo Redis, no Docker, no config to start: those two lines run in-memory until\nyou go multi-process. Not calling an LLM? The decorator alone does the same for\n[any dependency](#the-same-decorator-any-dependency).\n\nWhen the provider rate-limits you, dies, or just gets slow — mid-traffic:\n\n- **Your workers back off together.** A 429 or an \"overloaded\" answer installs\none wait every worker shares, at least as long as the provider asked: the\nfirst worker is refused and the rest wait, instead of each one discovering\nthe limit alone. A request the provider rejected (400, 422) is never retried.\n- **Your app keeps answering.** A hung request fails at the 60-second bound,\nthe circuit breaker opens, and calls fail fast — so a slow provider doesn't\ntake every worker down with it. Give the wrap`fallbacks=[...]` and a call\nmoves to the next endpoint instead (OpenAI SDK, Anthropic SDK, google-genai).\n- **Failed jobs are kept, not lost.** Every call that failed for good is\ncaptured with its arguments and listed in the built-in console at`http://127.0.0.1:9090/` . In a plain Python process, call`baldur.init()` once at startup to start the console; the Django, FastAPI, and Flask\nintegrations do that for you.\n- **They come back.**`replay=True` lets Baldur re-run a parked job from its\nstored arguments: from the console with a click, or automatically once the\nprovider recovers and the job's breaker closes — with a Celery worker running\nthe replay.\n\nSee it against the real `openai` SDK and a local fake provider — a rate limit,\nthen an outage, then every parked job replayed:\n\n*A real run. Playback shortens each pause to 3 seconds; the times on screen\nare the run's own. Run it yourself:*\n\n```\npip install \"baldur-framework[celery]\" openai\npython -m baldur.scripts.demo_llm_outage\n```\n\n**If Baldur breaks, does my call break?** No. A fault in Baldur's own\nbookkeeping, such as a failed dead-letter write, is logged: your call still\nreturns its own result or raises its own error. If Redis is unreachable, calls\nkeep running on each process's own state, and the workers stop sharing the 429\nwait and the breaker until it is back. One exception, by design: a call with\n`idempotency_key=` is refused with `IdempotencyUnavailableError` (after a few\nseconds' wait) rather than run without the shared record that keeps it from\nrunning twice.\n\n**Can a retry or a replay bill me twice?** Yes, so plan for it. A replay runs\nthe whole job again: a model call that had already succeeded inside the job is\nmade, and billed, a second time. A request that timed out may have finished on\nthe provider's side, so retrying it can bill twice too. Use `replay=True` on\njobs that are safe to run twice, or store each step's result yourself and skip\nthe steps that already finished.\n\nDjango, FastAPI, Flask, and Celery adapters included.\n\n**Already using your SDK's retries?** Around a decorated call, keep them.\nBaldur doesn't replace retry — it adds what retry can't: a breaker so one\nincident doesn't cost every request its retries, one wall-clock bound on what\nthe caller waits, a fallback, and the capture-and-replay no retry library gives\nyou. A `baldur.llm.wrap` client is the one exception: it turns the SDK's own\nretries off on its copy of the client, because there one coordinated retry\nreplaces them.\n\nThe Python package is `baldur` (you `import baldur`); the PyPI distribution is\n`baldur-framework`.\n\n```\npip install baldur-framework                 # framework-agnostic core\npip install baldur-framework[django]         # Django integration\npip install baldur-framework[django-api]     # Baldur's Django REST API (baldur.api.django.urls)\npip install baldur-framework[fastapi]        # FastAPI integration\npip install baldur-framework[flask]          # Flask integration\npip install baldur-framework[celery]         # Celery task protection\npip install baldur-framework[redis]          # Redis-backed shared state\npip install baldur-framework[prometheus]     # Prometheus metrics\n```\n\nA payment gateway, your database, an email provider — the call site never changes:\n\n``` python\n@baldur.protected(\"charge-customer\", dlq=True)\ndef charge(order_id: str, amount_cents: int) -> dict:\n    # Circuit breaker by default; dlq=True parks the call if it fails for\n    # good, with its arguments, and replays it once the gateway recovers.\n    return payment_gateway.charge(order_id, amount_cents)\n```\n\nWhen the gateway dies, the breaker opens and your service answers fast instead\nof stacking up timeouts; the charges that failed on the way out wait in the\ndead-letter queue and come back when it closes. (Replay is for work that failed\non the way out — never for a business rejection, and never for a checkout the\ncustomer already walked away from:\n[where that line sits](https://baldur.sh/concepts/foundations/dlq-replay/).)\n\n*The payment demo: the gateway goes unreachable mid-traffic, seven charges are\ncaptured with their arguments, and all seven are replayed on recovery. Zero\nlost. A real run, with the breaker states and DLQ tallies read live from the\nframework. The decorator itself is `pip install baldur-framework` and nothing\nelse; the demo adds the `celery` extra for its in-process stand-in worker —\nstill one process, no Redis, no broker. Run it yourself:*\n\n```\npip install \"baldur-framework[celery]\"\npython -m baldur.scripts.demo_self_healing\n```\n\nNeed more than the default? Compose the pipeline declaratively:\n\n```\n@baldur.protected(\n    \"summarize\",\n    timeout=30.0,                            # one bound on what the caller waits\n    fallback=lambda: last_good_summary(),    # graceful answer while OPEN\n    idempotency_key=\"doc_id\",                # a redelivered job pays once\n)\ndef summarize(doc_id: str) -> str:\n    return llm_api.summarize(doc_id)\n```\n\n**Notice what isn't there: `retry=`.** Your SDK almost certainly retries\nalready — `anthropic` and `openai` default to two attempts with backoff, boto3\nhas an adaptive mode — and it retries better than a generic wrapper can,\nbecause it knows which status codes are worth another attempt and honours\n`retry-after`. Keep it. What no SDK gives you is the rest: a breaker, so a\nprovider incident doesn't mean *every* request pays its retries before failing;\none wall-clock bound on what your caller waits, retries included (an SDK's own\nworst case is `timeout × (max_retries + 1)` — 30 minutes at `anthropic`'s\ndefaults); a fallback; and a dedup key that survives a job redelivery the SDK\nnever sees. `retry=True` is there for the calls that don't retry themselves.\n\nSync and async callables are both supported — the decorator auto-detects coroutine functions.\n\n| Capability | What it gives you | \n|---|---|\n| [Circuit breaker](https://baldur.sh/concepts/oss/circuit-breaker/) | Stops cascading failure; bounded half-open probes on recovery | \n| [Retry with backoff](https://baldur.sh/concepts/oss/retry/) | Exponential backoff with jitter and bounded attempts | \n| [Fallback & composition](https://baldur.sh/concepts/foundations/composition/) | One ordered pipeline for all resilience patterns | \n| [Idempotency](https://baldur.sh/concepts/oss/idempotency/) | Concurrent duplicate calls execute the side effect exactly once | \n| [Bulkhead isolation](https://baldur.sh/concepts/foundations/bulkhead/) | Each dependency gets a fixed slice of concurrency, so one slow dependency can't drain every worker | \n| [Dead-letter queue + replay](https://baldur.sh/concepts/foundations/dlq-replay/) | A call that fails for good is captured with its context and replayed once the dependency recovers | \n| [Health checks](https://baldur.sh/concepts/oss/health-check/) | Liveness/readiness that reflect real dependency state | \n| [Graceful shutdown](https://baldur.sh/concepts/oss/graceful-shutdown/) | Drain in-flight work cleanly on restart and deploy | \n| [Metrics](https://baldur.sh/concepts/oss/metrics/) | Prometheus and OpenTelemetry, emitted by default | \n| [System control](https://baldur.sh/concepts/oss/system-control/) | Instant kill switch and dry-run mode for Baldur's automation — no redeploy | \n| [Web console](https://baldur.sh/concepts/foundations/web-console/) | Built-in operations console: live breaker state, controls, recovery | \n| [Precomputed cache](https://baldur.sh/concepts/oss/precomputed-cache/) | Health/status endpoints answer from a warm cache, so constant probing stays cheap | \n\nThe read path heals the same way. Here a Django app under live HTTP traffic (recorded from a demo harness driving it) loses its network path to Redis for 21 seconds — every request keeps returning 200 off the in-memory cache tier, and the Redis tier resyncs itself on recovery:\n\nFull documentation lives at **[https://baldur.sh](https://baldur.sh)**.\n\n- [What is Baldur?](https://baldur.sh/what-is-baldur/) — the problem it solves and how\n- Getting started: [Django](https://baldur.sh/getting-started/django/) ·[FastAPI](https://baldur.sh/getting-started/fastapi/) ·[Flask](https://baldur.sh/getting-started/flask/) ·[Celery](https://baldur.sh/getting-started/celery/)\n- [Concept guides](https://baldur.sh) — one page per capability, linked\nthroughout this README\n- [API reference](https://baldur.sh/reference/)\n- [Troubleshooting](https://baldur.sh/troubleshooting/)\n- [Compatibility](https://baldur.sh/compatibility/)\n\nBuilding with an AI coding assistant (Claude Code, Cursor, Copilot, Codex)? Run\n`baldur init-ai` in your repo to drop an `AGENTS.md` (read by Cursor, Copilot,\nand Codex) plus a `CLAUDE.md` that imports it for Claude Code — together they\nteach the assistant to reach for `@baldur.protected(\"name\")` instead of\nhand-rolling a circuit breaker. See\n[Using Baldur with AI assistants](https://baldur.sh/getting-started/ai-assistants/).\n\n| Component | Minimum | Tested in CI | \n|---|---|---|\n| Python | 3.11 | 3.11 · 3.12 · 3.13 | \n| Django | 4.2 | 4.2 LTS · 5.2 LTS · 6.0 | \n| FastAPI | 0.100 | latest ≥ floor (smoke) | \n| Flask | 2.3 | latest ≥ floor (smoke) | \n| Celery | 5.3 | 5.4 | \n| Redis server | — | 7.x | \n\nSee [Compatibility](https://baldur.sh/compatibility/) for the full matrix, the\nPython × Django test grid, and the version support policy.\n\nBaldur PRO adds the fleet-level machinery on top of the same API — nothing in\nthe core gets relicensed or replaced:\n[DLQ at scale](https://baldur.sh/concepts/foundations/dlq-replay/) (batch replay from the\nconsole, success-rate-driven pacing, and archive/purge\nretention), a hash-chained [audit trail](https://baldur.sh/concepts/pro/audit/),\n[unified notifications](https://baldur.sh/concepts/pro/unified-notification/),\n[emergency mode](https://baldur.sh/concepts/pro/emergency-mode/),\n[bulkhead thread-pool isolation](https://baldur.sh/concepts/foundations/bulkhead/),\n[adaptive throttling](https://baldur.sh/concepts/pro/throttle/),\n[canary recovery](https://baldur.sh/concepts/pro/canary-recovery/),\n[governance gates](https://baldur.sh/concepts/pro/governance/), and a\n[meta-watchdog](https://baldur.sh/concepts/pro/meta-watchdog/) that watches Baldur itself.\nSee the full [OSS vs PRO capability matrix](https://baldur.sh/concepts/oss-vs-pro/) and\n[pricing](https://baldur.sh/pricing/).\n\nBaldur is in early access: the API is stable and the core is tested under\nsustained load with Sentinel failover, but the project is young — minor\nreleases may still ship breaking changes, always with a changelog entry. It is\nlooking for a small number of teams already running a Python service in\nproduction to work with directly. If that is you, the details and how to reach\nme are in [Discussions](https://github.com/baldurhq/baldur/discussions).\n\nHow the project got here — including why it was nearly shelved in September\n2026: [retrospective (Korean)](https://github.com/baldurhq/baldur/blob/main/POSTMORTEM.ko.md).\n\nBaldur is released under the Apache License 2.0 — see [LICENSE](https://github.com/baldurhq/baldur/blob/main/LICENSE) and\n[NOTICE](https://github.com/baldurhq/baldur/blob/main/NOTICE).\n\nContributions are welcome under the Apache License 2.0. Pull requests are\naccepted through a sign-off-based [DCO](https://developercertificate.org/) flow —\nsee [CONTRIBUTING.md](https://github.com/baldurhq/baldur/blob/main/CONTRIBUTING.md) for the full model.\n\n- **Ideas, or showing what you built** →[Discussions](https://github.com/baldurhq/baldur/discussions) .\n- **Bugs / feature requests / docs** → open an issue or a pull request.\n- **Security** → see[SECURITY.md](https://github.com/baldurhq/baldur/blob/main/SECURITY.md) (no public issues for vulnerabilities).\n- **Usage questions / commercial** →`support@baldur.sh` .", "url": "https://wpnews.pro/news/show-hn-one-line-so-llm-workers-share-a-429-wait-and-keep-jobs-an-outage-stops", "canonical_source": "https://github.com/baldurhq/baldur", "published_at": "2026-10-02 16:40:24+00:00", "updated_at": "2026-10-02 17:06:35.233333+00:00", "lang": "en", "topics": ["ai-infrastructure", "developer-tools", "mlops", "large-language-models"], "entities": ["Baldur", "baldurhq", "OpenAI", "Anthropic", "google-genai", "Celery", "Django", "FastAPI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-one-line-so-llm-workers-share-a-429-wait-and-keep-jobs-an-outage-stops", "markdown": "https://wpnews.pro/news/show-hn-one-line-so-llm-workers-share-a-429-wait-and-keep-jobs-an-outage-stops.md", "text": "https://wpnews.pro/news/show-hn-one-line-so-llm-workers-share-a-429-wait-and-keep-jobs-an-outage-stops.txt", "jsonld": "https://wpnews.pro/news/show-hn-one-line-so-llm-workers-share-a-429-wait-and-keep-jobs-an-outage-stops.jsonld"}}