English | νκ΅μ΄
Early access β feedback wanted. Trying Baldur on a real service? If anything gets in your way β installing, the docs, behavior you didn't expect β tell us in Discussions or open an issue.
An API you depend on goes down for an hour. What happens to your app?
Requests hang until they time out, every worker fills up, and the jobs that failed in that hour are gone. Whether it's OpenAI, your payment provider, or your email service β Baldur fixes all three with one decorator, for Python services that don't have anyone on call.
import baldur
from openai import OpenAI
llm = baldur.llm.wrap(OpenAI(), timeout=60.0)
@baldur.protected("summarize", replay=True)
def summarize(doc_id: str) -> str:
response = llm.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": load_document(doc_id)}],
)
return response.choices[0].message.content
No Redis, no Docker, no config to start: those two lines run in-memory until you go multi-process. Not calling an LLM? The decorator alone does the same for any dependency.
When the provider rate-limits you, dies, or just gets slow β mid-traffic:
- Your workers back off together. A 429 or an "overloaded" answer installs one wait every worker shares, at least as long as the provider asked: the first worker is refused and the rest wait, instead of each one discovering the limit alone. A request the provider rejected (400, 422) is never retried.
- Your app keeps answering. A hung request fails at the 60-second bound,
the circuit breaker opens, and calls fail fast β so a slow provider doesn't
take every worker down with it. Give the wrap
fallbacks=[...]and a call moves to the next endpoint instead (OpenAI SDK, Anthropic SDK, google-genai). - Failed jobs are kept, not lost. Every call that failed for good is
captured with its arguments and listed in the built-in console at
http://127.0.0.1:9090/. In a plain Python process, callbaldur.init()once at startup to start the console; the Django, FastAPI, and Flask integrations do that for you. - They come back.
replay=Truelets Baldur re-run a parked job from its stored arguments: from the console with a click, or automatically once the provider recovers and the job's breaker closes β with a Celery worker running the replay.
See it against the real openai SDK and a local fake provider β a rate limit,
then an outage, then every parked job replayed:
A real run. Playback shortens each to 3 seconds; the times on screen are the run's own. Run it yourself:
pip install "baldur-framework[celery]" openai
python -m baldur.scripts.demo_llm_outage
If Baldur breaks, does my call break? No. A fault in Baldur's own
bookkeeping, such as a failed dead-letter write, is logged: your call still
returns its own result or raises its own error. If Redis is unreachable, calls
keep running on each process's own state, and the workers stop sharing the 429
wait and the breaker until it is back. One exception, by design: a call with
idempotency_key= is refused with IdempotencyUnavailableError (after a few
seconds' wait) rather than run without the shared record that keeps it from
running twice.
Can a retry or a replay bill me twice? Yes, so plan for it. A replay runs
the whole job again: a model call that had already succeeded inside the job is
made, and billed, a second time. A request that timed out may have finished on
the provider's side, so retrying it can bill twice too. Use replay=True on
jobs that are safe to run twice, or store each step's result yourself and skip
the steps that already finished.
Django, FastAPI, Flask, and Celery adapters included.
Already using your SDK's retries? Around a decorated call, keep them.
Baldur doesn't replace retry β it adds what retry can't: a breaker so one
incident doesn't cost every request its retries, one wall-clock bound on what
the caller waits, a fallback, and the capture-and-replay no retry library gives
you. A baldur.llm.wrap client is the one exception: it turns the SDK's own
retries off on its copy of the client, because there one coordinated retry
replaces them.
The Python package is baldur (you import baldur); the PyPI distribution is
baldur-framework.
pip install baldur-framework # framework-agnostic core
pip install baldur-framework[django] # Django integration
pip install baldur-framework[django-api] # Baldur's Django REST API (baldur.api.django.urls)
pip install baldur-framework[fastapi] # FastAPI integration
pip install baldur-framework[flask] # Flask integration
pip install baldur-framework[celery] # Celery task protection
pip install baldur-framework[redis] # Redis-backed shared state
pip install baldur-framework[prometheus] # Prometheus metrics
A payment gateway, your database, an email provider β the call site never changes:
@baldur.protected("charge-customer", dlq=True)
def charge(order_id: str, amount_cents: int) -> dict:
return payment_gateway.charge(order_id, amount_cents)
When the gateway dies, the breaker opens and your service answers fast instead of stacking up timeouts; the charges that failed on the way out wait in the dead-letter queue and come back when it closes. (Replay is for work that failed on the way out β never for a business rejection, and never for a checkout the customer already walked away from: where that line sits.)
The payment demo: the gateway goes unreachable mid-traffic, seven charges are
captured with their arguments, and all seven are replayed on recovery. Zero
lost. A real run, with the breaker states and DLQ tallies read live from the
framework. The decorator itself is pip install baldur-framework and nothing
else; the demo adds the celery extra for its in-process stand-in worker β
still one process, no Redis, no broker. Run it yourself:
pip install "baldur-framework[celery]"
python -m baldur.scripts.demo_self_healing
Need more than the default? Compose the pipeline declaratively:
@baldur.protected(
"summarize",
timeout=30.0, # one bound on what the caller waits
fallback=lambda: last_good_summary(), # graceful answer while OPEN
idempotency_key="doc_id", # a redelivered job pays once
)
def summarize(doc_id: str) -> str:
return llm_api.summarize(doc_id)
Notice what isn't there: retry=. Your SDK almost certainly retries
already β anthropic and openai default to two attempts with backoff, boto3
has an adaptive mode β and it retries better than a generic wrapper can,
because it knows which status codes are worth another attempt and honours
retry-after. Keep it. What no SDK gives you is the rest: a breaker, so a
provider incident doesn't mean every request pays its retries before failing;
one wall-clock bound on what your caller waits, retries included (an SDK's own
worst case is timeout Γ (max_retries + 1) β 30 minutes at anthropic's
defaults); a fallback; and a dedup key that survives a job redelivery the SDK
never sees. retry=True is there for the calls that don't retry themselves.
Sync and async callables are both supported β the decorator auto-detects coroutine functions.
| Capability | What it gives you |
|---|---|
| Circuit breaker | Stops cascading failure; bounded half-open probes on recovery |
| Retry with backoff | Exponential backoff with jitter and bounded attempts |
| Fallback & composition | One ordered pipeline for all resilience patterns |
| Idempotency | Concurrent duplicate calls execute the side effect exactly once |
| Bulkhead isolation | Each dependency gets a fixed slice of concurrency, so one slow dependency can't drain every worker |
| Dead-letter queue + replay | A call that fails for good is captured with its context and replayed once the dependency recovers |
| Health checks | Liveness/readiness that reflect real dependency state |
| Graceful shutdown | Drain in-flight work cleanly on restart and deploy |
| Metrics | Prometheus and OpenTelemetry, emitted by default |
| System control | Instant kill switch and dry-run mode for Baldur's automation β no redeploy |
| Web console | Built-in operations console: live breaker state, controls, recovery |
| Precomputed cache | Health/status endpoints answer from a warm cache, so constant probing stays cheap |
The read path heals the same way. Here a Django app under live HTTP traffic (recorded from a demo harness driving it) loses its network path to Redis for 21 seconds β every request keeps returning 200 off the in-memory cache tier, and the Redis tier resyncs itself on recovery:
Full documentation lives at https://baldur.sh.
- What is Baldur? β the problem it solves and how
- Getting started: Django Β·FastAPI Β·Flask Β·Celery
- Concept guides β one page per capability, linked throughout this README
- API reference
- Troubleshooting
- Compatibility
Building with an AI coding assistant (Claude Code, Cursor, Copilot, Codex)? Run
baldur init-ai in your repo to drop an AGENTS.md (read by Cursor, Copilot,
and Codex) plus a CLAUDE.md that imports it for Claude Code β together they
teach the assistant to reach for @baldur.protected("name") instead of
hand-rolling a circuit breaker. See
Using Baldur with AI assistants.
| Component | Minimum | Tested in CI |
|---|---|---|
| Python | 3.11 | 3.11 Β· 3.12 Β· 3.13 |
| Django | 4.2 | 4.2 LTS Β· 5.2 LTS Β· 6.0 |
| FastAPI | 0.100 | latest β₯ floor (smoke) |
| Flask | 2.3 | latest β₯ floor (smoke) |
| Celery | 5.3 | 5.4 |
| Redis server | β | 7.x |
See Compatibility for the full matrix, the Python Γ Django test grid, and the version support policy.
Baldur PRO adds the fleet-level machinery on top of the same API β nothing in the core gets relicensed or replaced: DLQ at scale (batch replay from the console, success-rate-driven pacing, and archive/purge retention), a hash-chained audit trail, unified notifications, emergency mode, bulkhead thread-pool isolation, adaptive throttling, canary recovery, governance gates, and a meta-watchdog that watches Baldur itself. See the full OSS vs PRO capability matrix and pricing.
Baldur is in early access: the API is stable and the core is tested under sustained load with Sentinel failover, but the project is young β minor releases may still ship breaking changes, always with a changelog entry. It is looking for a small number of teams already running a Python service in production to work with directly. If that is you, the details and how to reach me are in Discussions.
How the project got here β including why it was nearly shelved in September 2026: retrospective (Korean).
Baldur is released under the Apache License 2.0 β see LICENSE and NOTICE.
Contributions are welcome under the Apache License 2.0. Pull requests are accepted through a sign-off-based DCO flow β see CONTRIBUTING.md for the full model.
- Ideas, or showing what you built βDiscussions .
- Bugs / feature requests / docs β open an issue or a pull request.
- Security β seeSECURITY.md (no public issues for vulnerabilities).
- Usage questions / commercial β
support@baldur.sh.