alreadyspent too much. LLM-budget-cap flips this by implementing an atomic spend cap using Redis, effectively acting as a circuit breaker for your API costs.
The core problem it solves is the latency and inconsistency of official provider dashboards. By the time a cloud provider's billing API updates, a runaway loop in your code could have already burned through hundreds of dollars. This tool tracks usage in real-time via Redis, ensuring that once your predefined limit is hit, the requests stop immediately.
Getting Started #
Since this is a lightweight utility, deployment is straightforward. You'll need a Redis instance to handle the atomic increments.
-
Install the package via your preferred manager.
-
Configure your Redis connection to track the budget keys.
-
Wrap your LLM API calls with the budget check logic.
llm_budget_cap.set_limit(user_id="user_123", limit=10.00)
if llm_budget_cap.has_budget("user_123"):
llm_budget_cap.add_spend("user_123", cost=0.02)
else:
Is it worth it? #
If you are building a SaaS where users have their own API keys or you're deploying an LLM agent that can loop autonomously, this is a mandatory addition to your AI workflow. It's far more reliable than writing your own custom counter in a relational database, which would likely suffer from race conditions.
The use of Redis ensures the check is fast enough that it doesn't add noticeable latency to the request pipeline. For a production-ready deployment, this is a solid way to implement a hard ceiling on costs without relying on the slow-updating billing portals of the model providers.
Full source code is available here:https://github.com/Rentheria/llm-budget-cap
Next Nunchaku 4-bit Diffusion in Diffusers →