How do you stop an AI agent from overspending on GPUs? A developer published a guide showing how to cap AI agent GPU spending at the platform level rather than in prompts, citing an example where a forgotten 8-GPU H100 run from Friday evening to Monday morning would cost $1,331.20 at $2.60/hr, versus stopping after roughly 1.9 hours under a $40 cap. The writeup compares budget controls across AWS, GCP, Azure, Modal and Nodus, noting that major cloud budget data updates only every 8 to 24 hours, and lays out five enforcement layers including per-run cost caps, blocking budgets, scoped credentials, dry-run approval, and timeouts. Short answer October 11, 2026 : put the limits where the money is spent, not in the prompt: a hard dollar cap on every job the agent launches, a budget that blocks not just alerts around all agent work, and a credential scoped so the agent cannot raise either one. Add a dry-run estimate that you approve before anything is created, plus timeouts and idle stops, and a runaway agent costs you the cap instead of a weekend of GPU time. Below: why cloud budget alerts are too slow for agents, the five layers that actually stop spend, and copy-paste config for Claude Code, Codex and Nodus. An agent that loops, asks for H100:8 instead of L4 , or forgets to stop a machine spends at GPU speed. An H100 is listed at $2.60/hr on the Nodus pricing page https://www.nodus-compute.ai/pricing/ as of October 11, 2026 , so one forgotten 8 GPU run from Friday 6 pm to Monday 10 am 64 hours comes to 8 x $2.60 x 64 = $1,331.20 . The same run with a $40 cap stops after about 1.9 hours. "Don't spend more than $20" in a system prompt is a request. Models misread it, lose it in a long session, or follow instructions injected through logs and files. You need a limit the platform enforces, and most cloud budgets were built to inform a finance team, not to stop a process within minutes. Checked against each platform's own docs on October 11, 2026. | Platform | Control | Stops work by itself? | How quickly | |---|---|---|---| | AWS | AWS Budgets with budget actions | Only through actions: apply a deny IAM policy or SCP, or stop targeted instances | Budget data updates up to 3 times a day, typically 8 to 12 hours apart | | GCP | Alerts-only budgets | No, alerts only | Follows billing data, which lags | | GCP | Spend cap budgets | Yes, pauses the service, but only for the Gemini API, Vertex AI, Cloud Run and Cloud Run functions not GPU VMs | Faster than billing reports, still not instantaneous | | Azure | Cost Management budgets | No: "Resources aren't affected, and your consumption isn't stopped" | Data in 8 to 24 hours, evaluated every 24 hours | | Modal | Workspace budget, spend limit, Environment budgets Team and Enterprise | Yes, as monthly caps; Functions also have timeouts default 300 s, up to 24 h | Not stated | | Nodus | maxCostUSD on each object, plus Budgets with action: Block | Yes, per run and per org, project, member or label | Checked before paid work starts, then every few minutes | Sources: AWS Budgets https://docs.aws.amazon.com/cost-management/latest/userguide/budgets-managing-costs.html , AWS budget actions https://docs.aws.amazon.com/cost-management/latest/userguide/budgets-controls.html , GCP budgets https://cloud.google.com/billing/docs/how-to/budgets , GCP spend caps https://cloud.google.com/billing/docs/how-to/budgets-spend-caps , Azure budgets https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/tutorial-acm-create-budgets , Modal budgets https://modal.com/docs/guide/budgets , Modal timeouts https://modal.com/docs/guide/timeouts , Nodus budgets https://www.nodus-compute.ai/docs/guides/billing/budgets/ . The pattern: on the big clouds, a GPU agent can run for hours before a budget reacts. The hard stop has to live with whatever launches the work. | Layer | What it prevents | On Nodus | |---|---|---| | 1. Per-run cap | One runaway job: a loop, the wrong GPU count, a hung step | maxCostUSD or --max-cost ; the Job checkpoints and suspends with MaxCostReached | | 2. Blocking budget | Many small jobs adding up over a day or a month | A Budget with action: Block , scoped to a project, member or label | | 3. Scoped credential | The agent creating kinds you never meant it to, or raising its own ceiling | Scopes picked on the MCP consent page; nodus create apikey --scopes --projects --expires | | 4. Estimate, then approval | Surprise launches | MCP apply returns a dry-run and an etag first, and creates only on a second call with confirmed: true and the same etag | | 5. Timeouts and idle stops | Machines nobody is using | --timeout , Sandbox idleTimeout and maxLifetime | Thresholds and webhooks budget.threshold , budget.exceeded , billing.low balance tell you when any of them fires. Set up once with pip install nodus-compute and nodus login . Everything an agent creates through the Nodus MCP server carries the label nodus.dev/launched-by: mcp , and a Budget counts work whose labels matched when it was created. Point a blocking Budget at that label and every agent-launched Job, Sandbox or TrainingJob shares one ceiling: apiVersion: nodus.dev/v1 kind: Budget metadata: name: agents-monthly spec: limitUSD: "100.00" period: Monthly UTC calendar month scope: selector: matchLabels: nodus.dev/launched-by: mcp action: Block Notify would only send notices thresholds: 50, 80, 100 notify: emails: you@example.com nodus apply -f budget.yaml nodus get budget agents-monthly spent, held for running work, remaining If the agent uses an API key instead of MCP, give it its own project and scope the Budget there: nodus create budget agents --limit 100 --period Monthly --scope-project agents --action Block . At the limit, new agent work is refused with 402 BudgetExceeded , naming the Budget, what is left and what the work needs. Running work stops gracefully with Funded=False , and checkpointed Jobs save first. Details: when money runs out https://www.nodus-compute.ai/docs/guides/billing/when-money-runs-out/ . When you connect Claude Code, Codex or Cursor to the hosted server, the consent page asks which org and scopes the client gets, and the tools it sees follow those scopes. For a headless agent or CI, mint a key that expires: nodus create project agents --display-name "Agents" KEY=$ nodus create apikey agent-runner --scopes jobs:write,volumes:read --projects agents --expires 168h NODUS API KEY=$KEY nodus auth can-i create workspaces should answer no Keep Budget write access out of the agent's grant or key. Caps can be raised but never lowered, so an agent that can update a Job can also raise its maxCostUSD ; the Budget is the wall it cannot move. Revoking is fast: nodus delete oauthgrant