Agent workflows are not expensive because the model is expensive. They are expensive because nothing in the loop stops the loop. A chatbot query is one request. An agent task is a chain: plan, read, edit, test, read the error, retry. Each step re-sends context. The context grows. The retries multiply. Nobody is counting.
KPMG's Q3 2026 AI Pulse survey put a number on this. 93% of respondents went over their AI budget, and the same coverage reports agent-based workflows running 5 to 30 times more computationally intensive than chatbot queries. Both of those figures are single-sourced in what I could find, so read them as direction rather than precision. The better-attested numbers point the same way: only 43% of organisations had put usage or token budgets in place, while 74% were requiring cost review at approval time. You cannot cap what you cannot see, and most teams built the agent before they built the meter.
Here is the fix I want in every agent codebase. It is small on purpose.
A developer ships an agent that refactors a module. The agent reads six files, writes a patch, runs the test suite, reads the failure, patches again, runs again. That is a normal, successful task. It also just cost more than a week of human review, and nobody can tell you the number because the loop never reported.
The reflex fix is to add a token ceiling at the provider dashboard. That caps spend. It does not tell you which step is expensive, and it fires after the money is gone. By the time the dashboard emails you, the task already finished or already failed.
The fix that works is a ceiling inside the loop, plus a record of where the tokens went.
Wrap the model call, not the agent. This is the whole mechanism. Fail loudly when a single task exceeds its ceiling, and attach the reason to the error so the caller knows whether to retry smaller or stop.
from dataclasses import dataclass, field
class BudgetExceeded(RuntimeError):
def __init__(self, task_id, spent, ceiling, breakdown):
self.task_id = task_id
self.spent = spent
self.ceiling = ceiling
self.breakdown = breakdown
top = sorted(breakdown.items(), key=lambda kv: -kv[1])[:3]
detail = ", ".join(f"{step}={tok}" for step, tok in top)
super().__init__(
f"task {task_id} used {spent} tokens, ceiling is {ceiling}. "
f"largest steps: {detail}"
)
@dataclass
class TaskBudget:
task_id: str
ceiling: int
spent: int = 0
breakdown: dict = field(default_factory=dict)
def record(self, step, tokens):
self.spent += tokens
self.breakdown[step] = self.breakdown.get(step, 0) + tokens
if self.spent > self.ceiling:
raise BudgetExceeded(
self.task_id, self.spent, self.ceiling, dict(self.breakdown)
)
Now route every model call through it.
def call_model(budget, step, messages, client, model="gpt-4o-mini"):
response = client.chat.completions.create(
model=model, messages=messages
)
usage = response.usage
budget.record(step, usage.total_tokens)
return response
Three properties matter here.
The budget is per task, not per process. A long-lived server process has no natural ceiling. A task does.
The check happens after recording, not before. Pre-flight checks cannot catch a retry loop, because the loop is inside a single call. Post-call accounting catches it on the iteration that crosses the line.
The error carries the breakdown. A bare "budget exceeded" makes the caller guess. Top three steps by token count tells them whether to shard the task, drop to a cheaper model, or stop.
Most of the cost is not the reasoning. It is the reading. A step that reads ten files and returns a list of edits does not need your largest model.
Split the loop into two tiers. Use a small model for retrieval, extraction, and summarisation steps. Use the large model only for the step that decides what to change.
CHEAP_STEPS = {"read_file", "list_dir", "summarise", "extract"}
def model_for(step):
return "gpt-4o-mini" if step in CHEAP_STEPS else "gpt-4o"
In a typical coding agent, the read and summarise steps outnumber the decide steps three to one, and they consume the bulk of the input tokens because they carry file contents. Routing them down is usually the single largest saving available, and it costs nothing in quality because those steps do not do reasoning.
The code above reads response.usage, and that field is not always populated. With streaming responses many providers return no usage block at all, because the count is only known once the stream closes. A guard that silently records zero is worse than no guard, because the ceiling never fires and you believe you are protected.
Force the count to exist before you trust the guard. If your provider supports it, request the usage field explicitly on the final chunk. Otherwise accumulate locally by counting the tokens you send and the tokens the model returns on each chunk.
def call_model_streaming(budget, step, messages, client, model="gpt-4o-mini"):
stream = client.chat.completions.create(
model=model, messages=messages, stream=True
)
out_tokens = 0
chunks = []
for event in stream:
if event.usage: # available when requested
out_tokens = event.usage.completion_tokens
delta = event.choices[0].delta.content
if delta:
chunks.append(delta)
in_tokens = count_tokens(messages) # local count, provider-independent
budget.record(step, in_tokens + out_tokens)
return "".join(chunks)
Count input tokens locally. That number depends only on your own messages, so you can always compute it, and it is usually the larger half in an agent loop where every step re-sends the whole history.
A budget guard raises an error when a task gets expensive. It does not tell you the task was stuck. Those are different failures and they need different instruments.
A retry loop has a signature: the same step, the same input, the same output, more than once. Detect it at the point where you have all three, and fail with a different error so it does not get confused with a spend problem.
def reject_repeat(task_id, step, fingerprint, seen, limit=2):
key = (step, fingerprint)
seen[key] = seen.get(key, 0) + 1
if seen[key] > limit:
raise LoopDetected(
f"task {task_id} repeated step {step} {seen[key]} times "
f"with identical input and output"
)
Fingerprint the step input, not the whole message list. Hash the action and its arguments. If two invocations of the read step produce the same fingerprint, you are not making progress and the model is not going to start.
This check is cheap and it catches the failure the budget guard cannot. A loop that retries nine times at low cost per step can stay under a generous ceiling the entire time.
A safety mechanism that has never fired is indistinguishable from a safety mechanism that does not work. Write the test before you need it.
def test_budget_raises_with_breakdown():
b = TaskBudget(task_id="t1", ceiling=1000)
b.record("read_file", 600)
with pytest.raises(BudgetExceeded) as exc:
b.record("summarise", 500)
assert exc.value.spent == 1100
assert exc.value.breakdown["read_file"] == 600
assert "read_file=600" in str(exc.value)
Assert on the message content, not just the exception type. The breakdown is the part your future self will rely on at 2am, so it is the part worth pinning down.
Add a second test that the guard does not fire under the ceiling, because a guard that always raises trains people to disable it.
Here is a worked example, using round numbers to make the shape easy to see rather than to report a specific run.
A refactoring task, agent with a read-plan-edit-test loop, ceiling of 400,000 tokens. The task uses 612,000 tokens, and the guard raises on the sixth test iteration. The breakdown reads read_file=310000, summarise=180000, run_tests=90000. The loop was not reasoning badly. It was re-reading the same four files after every failed test run, because the agent had no memory of what it had already read.
Two changes follow from reading that breakdown. The read step caches by file hash, removing roughly 200,000 tokens. The test step stops re-reading unchanged output, removing another 90,000. Final cost lands near 180,000 tokens, under the original ceiling, with no change to the quality of the patch.
Neither change is a model upgrade. Both come from reading the breakdown the guard produced. That is the whole argument for instrumenting before you optimise.
You do not need a dashboard to start. Log three numbers per completed task.
After about twenty tasks the pattern is obvious. The third number is the one that finds the bug: a task that retries four times on the same test failure is a loop with no exit condition, and no token ceiling in the world fixes that. The ceiling just tells you faster.
A budget guard tells you a task got expensive. It does not tell you the task was worth it. Those are different questions, and a guard answers only the first. Keep the two apart when you report to whoever funds the compute: the question "did this task succeed" needs a success signal, not a spend signal.
Gartner's forecast that more than 40% of agentic AI projects will be cancelled by the end of 2027 cites unclear value and governance failures, not cost alone. Cost is the visible symptom. Value is the actual complaint.
This post was written with AI assistance. The author is responsible for its content.