# I thought my cheap AI workflow was fine until I counted 3,000 tiny calls

> Source: <https://dev.to/lars_winstand/i-thought-my-cheap-ai-workflow-was-fine-until-i-counted-3000-tiny-calls-4eaa>
> Published: 2026-09-30 06:11:07+00:00

I didn’t get burned by one giant GPT-5 prompt.

I got burned by a workflow that looked cheap.

You know the kind:

Lead enrichment. Support triage. CRM cleanup. Scraped page normalization.

Nothing fancy. Nothing that looks like it should trigger budget panic.

And that’s exactly why it’s dangerous.

The mistake is simple: people inspect one AI call at a time.

They say:

All true.

But nobody multiplies.

Here’s a totally normal n8n flow:

``` php
Fetch 1,000 leads
-> Loop Over Items (Batch Size: 1)
-> OpenAI node: classify company
-> OpenAI node: extract pain points
-> OpenAI node: draft outreach angle
```

That is already:

```
1,000 records * 3 AI steps = 3,000 model calls
```

And that’s before:

This is not an edge case. This is a standard automation pattern.

If you build with n8n, Zapier, Make, or custom worker queues, you’ve probably done this already.

Not just in dollars.

Also in:

A lot of teams focus on token count because that’s what pricing pages train you to do.

But provider limits are usually not just about tokens.

OpenAI separates RPM and TPM for a reason. You can be nowhere near your token-per-minute cap and still hit request-per-minute limits. Their docs explicitly call out the idea that if your RPM is 20, then 20 requests of only 100 tokens each can still max you out.

That changes the architecture discussion.

If your workload is “analyze one giant contract,” token cost is the problem.

If your workload is “touch 8,000 CRM rows and make 3 tiny decisions on each,” request multiplication is usually the real problem.

Because staging lies.

A test run on 20 records looks cheap.

Then someone points the workflow at:

And suddenly the cost shape changes.

Not because prompts got bigger.

Because you turned on a machine that makes tiny calls thousands of times.

I now do this before I trust any AI automation:

```
items_per_day=1000
ai_steps_per_item=3
retry_rate=0.1
validation_calls_per_item=1

base_calls=$((items_per_day * ai_steps_per_item))
validation_calls=$((items_per_day * validation_calls_per_item))
retry_calls=$(python3 - <<'PY'
items=1000
steps=3
retry_rate=0.1
print(int(items * steps * retry_rate))
PY
)

echo "Base calls: $base_calls"
echo "Validation calls: $validation_calls"
echo "Retry calls: $retry_calls"
```

Even rough math is enough.

If the answer is “we’re making 4,000 to 10,000 model requests a day,” you do not have a tiny workflow.

You have a high-frequency AI system.

I’m not anti-OpenAI Batch API or anti-Anthropic prompt caching.

Both are good.

But they are not universal fixes for “my automation explodes into thousands of micro-calls.”

OpenAI’s Batch API is legitimately useful.

It offers a 50% discount versus synchronous calls, and OpenAI explicitly positions it for jobs like large-scale classification and embeddings.

That maps well to:

Example request shape:

```
{"custom_id":"request-1","method":"POST","url":"/v1/chat/completions","body":{"model":"gpt-3.5-turbo-0125","messages":[{"role":"system","content":"You are a helpful assistant."},{"role":"user","content":"Hello world!"}],"max_tokens":1000}}
```

But there are tradeoffs:

If your support workflow needs to classify a ticket now, Batch is not your answer.

Anthropic prompt caching is also very real.

When you have a large repeated prompt prefix, it can cut both latency and cost dramatically.

That’s excellent for:

Example:

``` python
import anthropic

client = anthropic.Anthropic()
response = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=1024,
    cache_control={"type": "ephemeral"},
    system="You are an AI assistant tasked with analyzing literary works.",
    messages=[
        {"role": "user", "content": "Analyze the major themes in Pride and Prejudice."}
    ],
)
```

But caching helps when requests share stable context.

It does not magically fix workloads where every row is different:

You can’t cache uniqueness.

Most teams skip this and jump straight to model comparisons.

That’s backwards.

First figure out the shape of your workload.

| Option | Best fit | 
|---|---|
| OpenAI synchronous API | Real-time flows where latency matters, but you still deal with per-token billing and RPM/TPM limits | 
| OpenAI Batch API | Large asynchronous jobs where lower cost matters more than immediate completion | 
| Anthropic prompt caching | Repeated prompt prefixes, long shared context, and workloads that benefit from cache reuse | 
| Per-item multi-step automation in n8n, Make, or Zapier | Easy to build, easy to underestimate, and very likely to multiply request count fast | 

If your workload is a few giant prompts, token pricing is the main issue.

If your workload is thousands of tiny calls, the issue is usually a mix of:

That last one matters more than people admit.

This is the hidden tax.

Teams start making worse technical decisions because they’re trying not to trigger more model calls.

I’ve seen teams do all of these:

That is not clean engineering.

That is workflow design shaped by billing anxiety.

And the model bill is not the only meter.

You may also be managing:

At some point this stops being a prompt engineering problem and becomes systems design.

This is the mental model that finally made it click for me.

A lot of record-level automations do not make one AI decision.

They make a committee.

For one item, you might do:

Every call is defensible.

Together, they behave like a swarm.

That’s why I’m increasingly opinionated about this:

For high-volume operational workflows, pricing model matters almost as much as model quality.

Not because GPT-5, Claude Opus, Grok, Qwen, or Llama are bad.

Because once your team stops fearing each micro-call, you build better automations:

If I’m building a real system, I’d break the problem down like this.

That third category is where a lot of agent and automation teams actually live.

If you have an n8n, Make, Zapier, or custom agent workflow, map it like this:

```
workflow_audit:
  items_per_day: 1000
  ai_steps_per_item: 3
  average_retries_per_100_calls: 12
  validation_calls_per_item: 1
  fallback_model_enabled: true
  real_time_steps:
    - classify_ticket
    - route_priority
  async_steps:
    - nightly_summary
    - enrichment_backfill
  repeated_prompt_prefixes:
    - support_policy_context
    - extraction_schema
```

Then answer these questions honestly:

That exercise usually reveals one of two stories.

If that’s true, prompt caching can help a lot.

If that’s true, you need to stop evaluating cost one prompt at a time.

You need to think in workflow volume.

This is exactly why products like Standard Compute are interesting for agent and automation workloads.

If you’re running lots of small calls across n8n, Make, Zapier, OpenClaw, or custom workers, flat-rate unlimited compute changes the design space.

Instead of asking:

You can build the workflow you actually want.

Standard Compute is a drop-in OpenAI-compatible API, so you can usually swap it into existing SDKs or HTTP clients without rebuilding your stack. Under the hood it routes across models like GPT-5.4, Claude Opus 4.6, and Grok 4.20, with batching, prompt optimization, and throttling designed for exactly this kind of high-frequency automation work.

That matters if your bottleneck is not one giant prompt.

It matters if your bottleneck is thousands of tiny decisions.

Before you optimize prompts, count calls.

Before you compare GPT-5 vs Claude, count calls.

Before you celebrate a cheap staging run, count calls.

Most teams think their cost problem is “big prompts are expensive.”

A surprising number actually have a different problem:

small prompts, repeated constantly, inside workflows that looked harmless.

That was my mistake.

If you’re building AI automations, don’t price a single request.

Price the swarm.
