{"slug": "i-thought-my-cheap-ai-workflow-was-fine-until-i-counted-3000-tiny-calls", "title": "I thought my cheap AI workflow was fine until I counted 3,000 tiny calls", "summary": "A developer detailed how a seemingly cheap n8n lead-enrichment workflow that loops over 1,000 records with three AI steps per item actually generates 3,000 model calls, warning that request-per-minute limits, not token counts, are the real constraint for high-frequency automation. The engineer recommends estimating base, validation and retry call volumes before trusting any AI workflow, and notes that OpenAI's Batch API and Anthropic prompt caching only partially address the problem since batch jobs can't serve real-time needs and caching can't help unique per-row requests.", "body_md": "I didn’t get burned by one giant GPT-5 prompt.\n\nI got burned by a workflow that looked cheap.\n\nYou know the kind:\n\nLead enrichment. Support triage. CRM cleanup. Scraped page normalization.\n\nNothing fancy. Nothing that looks like it should trigger budget panic.\n\nAnd that’s exactly why it’s dangerous.\n\nThe mistake is simple: people inspect one AI call at a time.\n\nThey say:\n\nAll true.\n\nBut nobody multiplies.\n\nHere’s a totally normal n8n flow:\n\n``` php\nFetch 1,000 leads\n-> Loop Over Items (Batch Size: 1)\n-> OpenAI node: classify company\n-> OpenAI node: extract pain points\n-> OpenAI node: draft outreach angle\n```\n\nThat is already:\n\n```\n1,000 records * 3 AI steps = 3,000 model calls\n```\n\nAnd that’s before:\n\nThis is not an edge case. This is a standard automation pattern.\n\nIf you build with n8n, Zapier, Make, or custom worker queues, you’ve probably done this already.\n\nNot just in dollars.\n\nAlso in:\n\nA lot of teams focus on token count because that’s what pricing pages train you to do.\n\nBut provider limits are usually not just about tokens.\n\nOpenAI separates RPM and TPM for a reason. You can be nowhere near your token-per-minute cap and still hit request-per-minute limits. Their docs explicitly call out the idea that if your RPM is 20, then 20 requests of only 100 tokens each can still max you out.\n\nThat changes the architecture discussion.\n\nIf your workload is “analyze one giant contract,” token cost is the problem.\n\nIf your workload is “touch 8,000 CRM rows and make 3 tiny decisions on each,” request multiplication is usually the real problem.\n\nBecause staging lies.\n\nA test run on 20 records looks cheap.\n\nThen someone points the workflow at:\n\nAnd suddenly the cost shape changes.\n\nNot because prompts got bigger.\n\nBecause you turned on a machine that makes tiny calls thousands of times.\n\nI now do this before I trust any AI automation:\n\n```\nitems_per_day=1000\nai_steps_per_item=3\nretry_rate=0.1\nvalidation_calls_per_item=1\n\nbase_calls=$((items_per_day * ai_steps_per_item))\nvalidation_calls=$((items_per_day * validation_calls_per_item))\nretry_calls=$(python3 - <<'PY'\nitems=1000\nsteps=3\nretry_rate=0.1\nprint(int(items * steps * retry_rate))\nPY\n)\n\necho \"Base calls: $base_calls\"\necho \"Validation calls: $validation_calls\"\necho \"Retry calls: $retry_calls\"\n```\n\nEven rough math is enough.\n\nIf the answer is “we’re making 4,000 to 10,000 model requests a day,” you do not have a tiny workflow.\n\nYou have a high-frequency AI system.\n\nI’m not anti-OpenAI Batch API or anti-Anthropic prompt caching.\n\nBoth are good.\n\nBut they are not universal fixes for “my automation explodes into thousands of micro-calls.”\n\nOpenAI’s Batch API is legitimately useful.\n\nIt offers a 50% discount versus synchronous calls, and OpenAI explicitly positions it for jobs like large-scale classification and embeddings.\n\nThat maps well to:\n\nExample request shape:\n\n```\n{\"custom_id\":\"request-1\",\"method\":\"POST\",\"url\":\"/v1/chat/completions\",\"body\":{\"model\":\"gpt-3.5-turbo-0125\",\"messages\":[{\"role\":\"system\",\"content\":\"You are a helpful assistant.\"},{\"role\":\"user\",\"content\":\"Hello world!\"}],\"max_tokens\":1000}}\n```\n\nBut there are tradeoffs:\n\nIf your support workflow needs to classify a ticket now, Batch is not your answer.\n\nAnthropic prompt caching is also very real.\n\nWhen you have a large repeated prompt prefix, it can cut both latency and cost dramatically.\n\nThat’s excellent for:\n\nExample:\n\n``` python\nimport anthropic\n\nclient = anthropic.Anthropic()\nresponse = client.messages.create(\n    model=\"claude-opus-5-5\",\n    max_tokens=1024,\n    cache_control={\"type\": \"ephemeral\"},\n    system=\"You are an AI assistant tasked with analyzing literary works.\",\n    messages=[\n        {\"role\": \"user\", \"content\": \"Analyze the major themes in Pride and Prejudice.\"}\n    ],\n)\n```\n\nBut caching helps when requests share stable context.\n\nIt does not magically fix workloads where every row is different:\n\nYou can’t cache uniqueness.\n\nMost teams skip this and jump straight to model comparisons.\n\nThat’s backwards.\n\nFirst figure out the shape of your workload.\n\n| Option | Best fit | \n|---|---|\n| OpenAI synchronous API | Real-time flows where latency matters, but you still deal with per-token billing and RPM/TPM limits | \n| OpenAI Batch API | Large asynchronous jobs where lower cost matters more than immediate completion | \n| Anthropic prompt caching | Repeated prompt prefixes, long shared context, and workloads that benefit from cache reuse | \n| Per-item multi-step automation in n8n, Make, or Zapier | Easy to build, easy to underestimate, and very likely to multiply request count fast | \n\nIf your workload is a few giant prompts, token pricing is the main issue.\n\nIf your workload is thousands of tiny calls, the issue is usually a mix of:\n\nThat last one matters more than people admit.\n\nThis is the hidden tax.\n\nTeams start making worse technical decisions because they’re trying not to trigger more model calls.\n\nI’ve seen teams do all of these:\n\nThat is not clean engineering.\n\nThat is workflow design shaped by billing anxiety.\n\nAnd the model bill is not the only meter.\n\nYou may also be managing:\n\nAt some point this stops being a prompt engineering problem and becomes systems design.\n\nThis is the mental model that finally made it click for me.\n\nA lot of record-level automations do not make one AI decision.\n\nThey make a committee.\n\nFor one item, you might do:\n\nEvery call is defensible.\n\nTogether, they behave like a swarm.\n\nThat’s why I’m increasingly opinionated about this:\n\nFor high-volume operational workflows, pricing model matters almost as much as model quality.\n\nNot because GPT-5, Claude Opus, Grok, Qwen, or Llama are bad.\n\nBecause once your team stops fearing each micro-call, you build better automations:\n\nIf I’m building a real system, I’d break the problem down like this.\n\nThat third category is where a lot of agent and automation teams actually live.\n\nIf you have an n8n, Make, Zapier, or custom agent workflow, map it like this:\n\n```\nworkflow_audit:\n  items_per_day: 1000\n  ai_steps_per_item: 3\n  average_retries_per_100_calls: 12\n  validation_calls_per_item: 1\n  fallback_model_enabled: true\n  real_time_steps:\n    - classify_ticket\n    - route_priority\n  async_steps:\n    - nightly_summary\n    - enrichment_backfill\n  repeated_prompt_prefixes:\n    - support_policy_context\n    - extraction_schema\n```\n\nThen answer these questions honestly:\n\nThat exercise usually reveals one of two stories.\n\nIf that’s true, prompt caching can help a lot.\n\nIf that’s true, you need to stop evaluating cost one prompt at a time.\n\nYou need to think in workflow volume.\n\nThis is exactly why products like Standard Compute are interesting for agent and automation workloads.\n\nIf you’re running lots of small calls across n8n, Make, Zapier, OpenClaw, or custom workers, flat-rate unlimited compute changes the design space.\n\nInstead of asking:\n\nYou can build the workflow you actually want.\n\nStandard Compute is a drop-in OpenAI-compatible API, so you can usually swap it into existing SDKs or HTTP clients without rebuilding your stack. Under the hood it routes across models like GPT-5.4, Claude Opus 4.6, and Grok 4.20, with batching, prompt optimization, and throttling designed for exactly this kind of high-frequency automation work.\n\nThat matters if your bottleneck is not one giant prompt.\n\nIt matters if your bottleneck is thousands of tiny decisions.\n\nBefore you optimize prompts, count calls.\n\nBefore you compare GPT-5 vs Claude, count calls.\n\nBefore you celebrate a cheap staging run, count calls.\n\nMost teams think their cost problem is “big prompts are expensive.”\n\nA surprising number actually have a different problem:\n\nsmall prompts, repeated constantly, inside workflows that looked harmless.\n\nThat was my mistake.\n\nIf you’re building AI automations, don’t price a single request.\n\nPrice the swarm.", "url": "https://wpnews.pro/news/i-thought-my-cheap-ai-workflow-was-fine-until-i-counted-3000-tiny-calls", "canonical_source": "https://dev.to/lars_winstand/i-thought-my-cheap-ai-workflow-was-fine-until-i-counted-3000-tiny-calls-4eaa", "published_at": "2026-09-30 06:11:07+00:00", "updated_at": "2026-09-30 06:16:33.887544+00:00", "lang": "en", "topics": ["ai-agents", "ai-infrastructure", "mlops", "ai-tools", "large-language-models"], "entities": ["OpenAI", "Anthropic", "n8n", "Zapier", "Make", "OpenAI Batch API", "GPT-5"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-thought-my-cheap-ai-workflow-was-fine-until-i-counted-3000-tiny-calls", "markdown": "https://wpnews.pro/news/i-thought-my-cheap-ai-workflow-was-fine-until-i-counted-3000-tiny-calls.md", "text": "https://wpnews.pro/news/i-thought-my-cheap-ai-workflow-was-fine-until-i-counted-3000-tiny-calls.txt", "jsonld": "https://wpnews.pro/news/i-thought-my-cheap-ai-workflow-was-fine-until-i-counted-3000-tiny-calls.jsonld"}}