{"slug": "my-background-ai-processes-worked-great-in-the-demo-and-then-fell-apart-the-they", "title": "My background AI processes worked great in the demo and then fell apart the second they had to wait", "summary": "A developer argues that AI automation workflows break down once they must wait for human approvals, rate-limited APIs, webhooks, retries or scheduled resumes, because the request/response model cannot survive beyond a single HTTP request. The piece contrasts execution limits such as AWS Lambda's 15-minute cap and Cloudflare Workers' CPU-time constraints with durable orchestration tools like Inngest, whose concurrency limits active steps rather than total runs, and Temporal, which replays event history so workflows survive crashes, deploys and long waits. It also notes that self-hosted n8n runs production executions unlimited by default unless N8N_CONCURRENCY_PRODUCTION_LIMIT is set, and that queue mode splits n8n into a distributed background job system.", "body_md": "I’ve watched this happen a bunch of times.\n\nThe demo is clean:\n\nThen someone adds one very normal requirement:\n\nThat’s the moment your “AI workflow” stops being a sequence of API calls and starts becoming distributed systems work.\n\nAnd a lot of AI automation demos completely hide that fact.\n\nIf your workflow can outlive one HTTP request, you need to design for background execution on purpose.\n\nYour first successful automation teaches the wrong lesson.\n\nIt teaches you that AI pipelines are just request/response code with a few extra steps:\n\nThat works until the workflow needs to wait.\n\nA human approval.\n\nA rate-limited API.\n\nA follow-up webhook.\n\nA retry after failure.\n\nA scheduled resume.\n\nNow your code has to survive outside the lifespan of the original request.\n\nThat’s where the wheels come off.\n\nA lot of people discover this through timeouts.\n\nAWS Lambda has a hard 15-minute max execution timeout.\n\nCloudflare Workers are even stranger for orchestration-heavy jobs because you also care about CPU time, not just wall-clock time.\n\nSo if your agent needs to:\n\n...your nice little webhook handler is no longer the right abstraction.\n\nAt that point you usually bolt on:\n\nThat’s how a neat Friday demo becomes a very real Monday incident.\n\nMore than most teams expect.\n\nTake n8n.\n\nIn the demo, it feels like one app running one flow.\n\nIn queue mode, it becomes a split architecture:\n\nThat is not “just automation tooling” anymore.\n\nThat is a background job system.\n\nA webhook-triggered AI flow that looked simple in the demo is now a distributed system wearing a low-code hat.\n\nThat sounds dramatic, but it’s also why production bugs show up in places the workflow builder never warned you about.\n\nThe logic was fine.\n\nThe runtime wasn’t.\n\nMost teams hear “concurrency” and assume it means “how many runs exist at once.”\n\nIt often does not.\n\nIn self-hosted n8n regular mode, production executions are unlimited by default unless you set a limit.\n\n```\nexport N8N_CONCURRENCY_PRODUCTION_LIMIT=20\n```\n\nIn queue mode, workers can be tuned directly:\n\n```\nn8n worker --concurrency=5\n```\n\nThat’s already more subtle than people expect.\n\nNow compare that to Inngest.\n\nIn Inngest, concurrency limits active steps, not total runs.\n\nThat matters a lot for AI workflows.\n\nIf a function is:\n\n...it doesn’t necessarily consume active concurrency the whole time.\n\nThat’s a much better fit for always-on agents than the classic model where every waiting job burns a worker slot.\n\nHere’s a tiny example:\n\n``` js\nimport { inngest } from \"./client\";\n\nexport default inngest.createFunction(\n  {\n    id: \"generate-ai-summary\",\n    concurrency: 10,\n  },\n  { event: \"ai/summary.requested\" },\n  async ({ event, step }) => {\n    const data = await step.run(\"get-data\", async () => {\n      return getDataFromExternalSource(event.data.id);\n    });\n\n    const summary = await step.run(\"generate-summary\", async () => {\n      return callLLM(data);\n    });\n\n    await step.run(\"save-summary\", async () => {\n      return db.summaries.insertOne({ id: event.data.id, summary });\n    });\n  }\n);\n```\n\nThe syntax is not the interesting part.\n\nThe checkpointing is.\n\nIf `save-summary` fails, you don’t want to rerun the whole thing from the top and burn another expensive model call if you don’t need to.\n\nFor AI automations, that difference matters a lot.\n\nThis is the part that usually changes how people think about orchestration.\n\nLong-running automation is not mainly a timeout problem.\n\nIt’s a state problem.\n\nTemporal gets this better than almost anyone.\n\nThe core idea is simple:\n\nWorkflows should survive crashes, deploys, pauses, and long waits by replaying event history, not by hoping some process memory stays alive forever.\n\nIf your agent pipeline depends on in-memory state surviving overnight, you built a trap.\n\nIf your workflow needs to continue after:\n\n...plain request/response code is the wrong abstraction.\n\nTemporal is opinionated in a useful way.\n\nWorkflow execution timeout is unlimited by default.\n\nWorkflows can run for a very long time.\n\nRetries and replay are built into the model instead of being bolted on later.\n\nA minimal TypeScript sketch looks like this:\n\n``` python\nimport { proxyActivities } from \"@temporalio/workflow\";\nimport type * as activities from \"./activities\";\n\nconst { generateSummary, postToDiscord } = proxyActivities<typeof activities>({\n  startToCloseTimeout: \"5 minutes\",\n  retry: {\n    initialInterval: \"1 second\",\n    backoffCoefficient: 2,\n    maximumInterval: \"10 minutes\",\n    maximumAttempts: 5,\n  },\n});\n\nexport async function aiApprovalWorkflow(input: { docId: string }) {\n  const summary = await generateSummary(input.docId);\n\n  // imagine waiting for approval signal here\n  // then continue after approval arrives\n\n  await postToDiscord(summary);\n}\n```\n\nThat isn’t overengineering if your workflow behaves more like a business process than a script.\n\nIt’s just being honest about the shape of the problem.\n\nThis is where a lot of AI teams get burned.\n\nThey put LLM calls directly inside workflow logic and then wonder why retries become messy and replay becomes dangerous.\n\nTemporal’s model is the right one here:\n\nPut failure-prone or non-deterministic work into Activities.\n\nThat includes things like:\n\nWhy?\n\nBecause workflow code needs to stay deterministic for replay.\n\nActivities are where retries belong.\n\nIf GPT-5 times out, you want the model call retried.\n\nYou do not want your orchestration logic inventing a new branch of reality because replay took a different path.\n\nFor AI systems, the flaky edges are the product:\n\nThat’s exactly where demos break in production.\n\nThere is no universal winner, but there are very clear tradeoffs.\n\n| Option | What it’s really good at | \n|---|---|\n| n8n queue mode | Best when you already live in n8n and need webhook/timer ingestion separated from execution; Redis brokers execution IDs, workers run jobs, and you need to manage operational details like shared encryption keys and durable DB state | \n| Temporal | Best when workflows are long-running, failure-prone, and need to survive crashes cleanly; event-history replay and automatic Activity retries are the point | \n| Inngest | Best for event-driven automations with waits, pauses, resumptions, and step-level checkpointing without as much queue plumbing | \n\nMy opinionated version:\n\nNot always.\n\nSome jobs are fine with a queue plus workers.\n\nIf a task:\n\n...a full workflow engine may be unnecessary.\n\nBut the more common mistake is the opposite one.\n\nTeams keep pretending their automation is “just a webhook” long after it has become a living process with:\n\nAt that point, avoiding durable execution is not simplicity.\n\nIt’s denial.\n\nBackground execution solves timeout pain.\n\nIt also creates a new pile of work.\n\nNow you need to care about:\n\nThis is the part that makes an always-on AI system feel less like “calling an LLM” and more like operating a small factory.\n\nThat framing is healthier.\n\nBecause once the workflow is always on, the hard part is rarely the prompt.\n\nIt’s the lifecycle.\n\nIf your AI automation can pause, retry, wait, or outlive a web request, ask these questions before you ship:\n\n```\n- Where does workflow state live?\n- What happens if the worker dies mid-step?\n- Can one failed model call resume from a checkpoint?\n- Does waiting consume concurrency?\n- Can the workflow survive a deploy?\n- Are downstream side effects idempotent?\n- Can I tell the difference between sleeping, retrying, and stuck?\n```\n\nIf you can’t answer those clearly, the demo is ahead of the architecture.\n\nThere’s another production problem people underestimate.\n\nOnce workflows become long-running and retry-heavy, token-based pricing gets annoying fast.\n\nA failed AI step that gets replayed or retried a few times is not just an engineering problem. It becomes a billing problem too.\n\nThat’s especially painful for teams running agents in:\n\nIf your automations run all day and you’re routing across models like GPT-5, Claude Opus, and Grok, per-token pricing turns every retry policy into a finance discussion.\n\nThat’s one reason Standard Compute is interesting for this kind of workload.\n\nIt gives you an OpenAI-compatible API with flat monthly pricing, so background agents can run continuously without somebody watching token spend like a hawk. If your orchestration layer is already complicated, removing billing unpredictability helps a lot.\n\nIf your AI workflow can outlive a request, design it like a background job from day one.\n\nNot after the first timeout.\n\nNot after the first failed approval flow.\n\nNot after the first customer says, “it worked yesterday.”\n\nStart with the boring questions:\n\nThe happy-path demo is still useful.\n\nIt proves the idea.\n\nBut always-on AI systems do not fail on the happy path.\n\nThey fail in the gaps between steps:\n\nThat hidden layer is most of the architecture.\n\nOnce you see that, you stop building AI automations like glorified request handlers.\n\nAnd your production incidents get a lot less surprising.", "url": "https://wpnews.pro/news/my-background-ai-processes-worked-great-in-the-demo-and-then-fell-apart-the-they", "canonical_source": "https://dev.to/lars_winstand/my-background-ai-processes-worked-great-in-the-demo-and-then-fell-apart-the-second-they-had-to-wait-5a12", "published_at": "2026-09-28 14:11:17+00:00", "updated_at": "2026-09-28 14:19:54.016696+00:00", "lang": "en", "topics": ["ai-agents", "ai-infrastructure", "mlops", "developer-tools", "ai-tools"], "entities": ["n8n", "Inngest", "Temporal", "AWS Lambda", "Cloudflare Workers"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/my-background-ai-processes-worked-great-in-the-demo-and-then-fell-apart-the-they", "markdown": "https://wpnews.pro/news/my-background-ai-processes-worked-great-in-the-demo-and-then-fell-apart-the-they.md", "text": "https://wpnews.pro/news/my-background-ai-processes-worked-great-in-the-demo-and-then-fell-apart-the-they.txt", "jsonld": "https://wpnews.pro/news/my-background-ai-processes-worked-great-in-the-demo-and-then-fell-apart-the-they.jsonld"}}