My background AI processes worked great in the demo and then fell apart the second they had to wait A developer argues that AI automation workflows break down once they must wait for human approvals, rate-limited APIs, webhooks, retries or scheduled resumes, because the request/response model cannot survive beyond a single HTTP request. The piece contrasts execution limits such as AWS Lambda's 15-minute cap and Cloudflare Workers' CPU-time constraints with durable orchestration tools like Inngest, whose concurrency limits active steps rather than total runs, and Temporal, which replays event history so workflows survive crashes, deploys and long waits. It also notes that self-hosted n8n runs production executions unlimited by default unless N8N_CONCURRENCY_PRODUCTION_LIMIT is set, and that queue mode splits n8n into a distributed background job system. I’ve watched this happen a bunch of times. The demo is clean: Then someone adds one very normal requirement: That’s the moment your “AI workflow” stops being a sequence of API calls and starts becoming distributed systems work. And a lot of AI automation demos completely hide that fact. If your workflow can outlive one HTTP request, you need to design for background execution on purpose. Your first successful automation teaches the wrong lesson. It teaches you that AI pipelines are just request/response code with a few extra steps: That works until the workflow needs to wait. A human approval. A rate-limited API. A follow-up webhook. A retry after failure. A scheduled resume. Now your code has to survive outside the lifespan of the original request. That’s where the wheels come off. A lot of people discover this through timeouts. AWS Lambda has a hard 15-minute max execution timeout. Cloudflare Workers are even stranger for orchestration-heavy jobs because you also care about CPU time, not just wall-clock time. So if your agent needs to: ...your nice little webhook handler is no longer the right abstraction. At that point you usually bolt on: That’s how a neat Friday demo becomes a very real Monday incident. More than most teams expect. Take n8n. In the demo, it feels like one app running one flow. In queue mode, it becomes a split architecture: That is not “just automation tooling” anymore. That is a background job system. A webhook-triggered AI flow that looked simple in the demo is now a distributed system wearing a low-code hat. That sounds dramatic, but it’s also why production bugs show up in places the workflow builder never warned you about. The logic was fine. The runtime wasn’t. Most teams hear “concurrency” and assume it means “how many runs exist at once.” It often does not. In self-hosted n8n regular mode, production executions are unlimited by default unless you set a limit. export N8N CONCURRENCY PRODUCTION LIMIT=20 In queue mode, workers can be tuned directly: n8n worker --concurrency=5 That’s already more subtle than people expect. Now compare that to Inngest. In Inngest, concurrency limits active steps, not total runs. That matters a lot for AI workflows. If a function is: ...it doesn’t necessarily consume active concurrency the whole time. That’s a much better fit for always-on agents than the classic model where every waiting job burns a worker slot. Here’s a tiny example: js import { inngest } from "./client"; export default inngest.createFunction { id: "generate-ai-summary", concurrency: 10, }, { event: "ai/summary.requested" }, async { event, step } = { const data = await step.run "get-data", async = { return getDataFromExternalSource event.data.id ; } ; const summary = await step.run "generate-summary", async = { return callLLM data ; } ; await step.run "save-summary", async = { return db.summaries.insertOne { id: event.data.id, summary } ; } ; } ; The syntax is not the interesting part. The checkpointing is. If save-summary fails, you don’t want to rerun the whole thing from the top and burn another expensive model call if you don’t need to. For AI automations, that difference matters a lot. This is the part that usually changes how people think about orchestration. Long-running automation is not mainly a timeout problem. It’s a state problem. Temporal gets this better than almost anyone. The core idea is simple: Workflows should survive crashes, deploys, pauses, and long waits by replaying event history, not by hoping some process memory stays alive forever. If your agent pipeline depends on in-memory state surviving overnight, you built a trap. If your workflow needs to continue after: ...plain request/response code is the wrong abstraction. Temporal is opinionated in a useful way. Workflow execution timeout is unlimited by default. Workflows can run for a very long time. Retries and replay are built into the model instead of being bolted on later. A minimal TypeScript sketch looks like this: python import { proxyActivities } from "@temporalio/workflow"; import type as activities from "./activities"; const { generateSummary, postToDiscord } = proxyActivities