cd /news/ai-agents/my-background-ai-processes-worked-gr… · home › topics › ai-agents › article
[ARTICLE · art-141026] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

My background AI processes worked great in the demo and then fell apart the second they had to wait

A developer argues that AI automation workflows break down once they must wait for human approvals, rate-limited APIs, webhooks, retries or scheduled resumes, because the request/response model cannot survive beyond a single HTTP request. The piece contrasts execution limits such as AWS Lambda's 15-minute cap and Cloudflare Workers' CPU-time constraints with durable orchestration tools like Inngest, whose concurrency limits active steps rather than total runs, and Temporal, which replays event history so workflows survive crashes, deploys and long waits. It also notes that self-hosted n8n runs production executions unlimited by default unless N8N_CONCURRENCY_PRODUCTION_LIMIT is set, and that queue mode splits n8n into a distributed background job system.

by read7 min views3 publishedSep 28, 2026

I’ve watched this happen a bunch of times.

The demo is clean:

Then someone adds one very normal requirement:

That’s the moment your “AI workflow” stops being a sequence of API calls and starts becoming distributed systems work.

And a lot of AI automation demos completely hide that fact.

If your workflow can outlive one HTTP request, you need to design for background execution on purpose.

Your first successful automation teaches the wrong lesson.

It teaches you that AI pipelines are just request/response code with a few extra steps:

That works until the workflow needs to wait.

A human approval.

A rate-limited API.

A follow-up webhook.

A retry after failure.

A scheduled resume.

Now your code has to survive outside the lifespan of the original request.

That’s where the wheels come off.

A lot of people discover this through timeouts.

AWS Lambda has a hard 15-minute max execution timeout.

Cloudflare Workers are even stranger for orchestration-heavy jobs because you also care about CPU time, not just wall-clock time.

So if your agent needs to:

...your nice little webhook handler is no longer the right abstraction.

At that point you usually bolt on:

That’s how a neat Friday demo becomes a very real Monday incident.

More than most teams expect.

Take n8n.

In the demo, it feels like one app running one flow.

In queue mode, it becomes a split architecture:

That is not “just automation tooling” anymore.

That is a background job system.

A webhook-triggered AI flow that looked simple in the demo is now a distributed system wearing a low-code hat.

That sounds dramatic, but it’s also why production bugs show up in places the workflow builder never warned you about.

The logic was fine.

The runtime wasn’t.

Most teams hear “concurrency” and assume it means “how many runs exist at once.”

It often does not.

In self-hosted n8n regular mode, production executions are unlimited by default unless you set a limit.

export N8N_CONCURRENCY_PRODUCTION_LIMIT=20

In queue mode, workers can be tuned directly:

n8n worker --concurrency=5

That’s already more subtle than people expect.

Now compare that to Inngest.

In Inngest, concurrency limits active steps, not total runs.

That matters a lot for AI workflows.

If a function is:

...it doesn’t necessarily consume active concurrency the whole time.

That’s a much better fit for always-on agents than the classic model where every waiting job burns a worker slot.

Here’s a tiny example:

import { inngest } from "./client";

export default inngest.createFunction(
  {
    id: "generate-ai-summary",
    concurrency: 10,
  },
  { event: "ai/summary.requested" },
  async ({ event, step }) => {
    const data = await step.run("get-data", async () => {
      return getDataFromExternalSource(event.data.id);
    });

    const summary = await step.run("generate-summary", async () => {
      return callLLM(data);
    });

    await step.run("save-summary", async () => {
      return db.summaries.insertOne({ id: event.data.id, summary });
    });
  }
);

The syntax is not the interesting part.

The checkpointing is.

If save-summary fails, you don’t want to rerun the whole thing from the top and burn another expensive model call if you don’t need to.

For AI automations, that difference matters a lot.

This is the part that usually changes how people think about orchestration.

Long-running automation is not mainly a timeout problem.

It’s a state problem.

Temporal gets this better than almost anyone.

The core idea is simple:

Workflows should survive crashes, deploys, s, and long waits by replaying event history, not by hoping some process memory stays alive forever.

If your agent pipeline depends on in-memory state surviving overnight, you built a trap.

If your workflow needs to continue after:

...plain request/response code is the wrong abstraction.

Temporal is opinionated in a useful way.

Workflow execution timeout is unlimited by default.

Workflows can run for a very long time.

Retries and replay are built into the model instead of being bolted on later.

A minimal TypeScript sketch looks like this:

import { proxyActivities } from "@temporalio/workflow";
import type * as activities from "./activities";

const { generateSummary, postToDiscord } = proxyActivities<typeof activities>({
  startToCloseTimeout: "5 minutes",
  retry: {
    initialInterval: "1 second",
    backoffCoefficient: 2,
    maximumInterval: "10 minutes",
    maximumAttempts: 5,
  },
});

export async function aiApprovalWorkflow(input: { docId: string }) {
  const summary = await generateSummary(input.docId);

  // imagine waiting for approval signal here
  // then continue after approval arrives

  await postToDiscord(summary);
}

That isn’t overengineering if your workflow behaves more like a business process than a script.

It’s just being honest about the shape of the problem.

This is where a lot of AI teams get burned.

They put LLM calls directly inside workflow logic and then wonder why retries become messy and replay becomes dangerous.

Temporal’s model is the right one here:

Put failure-prone or non-deterministic work into Activities.

That includes things like:

Why?

Because workflow code needs to stay deterministic for replay.

Activities are where retries belong.

If GPT-5 times out, you want the model call retried.

You do not want your orchestration logic inventing a new branch of reality because replay took a different path.

For AI systems, the flaky edges are the product:

That’s exactly where demos break in production.

There is no universal winner, but there are very clear tradeoffs.

Option What it’s really good at
n8n queue mode Best when you already live in n8n and need webhook/timer ingestion separated from execution; Redis brokers execution IDs, workers run jobs, and you need to manage operational details like shared encryption keys and durable DB state
Temporal Best when workflows are long-running, failure-prone, and need to survive crashes cleanly; event-history replay and automatic Activity retries are the point
Inngest Best for event-driven automations with waits, s, resumptions, and step-level checkpointing without as much queue plumbing

My opinionated version:

Not always.

Some jobs are fine with a queue plus workers.

If a task:

...a full workflow engine may be unnecessary.

But the more common mistake is the opposite one.

Teams keep pretending their automation is “just a webhook” long after it has become a living process with:

At that point, avoiding durable execution is not simplicity.

It’s denial.

Background execution solves timeout pain.

It also creates a new pile of work.

Now you need to care about:

This is the part that makes an always-on AI system feel less like “calling an LLM” and more like operating a small factory.

That framing is healthier.

Because once the workflow is always on, the hard part is rarely the prompt.

It’s the lifecycle.

If your AI automation can , retry, wait, or outlive a web request, ask these questions before you ship:

- Where does workflow state live?
- What happens if the worker dies mid-step?
- Can one failed model call resume from a checkpoint?
- Does waiting consume concurrency?
- Can the workflow survive a deploy?
- Are downstream side effects idempotent?
- Can I tell the difference between sleeping, retrying, and stuck?

If you can’t answer those clearly, the demo is ahead of the architecture.

There’s another production problem people underestimate.

Once workflows become long-running and retry-heavy, token-based pricing gets annoying fast.

A failed AI step that gets replayed or retried a few times is not just an engineering problem. It becomes a billing problem too.

That’s especially painful for teams running agents in:

If your automations run all day and you’re routing across models like GPT-5, Claude Opus, and Grok, per-token pricing turns every retry policy into a finance discussion.

That’s one reason Standard Compute is interesting for this kind of workload.

It gives you an OpenAI-compatible API with flat monthly pricing, so background agents can run continuously without somebody watching token spend like a hawk. If your orchestration layer is already complicated, removing billing unpredictability helps a lot.

If your AI workflow can outlive a request, design it like a background job from day one.

Not after the first timeout.

Not after the first failed approval flow.

Not after the first customer says, “it worked yesterday.”

Start with the boring questions:

The happy-path demo is still useful.

It proves the idea.

But always-on AI systems do not fail on the happy path.

They fail in the gaps between steps:

That hidden layer is most of the architecture.

Once you see that, you stop building AI automations like glorified request handlers.

And your production incidents get a lot less surprising.

── more in #ai-agents 4 stories · sorted by recency
── more on @n8n 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/my-background-ai-pro…] indexed:0 read:7min 2026-09-28 · —