# My headless AI server setup stopped dying when I got boring about queues and 30-second watchdogs

> Source: <https://dev.to/lars_winstand/my-headless-ai-server-setup-stopped-dying-when-i-got-boring-about-queues-and-30-second-watchdogs-3im5>
> Published: 2026-09-29 14:13:11+00:00

I realized our setup was fake-stable the first time a background agent stayed “running” while doing absolutely nothing.

No crash.

No stack trace.

Docker said the container was up. The process existed. CPU was low. Memory looked fine.

Meanwhile, webhook jobs were stacking up like a sink full of dishes.

That was the week I stopped trusting "container is running" as a useful definition of healthy.

If you’re moving from a laptop-and-good-vibes setup to a real headless AI server for always-on agents, this failure mode shows up fast.

Usually the stack looks something like this:

Everything looks fine until one worker hangs without technically dying.

The fix was not a new framework.

It was three boring things:

That’s the whole post.

The biggest mistake was letting webhook intake, cron scheduling, and LLM-heavy execution all happen in the same process.

That works on a laptop.

It gets weird in production.

A slow Claude call delays a Stripe webhook.

A giant PDF parse blocks the event loop.

A CPU spike from OCR makes your Discord bot miss heartbeats.

One stuck worker makes the whole service look haunted.

The clean split is the same pattern n8n uses in queue mode:

That separation matters more than most teams think.

Ingress should stay lightweight.

Execution should be allowed to be slow, retryable, and disposable.

```
[Stripe/GitHub/cron/webhooks]
            |
            v
    [main ingress service]
            |
            v
        [Redis queue]
            |
    -------------------
    |        |        |
| --- | --- |
    v        v        v
[worker] [worker] [worker]
    |        |        |
| --- | --- |
    ---------+--------
            |
            v
   [OpenAI-compatible endpoint]
            |
            v
 [GPT-5.4 / Claude Opus 4.6 / Grok 4.20]
```

If you’re using n8n, this is basically why queue mode exists.

Also: once you do this, local-only state starts breaking things.

That means:

This is the boring infrastructure tax for not losing jobs.

A queue gives you another chance.

That’s it.

Whether that second chance helps or hurts depends on idempotency.

If a worker does this:

...then the queue may redeliver the job.

If the task is not replay-safe, you just built a duplicate side-effect machine.

This is not theoretical. It’s how these systems behave.

With Celery, for example, tasks can be redelivered if the worker dies before acknowledgment. Features like `acks_late` help, but they do not magically make your side effects safe.

The real fix is application-level:

Stripe’s webhook model is basically a giant sign that says: expect duplicates, expect retries, design accordingly.

If your endpoint is down, Stripe retries undelivered events for up to 3 days.

You can also recover manually by listing undelivered events.

Example:

```
curl -G https://api.stripe.com/v1/events \
  -u "sk_test_...:" \
  -d ending_before=evt_001 \
  -d "types[]=payment_intent.succeeded" \
  -d "types[]=payment_intent.payment_failed" \
  -d delivery_success=false
```

The important part is not the curl command.

The important part is that your system has to know whether it already processed `evt_001`.

A simple pattern is a table like this:

```
CREATE TABLE processed_events (
  event_id TEXT PRIMARY KEY,
  event_type TEXT NOT NULL,
  processed_at TIMESTAMP NOT NULL DEFAULT NOW()
);
```

Then your worker logic becomes:

``` js
async function handleStripeEvent(event) {
  const alreadyProcessed = await db.processed_events.findUnique({
    where: { event_id: event.id }
  });

  if (alreadyProcessed) {
    return { ok: true, duplicate: true };
  }

  await doSideEffects(event);

  await db.processed_events.create({
    data: {
      event_id: event.id,
      event_type: event.type,
    }
  });

  return { ok: true };
}
```

That pattern is not glamorous.

It is also the difference between “retries are safe” and “why did we send 4 receipts?”

Yes, use restart policies.

They help with crashes.

```
docker run -d --name worker --restart unless-stopped my-worker-image
```

But restart policies only help when the process exits.

My actual problem was worse:

That’s a different class of failure.

Docker does not know your Node.js event loop is wedged.

Docker does not know your Python worker is stuck in a bad call.

Docker does not know your queue heartbeat stopped.

That’s why “container is up” is not enough.

BullMQ has a much more useful idea of health than Docker does.

It marks jobs as stalled when an active worker stops renewing its lock.

By default, that check is around 30 seconds.

That is exactly the kind of signal I needed.

Because a worker that is still running but hasn’t renewed a lock in 30 seconds is not healthy in any meaningful sense.

A basic BullMQ worker looks like this:

``` python
import { Queue, Worker } from 'bullmq';
import IORedis from 'ioredis';

const connection = new IORedis(process.env.REDIS_URL);

const queue = new Queue('jobs', { connection });

const worker = new Worker(
  'jobs',
  async job => {
    console.log(`Processing job ${job.id}`);

    // your LLM / automation work here
    return await runTask(job.data);
  },
  { connection }
);

worker.on('completed', job => {
  console.log(`Job ${job.id} completed`);
});

worker.on('failed', (job, err) => {
  console.error(`Job ${job?.id} failed`, err);
});
```

And a retry config with jitter should be your default, not an afterthought:

```
await queue.add('test-retry', { foo: 'bar' }, {
  attempts: 8,
  backoff: {
    type: 'fixed',
    delay: 1000,
    jitter: 0.5,
  },
});
```

That `jitter: 0.5` matters a lot.

Without jitter, a worker fleet tends to fail together and retry together.

That’s how a small outage turns into a retry storm.

This one bit me hard.

If you run CPU-heavy work in the same Node.js process that needs to keep queue heartbeats alive, you can create stalls without crashing anything.

Examples:

The worker process is technically alive.

But queue progress says otherwise.

That’s why sandboxed processors or separate worker processes are worth it.

If a task can block the event loop, isolate it.

The best answer I found was: use multiple layers, but give each layer one job.

| Option | What it’s actually good at | 
|---|---|
| systemd watchdog + service restart | Detects hung processes via heartbeat instead of only exit codes. Great for VM or bare-metal workers. | 
| Docker restart policy | Restarts exited containers. Good default. Weak for hung-but-still-running workers. | 
| BullMQ or Celery | Decouples ingress from execution, supports retries and redelivery, and tolerates worker death if jobs are replay-safe. | 

My ranking for real-world usefulness:

Container restarts save you from crashes.

Queues save you from reality.

This was a much bigger deal than I expected.

Before, different workers were calling different providers directly:

That gets ugly fast.

Especially when every layer retries differently.

OpenAI-style clients already need backoff for 429s and transient failures. If your queue retries jobs and your model client also retries aggressively, you can easily create a retry storm from both sides.

Putting all workers behind one OpenAI-compatible endpoint fixed a lot:

That architecture is a great fit for AI agents and automations.

It also makes provider swaps much less painful.

If you have:

...then the model layer should not be reimplemented in every service.

A single OpenAI-compatible endpoint gives you a clean boundary.

This is exactly why products like Standard Compute are interesting for this setup.

You point your workers at one endpoint, keep your existing OpenAI-compatible SDKs, and avoid wiring provider-specific behavior into every queue consumer.

The other practical benefit is cost predictability.

When agents run 24/7, per-token billing gets annoying fast. A flat monthly model is a lot easier to reason about when you’re scaling background jobs, retries, and automations across a worker fleet.

For this kind of architecture, that matters more than people admit.

A single endpoint also centralizes failure.

That is real.

If your gateway is down or misconfigured, every worker feels it.

So if you do this, you still need:

Still, I would take one well-observed gateway over a zoo of half-maintained provider clients every time.

Especially for teams that want interchangeable workers.

If your setup is tiny, don’t cargo-cult a distributed system.

...then a single process with systemd, strict timeouts, and a dead-letter path may be enough.

You do not need Redis, BullMQ, PostgreSQL, S3, and six workers to rename files once an hour.

But once you add:

...the “simple” setup stops being simple.

It becomes fragile.

And fragile always looks cheap right before it gets expensive.

If I were rebuilding a headless AI server setup today, this is the baseline:

If you want something actionable, start here:

```
# 1. run Redis

docker run -d \
  --name redis \
  --restart unless-stopped \
  -p 6379:6379 \
  redis
# 2. split your app into:
#    - ingress service
#    - worker service
#    - shared database
# 3. add health checks that measure progress, not just process existence
# 4. make every external side effect idempotent
# 5. point all workers at one OpenAI-compatible endpoint
```

If you do only those five things, your setup will already be less fragile than a surprising number of “production” agent stacks.

The most reliable headless AI server setup I’ve used is boring on purpose.

Not the fanciest agent framework.

Not the prettiest architecture diagram.

Not the stack with the most logos in it.

Just:

That setup assumes workers will freeze.

It assumes webhooks will retry.

It assumes jobs will replay.

It assumes somebody will eventually ship CPU-heavy nonsense into a process that was supposed to stay responsive.

That’s why it works.

And if you’re running always-on AI agents, “nothing broke at 3 a.m.” is a much better definition of success than “the container was still up.”
