{"slug": "my-headless-ai-server-setup-stopped-dying-when-i-got-boring-about-queues-and-30", "title": "My headless AI server setup stopped dying when I got boring about queues and 30-second watchdogs", "summary": "A developer described splitting a headless AI agent server into separate ingress, Redis queue, and worker processes after discovering that containers reported as running while background agents silently hung, causing webhook jobs to pile up. The setup relies on idempotent event handling, such as a processed_events table keyed on Stripe event IDs, so that queue redelivery and retries do not produce duplicate side effects. The developer notes that Celery's acks_late helps but does not make side effects safe on its own, and that Stripe retries undelivered events for up to three days.", "body_md": "I realized our setup was fake-stable the first time a background agent stayed “running” while doing absolutely nothing.\n\nNo crash.\n\nNo stack trace.\n\nDocker said the container was up. The process existed. CPU was low. Memory looked fine.\n\nMeanwhile, webhook jobs were stacking up like a sink full of dishes.\n\nThat was the week I stopped trusting \"container is running\" as a useful definition of healthy.\n\nIf you’re moving from a laptop-and-good-vibes setup to a real headless AI server for always-on agents, this failure mode shows up fast.\n\nUsually the stack looks something like this:\n\nEverything looks fine until one worker hangs without technically dying.\n\nThe fix was not a new framework.\n\nIt was three boring things:\n\nThat’s the whole post.\n\nThe biggest mistake was letting webhook intake, cron scheduling, and LLM-heavy execution all happen in the same process.\n\nThat works on a laptop.\n\nIt gets weird in production.\n\nA slow Claude call delays a Stripe webhook.\n\nA giant PDF parse blocks the event loop.\n\nA CPU spike from OCR makes your Discord bot miss heartbeats.\n\nOne stuck worker makes the whole service look haunted.\n\nThe clean split is the same pattern n8n uses in queue mode:\n\nThat separation matters more than most teams think.\n\nIngress should stay lightweight.\n\nExecution should be allowed to be slow, retryable, and disposable.\n\n```\n[Stripe/GitHub/cron/webhooks]\n            |\n            v\n    [main ingress service]\n            |\n            v\n        [Redis queue]\n            |\n    -------------------\n    |        |        |\n| --- | --- |\n    v        v        v\n[worker] [worker] [worker]\n    |        |        |\n| --- | --- |\n    ---------+--------\n            |\n            v\n   [OpenAI-compatible endpoint]\n            |\n            v\n [GPT-5.4 / Claude Opus 4.6 / Grok 4.20]\n```\n\nIf you’re using n8n, this is basically why queue mode exists.\n\nAlso: once you do this, local-only state starts breaking things.\n\nThat means:\n\nThis is the boring infrastructure tax for not losing jobs.\n\nA queue gives you another chance.\n\nThat’s it.\n\nWhether that second chance helps or hurts depends on idempotency.\n\nIf a worker does this:\n\n...then the queue may redeliver the job.\n\nIf the task is not replay-safe, you just built a duplicate side-effect machine.\n\nThis is not theoretical. It’s how these systems behave.\n\nWith Celery, for example, tasks can be redelivered if the worker dies before acknowledgment. Features like `acks_late` help, but they do not magically make your side effects safe.\n\nThe real fix is application-level:\n\nStripe’s webhook model is basically a giant sign that says: expect duplicates, expect retries, design accordingly.\n\nIf your endpoint is down, Stripe retries undelivered events for up to 3 days.\n\nYou can also recover manually by listing undelivered events.\n\nExample:\n\n```\ncurl -G https://api.stripe.com/v1/events \\\n  -u \"sk_test_...:\" \\\n  -d ending_before=evt_001 \\\n  -d \"types[]=payment_intent.succeeded\" \\\n  -d \"types[]=payment_intent.payment_failed\" \\\n  -d delivery_success=false\n```\n\nThe important part is not the curl command.\n\nThe important part is that your system has to know whether it already processed `evt_001`.\n\nA simple pattern is a table like this:\n\n```\nCREATE TABLE processed_events (\n  event_id TEXT PRIMARY KEY,\n  event_type TEXT NOT NULL,\n  processed_at TIMESTAMP NOT NULL DEFAULT NOW()\n);\n```\n\nThen your worker logic becomes:\n\n``` js\nasync function handleStripeEvent(event) {\n  const alreadyProcessed = await db.processed_events.findUnique({\n    where: { event_id: event.id }\n  });\n\n  if (alreadyProcessed) {\n    return { ok: true, duplicate: true };\n  }\n\n  await doSideEffects(event);\n\n  await db.processed_events.create({\n    data: {\n      event_id: event.id,\n      event_type: event.type,\n    }\n  });\n\n  return { ok: true };\n}\n```\n\nThat pattern is not glamorous.\n\nIt is also the difference between “retries are safe” and “why did we send 4 receipts?”\n\nYes, use restart policies.\n\nThey help with crashes.\n\n```\ndocker run -d --name worker --restart unless-stopped my-worker-image\n```\n\nBut restart policies only help when the process exits.\n\nMy actual problem was worse:\n\nThat’s a different class of failure.\n\nDocker does not know your Node.js event loop is wedged.\n\nDocker does not know your Python worker is stuck in a bad call.\n\nDocker does not know your queue heartbeat stopped.\n\nThat’s why “container is up” is not enough.\n\nBullMQ has a much more useful idea of health than Docker does.\n\nIt marks jobs as stalled when an active worker stops renewing its lock.\n\nBy default, that check is around 30 seconds.\n\nThat is exactly the kind of signal I needed.\n\nBecause a worker that is still running but hasn’t renewed a lock in 30 seconds is not healthy in any meaningful sense.\n\nA basic BullMQ worker looks like this:\n\n``` python\nimport { Queue, Worker } from 'bullmq';\nimport IORedis from 'ioredis';\n\nconst connection = new IORedis(process.env.REDIS_URL);\n\nconst queue = new Queue('jobs', { connection });\n\nconst worker = new Worker(\n  'jobs',\n  async job => {\n    console.log(`Processing job ${job.id}`);\n\n    // your LLM / automation work here\n    return await runTask(job.data);\n  },\n  { connection }\n);\n\nworker.on('completed', job => {\n  console.log(`Job ${job.id} completed`);\n});\n\nworker.on('failed', (job, err) => {\n  console.error(`Job ${job?.id} failed`, err);\n});\n```\n\nAnd a retry config with jitter should be your default, not an afterthought:\n\n```\nawait queue.add('test-retry', { foo: 'bar' }, {\n  attempts: 8,\n  backoff: {\n    type: 'fixed',\n    delay: 1000,\n    jitter: 0.5,\n  },\n});\n```\n\nThat `jitter: 0.5` matters a lot.\n\nWithout jitter, a worker fleet tends to fail together and retry together.\n\nThat’s how a small outage turns into a retry storm.\n\nThis one bit me hard.\n\nIf you run CPU-heavy work in the same Node.js process that needs to keep queue heartbeats alive, you can create stalls without crashing anything.\n\nExamples:\n\nThe worker process is technically alive.\n\nBut queue progress says otherwise.\n\nThat’s why sandboxed processors or separate worker processes are worth it.\n\nIf a task can block the event loop, isolate it.\n\nThe best answer I found was: use multiple layers, but give each layer one job.\n\n| Option | What it’s actually good at | \n|---|---|\n| systemd watchdog + service restart | Detects hung processes via heartbeat instead of only exit codes. Great for VM or bare-metal workers. | \n| Docker restart policy | Restarts exited containers. Good default. Weak for hung-but-still-running workers. | \n| BullMQ or Celery | Decouples ingress from execution, supports retries and redelivery, and tolerates worker death if jobs are replay-safe. | \n\nMy ranking for real-world usefulness:\n\nContainer restarts save you from crashes.\n\nQueues save you from reality.\n\nThis was a much bigger deal than I expected.\n\nBefore, different workers were calling different providers directly:\n\nThat gets ugly fast.\n\nEspecially when every layer retries differently.\n\nOpenAI-style clients already need backoff for 429s and transient failures. If your queue retries jobs and your model client also retries aggressively, you can easily create a retry storm from both sides.\n\nPutting all workers behind one OpenAI-compatible endpoint fixed a lot:\n\nThat architecture is a great fit for AI agents and automations.\n\nIt also makes provider swaps much less painful.\n\nIf you have:\n\n...then the model layer should not be reimplemented in every service.\n\nA single OpenAI-compatible endpoint gives you a clean boundary.\n\nThis is exactly why products like Standard Compute are interesting for this setup.\n\nYou point your workers at one endpoint, keep your existing OpenAI-compatible SDKs, and avoid wiring provider-specific behavior into every queue consumer.\n\nThe other practical benefit is cost predictability.\n\nWhen agents run 24/7, per-token billing gets annoying fast. A flat monthly model is a lot easier to reason about when you’re scaling background jobs, retries, and automations across a worker fleet.\n\nFor this kind of architecture, that matters more than people admit.\n\nA single endpoint also centralizes failure.\n\nThat is real.\n\nIf your gateway is down or misconfigured, every worker feels it.\n\nSo if you do this, you still need:\n\nStill, I would take one well-observed gateway over a zoo of half-maintained provider clients every time.\n\nEspecially for teams that want interchangeable workers.\n\nIf your setup is tiny, don’t cargo-cult a distributed system.\n\n...then a single process with systemd, strict timeouts, and a dead-letter path may be enough.\n\nYou do not need Redis, BullMQ, PostgreSQL, S3, and six workers to rename files once an hour.\n\nBut once you add:\n\n...the “simple” setup stops being simple.\n\nIt becomes fragile.\n\nAnd fragile always looks cheap right before it gets expensive.\n\nIf I were rebuilding a headless AI server setup today, this is the baseline:\n\nIf you want something actionable, start here:\n\n```\n# 1. run Redis\n\ndocker run -d \\\n  --name redis \\\n  --restart unless-stopped \\\n  -p 6379:6379 \\\n  redis\n# 2. split your app into:\n#    - ingress service\n#    - worker service\n#    - shared database\n# 3. add health checks that measure progress, not just process existence\n# 4. make every external side effect idempotent\n# 5. point all workers at one OpenAI-compatible endpoint\n```\n\nIf you do only those five things, your setup will already be less fragile than a surprising number of “production” agent stacks.\n\nThe most reliable headless AI server setup I’ve used is boring on purpose.\n\nNot the fanciest agent framework.\n\nNot the prettiest architecture diagram.\n\nNot the stack with the most logos in it.\n\nJust:\n\nThat setup assumes workers will freeze.\n\nIt assumes webhooks will retry.\n\nIt assumes jobs will replay.\n\nIt assumes somebody will eventually ship CPU-heavy nonsense into a process that was supposed to stay responsive.\n\nThat’s why it works.\n\nAnd if you’re running always-on AI agents, “nothing broke at 3 a.m.” is a much better definition of success than “the container was still up.”", "url": "https://wpnews.pro/news/my-headless-ai-server-setup-stopped-dying-when-i-got-boring-about-queues-and-30", "canonical_source": "https://dev.to/lars_winstand/my-headless-ai-server-setup-stopped-dying-when-i-got-boring-about-queues-and-30-second-watchdogs-3im5", "published_at": "2026-09-29 14:13:11+00:00", "updated_at": "2026-09-29 14:16:43.312469+00:00", "lang": "en", "topics": ["ai-agents", "ai-infrastructure", "mlops", "developer-tools"], "entities": ["Stripe", "Redis", "Celery", "n8n", "Docker", "Claude", "OpenAI", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/my-headless-ai-server-setup-stopped-dying-when-i-got-boring-about-queues-and-30", "markdown": "https://wpnews.pro/news/my-headless-ai-server-setup-stopped-dying-when-i-got-boring-about-queues-and-30.md", "text": "https://wpnews.pro/news/my-headless-ai-server-setup-stopped-dying-when-i-got-boring-about-queues-and-30.txt", "jsonld": "https://wpnews.pro/news/my-headless-ai-server-setup-stopped-dying-when-i-got-boring-about-queues-and-30.jsonld"}}