{"slug": "our-chatbot-not-replying-in-groups-turned-out-to-be-a-slack-plumbing-bug-not-an", "title": "Our chatbot not replying in groups turned out to be a Slack plumbing bug, not an AI bug", "summary": "A developer found that their Slack chatbot's silence in group channels was caused by Slack event-handling plumbing, not the underlying AI model. The bot worked in direct messages but failed in channels because DMs arrive as `message` events with `channel_type=\"im\"` while channel mentions arrive as separate `app_mention` events, and because slow model and tool calls exceeded Slack's roughly 3-second acknowledgment window, triggering retries and duplicate processing. The fix was to ack immediately and move slow work into a background queue with deduplication by event_id.", "body_md": "If your bot works in DMs but goes weirdly silent in Slack channels, start by blaming your event handling.\n\nNot GPT-5.\n\nNot Claude.\n\nNot your prompt.\n\nWe lost a bunch of time learning this the dumb way.\n\nOur bot was great in 1:1 chats. In Slack DMs it could hold context, call tools, summarize results, and generally look like a competent AI teammate.\n\nThen we dropped the exact same bot into a busy channel.\n\nIt became a ghost.\n\nSometimes it replied.\n\nSometimes it replied twice.\n\nSometimes the backend finished successfully and Slack showed nothing.\n\nSometimes it answered in the wrong thread.\n\nThat kind of failure is extra annoying because it makes you debug the wrong layer first.\n\nWe blamed the model.\n\nWe swapped GPT-5 for Claude.\n\nWe trimmed prompts.\n\nWe argued about whether Llama or Qwen would be more reliable in channels.\n\nNone of that mattered.\n\nThe real problem was that we treated Slack group chats like DMs, and Slack absolutely does not work that way.\n\nThis is the first thing I’d check in any Slack bot.\n\nA direct message to your app comes in as a `message` event with `channel_type=\"im\"`.\n\nA mention in a channel comes in as `app_mention`.\n\nThose are different event streams with different behavior.\n\nDM example:\n\n```\n{\n  \"event\": {\n    \"type\": \"message\",\n    \"channel_type\": \"im\",\n    \"channel\": \"D024BE91L\",\n    \"text\": \"Hello hello can you hear me?\"\n  }\n}\n```\n\nChannel mention example:\n\n```\n{\n  \"event\": {\n    \"type\": \"app_mention\",\n    \"channel\": \"C123ABC456\",\n    \"text\": \"<@U0LAN0Z89> is it everything a river should be?\",\n    \"ts\": \"1515449522.000016\"\n  }\n}\n```\n\nThat one distinction explains a lot of “works in DMs, fails in channels” bugs.\n\nIn DMs, the shape is simple: user talks to bot.\n\nIn channels, everything gets more fragile:\n\nIf your code has one generic `handleIncomingMessage()` path for everything, there’s a decent chance that’s your bug.\n\nQuiet channels hide bad architecture.\n\nIf your bot mostly lives in DMs or low-traffic channels, a synchronous handler can look fine for weeks.\n\nThen one incident channel gets busy.\n\nFive people mention the bot.\n\nOne tool call takes 8 seconds.\n\nOne GitHub request stalls.\n\nOne retry comes in.\n\nAnd suddenly your “AI reliability issue” is obviously just broken event lifecycle handling.\n\nSlack expects a fast HTTP 200 acknowledgment.\n\nThe practical rule is simple:\n\n**you have about 3 seconds to ack**\n\nIf you don’t ack quickly, Slack may retry the event, and now you’re in duplicate-processing land.\n\nThis is the dangerous flow:\n\n`app_mention`\nThat feels natural when you first build it.\n\nIt’s also how you build a bot that disappears under load.\n\nWe moved all slow work out of the request path.\n\nThat means:\n\nHere’s the shape in Slack Bolt for Python:\n\n``` python\nfrom slack_bolt import App\n\napp = App(process_before_response=True)\n\ndef ack_fast(ack):\n    ack(\"Accepted\")\n\ndef run_long_process(respond, body, logger):\n    user_text = body[\"event\"][\"text\"]\n\n    # slow work goes here\n    # call model\n    # call tools\n    # build final answer\n\n    respond(\"Completed\")\n```\n\nThat’s the idea, but in production I’d usually push this into a queue instead of doing everything inside the listener.\n\nSomething more like this:\n\n``` python\nfrom fastapi import FastAPI, Request, BackgroundTasks\nimport os\nimport json\n\napp = FastAPI()\n\n@app.post(\"/slack/events\")\nasync def slack_events(req: Request, background_tasks: BackgroundTasks):\n    payload = await req.json()\n\n    # Slack URL verification\n    if payload.get(\"type\") == \"url_verification\":\n        return {\"challenge\": payload[\"challenge\"]}\n\n    event = payload.get(\"event\", {})\n    event_id = payload.get(\"event_id\")\n\n    # 1. dedupe by event_id\n    # 2. enqueue background work\n    background_tasks.add_task(process_event, event_id, event)\n\n    # ack immediately\n    return {\"ok\": True}\n\ndef process_event(event_id: str, event: dict):\n    # idempotency check here\n    # model call here\n    # tool calls here\n    # post reply here\n    pass\n```\n\nThe important part is not the framework.\n\nThe important part is the split:\n\nReliable Slack bots are event systems first and AI apps second.\n\nThis one was nastier.\n\nWe assumed that if our worker finished and called `chat.postMessage`, the message would show up.\n\nThat is not safe in busy channels.\n\nSlack rate-limits `chat.postMessage` at roughly **1 message per second per channel**.\n\nPer channel.\n\nThat means if your bot posts:\n\nand three people mention it in the same thread at once, you’ve basically built a rate-limit machine.\n\nThe ugly part is that your logs can still look fine while users see silence.\n\nWe stopped treating outbound messages like free writes.\n\nThe new rules were:\n\nThat improved reliability more than any prompt tweak.\n\nNot because the model got smarter.\n\nBecause the plumbing stopped fighting the app.\n\nIn DMs, state is easy.\n\nIn channels, state is thread-shaped.\n\nIf you reply without the right `thread_ts`, your answer lands in the wrong place or becomes useless noise in the channel.\n\nA lot of teams do this:\n\n`conversations.replies`\nThat works at first.\n\nIt’s also expensive, slow, and increasingly brittle.\n\nA better pattern is to keep your own compact thread state.\n\nStore:\n\n`channel`` thread_ts`\nUse Slack history as recovery, not as your primary memory system.\n\nHere’s the tradeoff:\n\n| Approach | What actually happens | \n|---|---|\n| Rebuild thread from Slack every turn | Easy to start, but slower, noisier, and more fragile under load | \n| Maintain your own thread state | Slightly more engineering, much more reliable for long-running agents | \n\nIf you’re running workflows in n8n, Make, Zapier, OpenClaw, or your own worker stack, this matters even more.\n\nThose tools make it easy to ship a bot quickly.\n\nThey also make it easy to hide state bugs until traffic spikes.\n\nIf I were rebuilding this from scratch, I’d do it like this.\n\nHandle DMs and mentions separately.\n\n``` python\ndef route_event(event: dict):\n    event_type = event.get(\"type\")\n    channel_type = event.get(\"channel_type\")\n\n    if event_type == \"message\" and channel_type == \"im\":\n        return \"dm\"\n\n    if event_type == \"app_mention\":\n        return \"channel_mention\"\n\n    return \"ignore\"\n```\n\nDifferent paths should have different logic for:\n\nAnything slow goes to background work.\n\nThat includes:\n\nYour webhook should be boring.\n\nBoring is good.\n\nSlack retries are not edge cases.\n\nStore `event_id` and make processing idempotent.\n\nPseudo-code:\n\n``` python\ndef process_event(event_id: str, event: dict):\n    if already_processed(event_id):\n        return\n\n    mark_processing(event_id)\n\n    try:\n        handle_event(event)\n        mark_processed(event_id)\n    except Exception:\n        mark_failed(event_id)\n        raise\n```\n\nIf you skip this, duplicate replies are just a matter of time.\n\nOne queue per channel is a sane default.\n\n``` python\nfrom collections import defaultdict\nfrom queue import Queue\nimport time\n\nchannel_queues = defaultdict(Queue)\n\ndef post_with_throttle(channel_id: str, message: dict):\n    q = channel_queues[channel_id]\n    q.put(message)\n\n    while not q.empty():\n        next_msg = q.get()\n        slack_client.chat_postMessage(**next_msg)\n        time.sleep(1.0)\n```\n\nIn real code you’d use a proper worker, lock, or async queue, but the design point stands:\n\n**throttle per channel, not globally**\n\nIf you want progress updates, update one message instead of posting five new ones.\n\nThat usually looks cleaner for users and is friendlier to rate limits.\n\nSlack is a transport layer.\n\nIt should not be your source of truth.\n\nMinimal thread record example:\n\n```\n{\n  \"channel\": \"C123ABC456\",\n  \"thread_ts\": \"1712345678.123456\",\n  \"participants\": [\"U111\", \"U222\"],\n  \"messages\": [\n    {\"role\": \"user\", \"text\": \"check prod errors\"},\n    {\"role\": \"assistant\", \"text\": \"Looking into GitHub and Datadog\"}\n  ],\n  \"summary\": \"Investigating production error spike after deploy\",\n  \"last_updated\": \"2026-10-10T12:00:00Z\"\n}\n```\n\nThat gives you a stable context source for long-running agents.\n\nThe labels differ, but the architecture lesson is the same.\n\nSlack has:\n\n`chat.postMessage`\nDiscord has:\n\nDifferent APIs, same rule:\n\n**respond fast, queue slow work, dedupe retries, and control outbound writes**\n\nIf you do long-running agent work inline, both platforms will punish you.\n\nWhen a bot works in DMs but not in channels, this is the checklist I’d run in order.\n\nLog the event shape.\n\n``` python\nimport json\n\ndef debug_event(payload):\n    print(json.dumps(payload, indent=2))\n```\n\nCheck whether you’re actually receiving `app_mention` and not assuming all traffic is `message`.\n\nMake sure the app is actually in the channel and has the scopes it needs.\n\nLog how long your webhook takes before returning 200.\n\n``` python\nimport time\n\nstart = time.time()\n# validate + enqueue\nelapsed = time.time() - start\nprint(f\"ack path took {elapsed:.3f}s\")\n```\n\nIf that number is drifting upward, you’re moving work back into the request path.\n\nLog `event_id` and retry headers.\n\nIf duplicate replies exist, you probably don’t have idempotency under control.\n\nMake sure replies use the correct `thread_ts`.\n\nIf the bot is chatty in one channel, assume rate limits are part of the problem until proven otherwise.\n\nThis bug looked like a model problem because model calls were the most visible slow step.\n\nThat’s common in agent systems.\n\nPeople blame GPT-5, Claude, Grok, or tool-calling reliability when the real issue is orchestration:\n\nThat’s also why predictable AI infrastructure matters.\n\nWhen you’re building bots, automations, or long-running agents, you want the freedom to offload the slow work without obsessing over every token or every retry path.\n\nThat’s the appeal of Standard Compute: you can keep your app OpenAI-compatible, swap in a flat-rate endpoint, and let agents run in the background without per-token cost anxiety creeping into every architecture decision.\n\nEspecially if you’re wiring together Slack, GitHub, Jira, Notion, n8n, Make, Zapier, or custom workers, the expensive part is usually not one single prompt. It’s the whole loop.\n\nOur bot never needed a better personality.\n\nIt needed better manners.\n\nThe real fix was not prompt engineering.\n\nIt was:\n\n`message.im` and If your chatbot stops replying in groups while behaving perfectly in DMs, start with the boring plumbing.\n\nThat’s the stuff that actually fixes production bots.", "url": "https://wpnews.pro/news/our-chatbot-not-replying-in-groups-turned-out-to-be-a-slack-plumbing-bug-not-an", "canonical_source": "https://dev.to/lars_winstand/our-chatbot-not-replying-in-groups-turned-out-to-be-a-slack-plumbing-bug-not-an-ai-bug-1dg6", "published_at": "2026-10-10 22:11:44+00:00", "updated_at": "2026-10-10 22:16:06.558451+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools"], "entities": ["Slack", "GPT-5", "Claude", "Llama", "Qwen", "Slack Bolt"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/our-chatbot-not-replying-in-groups-turned-out-to-be-a-slack-plumbing-bug-not-an", "markdown": "https://wpnews.pro/news/our-chatbot-not-replying-in-groups-turned-out-to-be-a-slack-plumbing-bug-not-an.md", "text": "https://wpnews.pro/news/our-chatbot-not-replying-in-groups-turned-out-to-be-a-slack-plumbing-bug-not-an.txt", "jsonld": "https://wpnews.pro/news/our-chatbot-not-replying-in-groups-turned-out-to-be-a-slack-plumbing-bug-not-an.jsonld"}}