Our chatbot not replying in groups turned out to be a Slack plumbing bug, not an AI bug A developer found that their Slack chatbot's silence in group channels was caused by Slack event-handling plumbing, not the underlying AI model. The bot worked in direct messages but failed in channels because DMs arrive as `message` events with `channel_type="im"` while channel mentions arrive as separate `app_mention` events, and because slow model and tool calls exceeded Slack's roughly 3-second acknowledgment window, triggering retries and duplicate processing. The fix was to ack immediately and move slow work into a background queue with deduplication by event_id. If your bot works in DMs but goes weirdly silent in Slack channels, start by blaming your event handling. Not GPT-5. Not Claude. Not your prompt. We lost a bunch of time learning this the dumb way. Our bot was great in 1:1 chats. In Slack DMs it could hold context, call tools, summarize results, and generally look like a competent AI teammate. Then we dropped the exact same bot into a busy channel. It became a ghost. Sometimes it replied. Sometimes it replied twice. Sometimes the backend finished successfully and Slack showed nothing. Sometimes it answered in the wrong thread. That kind of failure is extra annoying because it makes you debug the wrong layer first. We blamed the model. We swapped GPT-5 for Claude. We trimmed prompts. We argued about whether Llama or Qwen would be more reliable in channels. None of that mattered. The real problem was that we treated Slack group chats like DMs, and Slack absolutely does not work that way. This is the first thing I’d check in any Slack bot. A direct message to your app comes in as a message event with channel type="im" . A mention in a channel comes in as app mention . Those are different event streams with different behavior. DM example: { "event": { "type": "message", "channel type": "im", "channel": "D024BE91L", "text": "Hello hello can you hear me?" } } Channel mention example: { "event": { "type": "app mention", "channel": "C123ABC456", "text": "<@U0LAN0Z89 is it everything a river should be?", "ts": "1515449522.000016" } } That one distinction explains a lot of “works in DMs, fails in channels” bugs. In DMs, the shape is simple: user talks to bot. In channels, everything gets more fragile: If your code has one generic handleIncomingMessage path for everything, there’s a decent chance that’s your bug. Quiet channels hide bad architecture. If your bot mostly lives in DMs or low-traffic channels, a synchronous handler can look fine for weeks. Then one incident channel gets busy. Five people mention the bot. One tool call takes 8 seconds. One GitHub request stalls. One retry comes in. And suddenly your “AI reliability issue” is obviously just broken event lifecycle handling. Slack expects a fast HTTP 200 acknowledgment. The practical rule is simple: you have about 3 seconds to ack If you don’t ack quickly, Slack may retry the event, and now you’re in duplicate-processing land. This is the dangerous flow: app mention That feels natural when you first build it. It’s also how you build a bot that disappears under load. We moved all slow work out of the request path. That means: Here’s the shape in Slack Bolt for Python: python from slack bolt import App app = App process before response=True def ack fast ack : ack "Accepted" def run long process respond, body, logger : user text = body "event" "text" slow work goes here call model call tools build final answer respond "Completed" That’s the idea, but in production I’d usually push this into a queue instead of doing everything inside the listener. Something more like this: python from fastapi import FastAPI, Request, BackgroundTasks import os import json app = FastAPI @app.post "/slack/events" async def slack events req: Request, background tasks: BackgroundTasks : payload = await req.json Slack URL verification if payload.get "type" == "url verification": return {"challenge": payload "challenge" } event = payload.get "event", {} event id = payload.get "event id" 1. dedupe by event id 2. enqueue background work background tasks.add task process event, event id, event ack immediately return {"ok": True} def process event event id: str, event: dict : idempotency check here model call here tool calls here post reply here pass The important part is not the framework. The important part is the split: Reliable Slack bots are event systems first and AI apps second. This one was nastier. We assumed that if our worker finished and called chat.postMessage , the message would show up. That is not safe in busy channels. Slack rate-limits chat.postMessage at roughly 1 message per second per channel . Per channel. That means if your bot posts: and three people mention it in the same thread at once, you’ve basically built a rate-limit machine. The ugly part is that your logs can still look fine while users see silence. We stopped treating outbound messages like free writes. The new rules were: That improved reliability more than any prompt tweak. Not because the model got smarter. Because the plumbing stopped fighting the app. In DMs, state is easy. In channels, state is thread-shaped. If you reply without the right thread ts , your answer lands in the wrong place or becomes useless noise in the channel. A lot of teams do this: conversations.replies That works at first. It’s also expensive, slow, and increasingly brittle. A better pattern is to keep your own compact thread state. Store: channel thread ts Use Slack history as recovery, not as your primary memory system. Here’s the tradeoff: | Approach | What actually happens | |---|---| | Rebuild thread from Slack every turn | Easy to start, but slower, noisier, and more fragile under load | | Maintain your own thread state | Slightly more engineering, much more reliable for long-running agents | If you’re running workflows in n8n, Make, Zapier, OpenClaw, or your own worker stack, this matters even more. Those tools make it easy to ship a bot quickly. They also make it easy to hide state bugs until traffic spikes. If I were rebuilding this from scratch, I’d do it like this. Handle DMs and mentions separately. python def route event event: dict : event type = event.get "type" channel type = event.get "channel type" if event type == "message" and channel type == "im": return "dm" if event type == "app mention": return "channel mention" return "ignore" Different paths should have different logic for: Anything slow goes to background work. That includes: Your webhook should be boring. Boring is good. Slack retries are not edge cases. Store event id and make processing idempotent. Pseudo-code: python def process event event id: str, event: dict : if already processed event id : return mark processing event id try: handle event event mark processed event id except Exception: mark failed event id raise If you skip this, duplicate replies are just a matter of time. One queue per channel is a sane default. python from collections import defaultdict from queue import Queue import time channel queues = defaultdict Queue def post with throttle channel id: str, message: dict : q = channel queues channel id q.put message while not q.empty : next msg = q.get slack client.chat postMessage next msg time.sleep 1.0 In real code you’d use a proper worker, lock, or async queue, but the design point stands: throttle per channel, not globally If you want progress updates, update one message instead of posting five new ones. That usually looks cleaner for users and is friendlier to rate limits. Slack is a transport layer. It should not be your source of truth. Minimal thread record example: { "channel": "C123ABC456", "thread ts": "1712345678.123456", "participants": "U111", "U222" , "messages": {"role": "user", "text": "check prod errors"}, {"role": "assistant", "text": "Looking into GitHub and Datadog"} , "summary": "Investigating production error spike after deploy", "last updated": "2026-10-10T12:00:00Z" } That gives you a stable context source for long-running agents. The labels differ, but the architecture lesson is the same. Slack has: chat.postMessage Discord has: Different APIs, same rule: respond fast, queue slow work, dedupe retries, and control outbound writes If you do long-running agent work inline, both platforms will punish you. When a bot works in DMs but not in channels, this is the checklist I’d run in order. Log the event shape. python import json def debug event payload : print json.dumps payload, indent=2 Check whether you’re actually receiving app mention and not assuming all traffic is message . Make sure the app is actually in the channel and has the scopes it needs. Log how long your webhook takes before returning 200. python import time start = time.time validate + enqueue elapsed = time.time - start print f"ack path took {elapsed:.3f}s" If that number is drifting upward, you’re moving work back into the request path. Log event id and retry headers. If duplicate replies exist, you probably don’t have idempotency under control. Make sure replies use the correct thread ts . If the bot is chatty in one channel, assume rate limits are part of the problem until proven otherwise. This bug looked like a model problem because model calls were the most visible slow step. That’s common in agent systems. People blame GPT-5, Claude, Grok, or tool-calling reliability when the real issue is orchestration: That’s also why predictable AI infrastructure matters. When you’re building bots, automations, or long-running agents, you want the freedom to offload the slow work without obsessing over every token or every retry path. That’s the appeal of Standard Compute: you can keep your app OpenAI-compatible, swap in a flat-rate endpoint, and let agents run in the background without per-token cost anxiety creeping into every architecture decision. Especially if you’re wiring together Slack, GitHub, Jira, Notion, n8n, Make, Zapier, or custom workers, the expensive part is usually not one single prompt. It’s the whole loop. Our bot never needed a better personality. It needed better manners. The real fix was not prompt engineering. It was: message.im and If your chatbot stops replying in groups while behaving perfectly in DMs, start with the boring plumbing. That’s the stuff that actually fixes production bots.