A token is not a message: stop storing AI responses like event logs Storing streamed AI responses as individual tokens rather than as messages forces every late-joining client — a refreshed browser, a second device, or a human agent taking over a support conversation — to rebuild the response itself, while rewriting a stored message for every token adds a database write per token, according to an analysis by Ably. Ably argues that message appends is the only pattern that delivers the full response so far to a late joiner in one request and then streams the rest live without a database write per token. Colin Kennedy, Principal Product Engineer at Fin (formerly Intercom), said "A dropped connection can mean that users lose entire AI responses with detailed information. How you store a streamed AI response decides what a client sees when the client joins partway through. A human agent taking over a support conversation, a refreshed browser, and a second device all need the response generated so far, and depending on how the response was stored, that late joiner gets the full response, fragments, or nothing. This article compares four common ways to store a streamed response, what each one gives a late joiner, and what each one costs your team. Key takeaways - To give a late joiner such as a refreshed browser, a second device, or a human agent taking over the response so far, store the message, not the tokens. Storing nothing leaves a late joiner with nothing. Storing tokens makes every client rebuild the response. Rewriting a stored message for every token loads your database with a write per token. - Message appends is the only pattern that gives a late joiner the full response so far in one request, then the rest live, without a database write for every token. - Building message appends yourself means building and running a distributed append store. With Ably, your team writes only short client code. The late-joiner test A late joiner is any client that connects after a response has started: a refreshed browser, a second device, a human agent taking over, or a supervisor opening a live conversation. A reconnecting client is also a late joiner https://ably.com/blog/token-streaming-for-ai-ux , because a reconnecting client cannot be sure which fragments arrived before the connection dropped. Late joining happens every time a user refreshes, switches device, or a conversation passes to a human agent. A storage pattern passes the late-joiner test when a client that connects mid-response can get three things: 1. The full response generated so far, as a single piece of content. 2. The current status of the response: still being generated, completed, or cancelled. 3. The rest of the response, delivered live, with no gap and no duplicated text. A storage pattern that fails the late-joiner test causes one or more of three problems: - Users lose the response. A refreshed browser shows a blank or restarted response. - Handoffs start blind. In a customer-support product, a conversation can escalate to a human agent while the AI is still responding https://ably.com/blog/ai-agent-streaming-in-action . If the agent console cannot show what the AI has already said, the customer repeats the details and the handoff takes longer. - Engineering cost multiplies. Every client that reads responses web, mobile, agent console, analytics needs its own logic to rebuild the response. Every copy of the rebuilding logic is another place where duplicated or missing text can appear. Colin Kennedy, Principal Product Engineer at Fin formerly Intercom https://ably.com/case-studies/fin-intercom , describes the stakes for AI support products: "A dropped connection can mean that users lose entire AI responses with detailed information." Each pattern section below covers how the pattern works, what a late joiner receives, and what your team must build on top of the pattern to pass the late-joiner test. Each section ends with a verdict: whether the pattern stores the message, the tokens, or nothing, and what that choice costs. The section after the four patterns brings the verdicts together and shows where each pattern still fits. Pattern 1: Ephemeral delivery over SSE or WebSockets How it works. The server forwards each token the smallest unit of text a model outputs, often part of a word to the connected client over SSE or WebSockets https://ably.com/topic/ai-stack/websockets-vs-sse-ai-chat , and the browser joins the tokens together in memory. The browser's memory is the only place the assembled response exists. What a late joiner receives. Nothing that was generated before the late joiner connected, because the server forwards each token once and keeps no copy. A refresh erases the browser's copy of the response, and a human agent taking over has no copy to read. If the application regenerates the response instead, then the team pays for a second model call, and the new response may not match the one the user was reading. What your team must build on top of ephemeral delivery to pass the late-joiner test. Your team must build a separate store for partial responses, plus logic that keeps the store consistent with the live stream. SSE's Last-Event-ID header https://ably.com/blog/resume-tokens-last-event-id-llm-streaming-reconnection can resume a stream only if the server keeps earlier events, and a server that keeps earlier events has become a chunk log Pattern 2 . The Vercel AI SDK's resumable streams https://github.com/vercel/resumable-stream , for example, keep earlier events in the memory of the server producing the stream, and use Redis to route a reconnecting client to that server, so the events are lost if that server stops. Storing only the completed response does not help, because late joiners then see nothing until the model finishes the response. Verdict. Ephemeral delivery stores nothing, which makes it a transport rather than a storage pattern. Without storage, a late joiner gets nothing, and adding a Redis store beside the stream turns the setup into a chunk log Pattern 2 . Pattern 2: Chunk logs Redis Streams, NATS JetStream, Electric Durable Streams How it works. Your backend, the service that calls the model, appends each chunk of output one token or a small group of tokens to a durable, append-only log the chunk log , where every position has an ID or offset. Because the log keeps every chunk and records its position, a client that disconnects can resume from the last position it received, without asking the model to generate the response again. Surviving a disconnect this way is the main advantage a chunk log has over ephemeral delivery. What a late joiner receives. A list of fragments, not a readable response, because the log stores each chunk as a separate entry and never joins the entries together. To show the response so far, a refreshed browser must find where the response starts in the log, read every entry since then, join the entries together in order, and watch for a completion event. A long response can mean replaying hundreds or thousands of entries before the browser renders the first word. Chunk logs are usually set to delete old entries after a fixed time or once the log reaches a maximum length. If the log has already deleted the start of a response, the browser cannot rebuild the response at all. What your team must build on top of a chunk log to pass the late-joiner test. A chunk log fails the late-joiner test because the log holds fragments rather than a readable response. To pass, every consumer each browser, app, or service that reads the log needs reassembly logic. Electric, for example, ships client libraries for Durable Streams https://electric.ax/docs/streams/clients/typescript that rebuild messages from stored chunks, which saves writing the reassembly logic but still runs it in every client. The alternative is to run a projection service: a separate backend service that reads the log, rebuilds each response, and stores the latest version for clients to read. A projection service adds a second stored copy of every response, with its own synchronization and recovery work. The reassembly logic, whether it runs in each consumer or in a projection service, also has to handle duplicates. After a failure, some chunk logs deliver the same chunk a second time, and reassembly logic that adds every chunk it receives would then show the same words twice. Redis Streams consumer groups, for instance, can pass a chunk https://redis.io/docs/latest/commands/xclaim/ that one reader never confirmed to a second reader, and NATS JetStream redelivers https://docs.nats.io/learn/jetstream/acknowledgment the-server-controls any chunk whose receipt is not confirmed in time. Whichever chunk log you use, the reassembly logic must recognize a chunk it has already added and skip it. Verdict. A chunk log stores the tokens, not the message. A chunk log solves disconnects but multiplies engineering cost: every client, or a projection service you run, must rebuild the response, and every copy of that logic is another place for duplicated or missing text to appear. Pattern 3: Repeated record updates Postgres, Supabase, Firestore, Sendbird How it works. Your backend stores each response as one database row or chat message and rewrites the stored content as new tokens arrive. Postgres, and Supabase on top of Postgres, run each rewrite as a transaction that writes a new version of the row https://www.postgresql.org/docs/current/routine-vacuuming.html VACUUM-FOR-SPACE-RECOVERY . Firestore behaves the same way when your backend updates a document for each new token, and chat APIs such as Sendbird offer an update-message endpoint for the same purpose. A variant of this pattern, used by Stream https://getstream.io/chat/docs/node/ai-message-streaming/ , keeps the in-progress rewrites out of the database. Each rewrite reaches every watching client as an in-memory update, and only the finished response is saved. This avoids per-token writes, but a client that loads the conversation mid-response finds an empty placeholder until the next in-memory update arrives, and the partial response is lost if your backend fails. What a late joiner receives. A readable response, possibly slightly behind the live stream. A late joiner reads one row and gets the whole response so far in a single request, which makes repeated record updates the first pattern in this article to pass the late-joiner test. But the stored row can trail the live stream by any tokens not yet written to the database. What your team must build on top of repeated record updates to keep passing the late-joiner test. Repeated record updates pass the late-joiner test only while the database keeps up with a write for every new token. Write volume grows with the number of responses streaming at once: 1,000 responses streaming at 100 tokens per second means 100,000 row updates per second, so the database must be sized for peak streaming traffic. To avoid sizing the database for peaks, your team must build write-load control: logic that limits how often each response is written, usually by batching tokens. Batching buffers several tokens and writes them together, which reduces the write rate but adds lag. Flushing every 20 tokens cuts writes twentyfold, but a human agent taking over then reads text up to 20 tokens about 200ms at 100 tokens per second behind what the customer sees. On hosted chat APIs, the platform also limits writes: Sendbird rate-limits its Platform API https://sendbird.com/docs/chat/platform-api/v3/application/understanding-rate-limits/rate-limits 2-plan-based-rate-limits per endpoint and per plan, which caps how often one message can be rewritten and forces larger batches with longer lag. Firestore's documentation warns https://firebase.google.com/docs/firestore/best-practices updates to a single document that a single document cannot be updated at an unlimited rate, and that high write rates cause contention, higher latency, or errors. Every write also needs a version check, so that a delayed retry cannot overwrite newer text or mark a response complete before the final content lands. Verdict. Repeated record updates store the message, but at the cost of a database write for every token. You either size your database for streaming peaks, or batch the writes and accept that a human agent taking over sees text slightly behind the customer. The Stream variant cuts the writes, but gives up the stored partial response. Pattern 4: Message appends Ably How it works. Message appends treat one AI response as one message that grows. When the model starts generating a response, your backend creates a message on Ably and receives a serial, a unique ID for that message. For each new token, your backend sends an append: a small update that tells Ably to add the token to the end of the message with that serial. Ably joins the appends together as they arrive, which is why no client has to rebuild the response from stored fragments. In practice, Ably holds one up-to-date copy of each response. Your backend only sends small appends and never rewrites the full text, and your database takes no writes while the response streams. When the model finishes, your backend saves the completed response to your database in one write. Because the response lives on Ably rather than in your backend's memory, the partial response survives if your backend fails mid-response. Each conversation's messages travel on an Ably channel, and the channel keeps a history: the record of past messages that any client can read back after connecting, which is where a late joiner finds the response so far. What a late joiner receives. The full response so far, as one message. With Ably, a client that joins the channel mid-response with rewind an option that replays recent messages as the client joins , or reads channel history, receives the accumulated content as one message, then each later append as the append arrives. The message also carries the status of the response https://ably.com/docs/ai-transport/streaming/token-streaming streaming, complete, or cancelled when you use Ably AI Transport https://ably.com/ai-transport , Ably's SDK for streaming AI responses, so the late joiner knows whether the response is still being generated. Live delivery and history refer to the same message serial, so the last update a client receives matches what history eventually returns. What your team must build on top of message appends to pass the late-joiner test. Your team only builds client code, because Ably assembles the message on the server. If you use Ably's core Pub/Sub SDKs directly, the client code must handle two kinds of message: - Appends: carry only the new text. The client adds that text to the end of the response it is showing. Most messages are appends. - Updates: carry the whole response so far. The client shows the update in place of what it had. Ably sends an update when a client may be missing earlier text, for example, as the first message after the client joins the channel or reconnects, and can send one at other times too. Because either kind can arrive at any time, the client code checks whether each message is an append or an update before showing it. The Ably AI Transport SDK handles both kinds for you, and the client example in the token streaming docs https://ably.com/docs/ai-transport/streaming/token-streaming is about six lines. Verdict. Message appends store the message, with no rebuilding in clients and no per-token database writes. A late joiner gets the full response so far in one request, and your database takes one write per completed response. The verdict: store the message, not the tokens Line the four patterns up, and the difference is what each one stores: nothing, tokens, a message rewritten at token rate, or a message that grows. The table sets the four patterns side by side, with the situation where each one still fits in the last column. Only repeated record updates Pattern 3 and message appends Pattern 4 store a readable partial response while the model is still generating, and of those two, only message appends avoids per-token database writes. Ephemeral delivery and chunk logs can still be the right choice, but only when responses are short, or every client reads each response from the start. Can I build message appends myself? Yes, but building message appends means building a distributed append store, not adding a table or a Redis key. The append store must get four things right at token rate, through node failures, retries, and reconnects: 1. Ordering every append. One component, often called a sequencer, must give each append a position, survive failover without reusing or skipping a position, and recognize retries. When this breaks, users see words out of order or the same sentence twice. 2. Keeping a readable copy without write amplification. Each append should store only the new token, not rewrite the whole response, yet a client must still get the full response in one request. When the readable copy falls behind, users who refresh see a shorter response than other clients show. 3. Handing late joiners from history to the live stream. The stored copy must end exactly where the live stream begins, after every reconnect. When the handover breaks, users see text disappear after they refresh, or appear twice. 4. Deciding how each response ends. When the completion signal, the final append, and a user's cancel arrive at the same moment, the append store must pick one outcome and keep to it. When the wrong outcome is chosen, users see a response stuck on "generating", or a response missing its last sentence. Once built, the append store sits on the path of every AI response your product sends, so your team also owns capacity planning, failover, monitoring, and on-call. Every new feature, such as cancellation or human handoff, then depends on the append store. If the team does not fully trust that foundation, each change feels risky, and the realtime layer becomes the part of the system nobody wants to touch. That was the situation at Doxy.me https://ably.com/case-studies/doxyme before the team moved to Ably. Ben Anderson-Dukes, VP of Engineering, said: "Our engineers were so afraid to touch this part of the stack. It had become this scary black box that was like a house of cards." Why teams use Ably for message appends Message appends is the only pattern in this article that gives a late joiner everything the late joiner needs: the full response so far in one request, the status of the response, and the rest of the response live, with no gaps or duplicates. A human agent taking over a support conversation sees exactly what the AI has already told the customer, on any device. You can build message appends yourself, but as the previous section shows, that means building and running a distributed append store that has to keep ordering, storage, handover, and completion correct at token rate, through every failure. Ably takes on that work by running message appends on Ably channels. Ably also guarantees message order from a single publisher https://ably.com/docs/platform/architecture/message-ordering and is designed for 99.999% global service availability https://ably.com/docs/platform/architecture . AI Transport wraps the create and append calls and adds cancellation and recovery. Text is appended as the model streams it, and tool calls and other events travel as separate messages on the same channel. The token streaming docs https://ably.com/docs/ai-transport/streaming/token-streaming show the code and setup. A fast model can stream more than 150 tokens a second. Ably publishes the first token of a response immediately, then rolls up later appends within a 40-millisecond window into a single message, so by default each response produces at most 25 messages a second, however fast the model runs. You can set the window anywhere from 0 to 500 milliseconds. Because Ably bills per message rather than per token, that cap also keeps the cost of streaming a response predictable AI Transport pricing docs https://ably.com/docs/ai-transport/pricing . A few practical points are worth planning for: - Enabling appends stores every message on the channel namespace https://ably.com/docs/ai-transport/setup/channel-rules , so keep AI responses in a dedicated namespace. - Ably applies appends in the order a single publisher sends them https://ably.com/docs/messages/updates-deletes , so each response should come from one publisher. If several workers contribute to a response, route their output through a single component. - Your longest responses should fit within Ably's message size limit https://ably.com/docs/platform/pricing/limits message-limits- : 64 KiB on Free and Standard packages, 256 KiB on Pro and Enterprise. - The AI Transport SDK is JavaScript and React today. Backends in other languages can use appendMessage https://ably.com/docs/pub-sub/api/javascript/realtime/realtime-channel append-to-a-message- in Ably's core Pub/Sub SDKs. To see message appends in action, follow the AI Transport token streaming guide https://ably.com/docs/ai-transport/streaming/token-streaming on a free Ably account https://ably.com/sign-up . If you are building an AI customer-support product and want to review your architecture, talk to Ably's engineering team https://ably.com/contact . FAQs Is Redis Streams enough on its own for storing AI responses? Redis Streams stores and replays chunks, but Redis Streams does not return an assembled response. Without a projection service https://ably.com/blog/ai-chat-stream-resumption , every consumer must order, deduplicate, and join chunks together, then detect completion. Does message appends replace SSE or WebSockets? No. Appends are a data model, and SSE and WebSockets are transports. Ably delivers appends over WebSockets, with fallback to HTTP on networks that block WebSockets. If I use message appends, do I still need my own database for conversation history? Yes, in most production products. Ably keeps stored messages for 24 hours on the Free package https://ably.com/docs/storage-history/storage , and for 72 hours by default on paid packages, configurable up to 30 days on Standard, 365 days on Pro, and custom periods on Enterprise. Conversations that must last longer, or that feed search, analytics, or reporting, need your own database, which takes one write per completed response. Is Ably message appends production-ready? The AI Transport roadmap https://ably.com/docs/ai-transport/roadmap lists resumable token streaming, which is built on message appends, as available today. The Ably docs page that covers message updates, deletes, and appends carries a Public Preview label https://ably.com/docs/messages/updates-deletes , so check the docs for the current status of each SDK you ship. I already stream AI responses over SSE. What changes if I move to message appends? Your backend publishes each response to an Ably channel as appends instead of writing to an SSE stream, and clients subscribe to the channel instead of holding an HTTP request open. If you use the Vercel AI SDK, AI Transport replaces only the transport that useChat uses https://ably.com/docs/ai-transport/frameworks/vercel-ai-sdk-ui , so the rest of your AI SDK code stays in place. How does Ably control who can read a conversation? Your server signs a short-lived JWT for each user https://ably.com/docs/ai-transport/setup/authentication , scoped to the channels that user can reach, such as the user's own conversations. Ably SDKs use TLS by default, and Ably is SOC 2 Type II certified and HIPAA compliant https://ably.com/docs/ai-transport .