Stream AI answers word by word in 15 lines A tutorial published on markodenic.tech shows developers how to stream AI-generated answers word by word in roughly 15 lines of browser JavaScript, using fetch with response.body.pipeThrough(new TextDecoderStream()).getReader() instead of waiting for response.json(). The approach requires the server endpoint to forward the AI provider's streaming SDK chunks as plain text via a ReadableStream, and the post notes that an AbortController can cancel a request while a reverse proxy may need the header 'X-Accel-Buffering': 'no' to stop buffering the response. Stream AI answers word by word in 15 lines Set https://www.google.com/preferences/source?q=markodenic.tech markodenic.tech as your preferred Google source Sponsor this newsletter · reach 9,500+ active developers https://markodenic.tech/sponsorship/ You add an “Ask AI” box to a client’s site. The user types a question, clicks the button, and stares at a spinner. Eight seconds later, a wall of text appears all at once. Half of them already clicked again, or left. The model was writing the answer the whole time. Your code just waited for the last word before showing the first one. Streaming fixes that: the answer appears word by word as it is generated, the way AI chat apps do it. The fix js async function ask question { const response = await fetch '/api/ask', { method: 'POST', body: JSON.stringify { question } , } ; const reader = response.body.pipeThrough new TextDecoderStream .getReader ; output.textContent = ''; while true { const { value, done } = await reader.read ; if done break; output.textContent += value; } } response.body is a stream. Instead of waiting for response.json , you read it chunk by chunk, and TextDecoderStream turns the raw bytes into text. Every chunk goes on screen the moment it arrives. Q: Why should I stream AI answers? Click both buttons and compare how long you wait before you can start reading. Your endpoint has to stream too The browser can only show what the server sends. Every major AI provider’s SDK has a streaming mode that hands you the answer piece by piece. Forward those pieces as plain text: export async function POST request { const { question } = await request.json ; const chunks = await askModel question ; // your AI provider's streaming call const body = new ReadableStream { async start controller { for await const text of chunks { controller.enqueue new TextEncoder .encode text ; } controller.close ; }, } ; return new Response body, { headers: { 'Content-Type': 'text/plain; charset=utf-8' } } ; } askModel stands for your provider’s SDK call with streaming turned on. Check its docs for the exact shape of each chunk and pull out the text. The handler uses the standard Request and Response objects, so it runs in most modern server runtimes and frameworks. The API key stays on the server, and the browser receives plain text. Switch providers and only this file changes. The 15 lines in the browser stay exactly the same. Let users stop it A long answer to the wrong question is worth cancelling. Pass a signal to fetch : js const controller = new AbortController ; stopButton.addEventListener 'click', = controller.abort ; const response = await fetch '/api/ask', { method: 'POST', body: JSON.stringify { question } , signal: controller.signal, } ; Aborting rejects the fetch or the pending reader.read with an AbortError , depending on when the user clicks, so wrap both in try/catch . Aborting stops the answer in the browser right away. To also stop paying for tokens nobody will read, your server has to notice the closed connection and cancel its own request to the AI provider. When it still arrives all at once If your code is right but the text still shows up in one piece, something between the server and the browser is buffering it. The usual suspect is a reverse proxy in front of your app. Check its buffering settings. On a very common setup, this response header turns buffering off for just this response: headers: { 'Content-Type': 'text/plain; charset=utf-8', 'X-Accel-Buffering': 'no', } Where this comes up - “Ask AI” boxes and chat widgets: the obvious one. - “Explain this” or “Summarize” buttons: anywhere a user waits for generated text. - Admin tools: AI-written product descriptions, email drafts, alt text. - Anything slow that produces text: streaming is not AI-specific. Large exports and log viewers work the same way. Why this matters - It feels fast without being faster: the total time is the same, but reading starts almost immediately. - Nobody double-clicks: visible progress stops the “is it broken?” retries that double your API bill. - It works everywhere: fetch streaming and TextDecoderStream are supported in every modern browser. No library, no WebSockets. Find the await response.json behind your AI feature and swap it for the loop above. Happy coding Marko