markodenic.tech as your preferred Google source Sponsor this newsletter · reach 9,500+ active developers
You add an “Ask AI” box to a client’s site. The user types a question, clicks the button, and stares at a spinner. Eight seconds later, a wall of text appears all at once. Half of them already clicked again, or left.
The model was writing the answer the whole time. Your code just waited for the last word before showing the first one. Streaming fixes that: the answer appears word by word as it is generated, the way AI chat apps do it.
The fix #
async function ask(question) {
const response = await fetch('/api/ask', {
method: 'POST',
body: JSON.stringify({ question }),
});
const reader = response.body.pipeThrough(new TextDecoderStream()).getReader();
output.textContent = '';
while (true) {
const { value, done } = await reader.read();
if (done) break;
output.textContent += value;
}
}
response.body is a stream. Instead of waiting for response.json(), you read it chunk by chunk, and TextDecoderStream turns the raw bytes into text. Every chunk goes on screen the moment it arrives.
Q: Why should I stream AI answers?
Click both buttons and compare how long you wait before you can start reading.
Your endpoint has to stream too #
The browser can only show what the server sends. Every major AI provider’s SDK has a streaming mode that hands you the answer piece by piece. Forward those pieces as plain text:
export async function POST(request) {
const { question } = await request.json();
const chunks = await askModel(question); // your AI provider's streaming call
const body = new ReadableStream({
async start(controller) {
for await (const text of chunks) {
controller.enqueue(new TextEncoder().encode(text));
}
controller.close();
},
});
return new Response(body, { headers: { 'Content-Type': 'text/plain; charset=utf-8' } });
}
askModel stands for your provider’s SDK call with streaming turned on. Check its docs for the exact shape of each chunk and pull out the text. The handler uses the standard Request and Response objects, so it runs in most modern server runtimes and frameworks.
The API key stays on the server, and the browser receives plain text. Switch providers and only this file changes. The 15 lines in the browser stay exactly the same.
Let users stop it #
A long answer to the wrong question is worth cancelling. Pass a signal to fetch:
const controller = new AbortController();
stopButton.addEventListener('click', () => controller.abort());
const response = await fetch('/api/ask', {
method: 'POST',
body: JSON.stringify({ question }),
signal: controller.signal,
});
Aborting rejects the fetch or the pending reader.read() with an AbortError, depending on when the user clicks, so wrap both in try/catch. Aborting stops the answer in the browser right away. To also stop paying for tokens nobody will read, your server has to notice the closed connection and cancel its own request to the AI provider.
When it still arrives all at once #
If your code is right but the text still shows up in one piece, something between the server and the browser is buffering it. The usual suspect is a reverse proxy in front of your app. Check its buffering settings. On a very common setup, this response header turns buffering off for just this response:
headers: {
'Content-Type': 'text/plain; charset=utf-8',
'X-Accel-Buffering': 'no',
}
Where this comes up #
- “Ask AI” boxes and chat widgets: the obvious one.
- “Explain this” or “Summarize” buttons: anywhere a user waits for generated text.
- Admin tools: AI-written product descriptions, email drafts, alt text.
- Anything slow that produces text: streaming is not AI-specific. Large exports and log viewers work the same way.
Why this matters #
- It feels fast without being faster: the total time is the same, but reading starts almost immediately.
- Nobody double-clicks: visible progress stops the “is it broken?” retries that double your API bill.
- It works everywhere:
fetchstreaming andTextDecoderStreamare supported in every modern browser. No library, no WebSockets.
Find the await response.json() behind your AI feature and swap it for the loop above.
Happy coding!
Marko