# Stream AI answers word by word in 15 lines

> Source: <https://markodenic.tech/stream-ai-answers-word-by-word-in-15-lines/>
> Published: 2026-10-09 18:53:46+00:00

# Stream AI answers word by word in 15 lines

[Set](https://www.google.com/preferences/source?q=markodenic.tech)

**markodenic.tech** as your preferred Google source
[Sponsor this newsletter · reach 9,500+ active developers](https://markodenic.tech/sponsorship/)

You add an “Ask AI” box to a client’s site. The user types a question, clicks the button, and stares at a spinner. Eight seconds later, a wall of text appears all at once. Half of them already clicked again, or left.

The model was writing the answer the whole time. Your code just waited for the last word before showing the first one. Streaming fixes that: the answer appears word by word as it is generated, the way AI chat apps do it.

## The fix

``` js
async function ask(question) {
  const response = await fetch('/api/ask', {
    method: 'POST',
    body: JSON.stringify({ question }),
  });

  const reader = response.body.pipeThrough(new TextDecoderStream()).getReader();
  output.textContent = '';

  while (true) {
    const { value, done } = await reader.read();
    if (done) break;
    output.textContent += value;
  }
}
```

`response.body` is a stream. Instead of waiting for `response.json()`, you read it chunk by chunk, and `TextDecoderStream` turns the raw bytes into text. Every chunk goes on screen the moment it arrives.

Q: Why should I stream AI answers?

Click both buttons and compare how long you wait before you can start reading.

## Your endpoint has to stream too

The browser can only show what the server sends. Every major AI provider’s SDK has a streaming mode that hands you the answer piece by piece. Forward those pieces as plain text:

```
export async function POST(request) {
  const { question } = await request.json();
  const chunks = await askModel(question); // your AI provider's streaming call

  const body = new ReadableStream({
    async start(controller) {
      for await (const text of chunks) {
        controller.enqueue(new TextEncoder().encode(text));
      }
      controller.close();
    },
  });

  return new Response(body, { headers: { 'Content-Type': 'text/plain; charset=utf-8' } });
}
```

`askModel` stands for your provider’s SDK call with streaming turned on. Check its docs for the exact shape of each chunk and pull out the text. The handler uses the standard `Request` and `Response` objects, so it runs in most modern server runtimes and frameworks.

The API key stays on the server, and the browser receives plain text. Switch providers and only this file changes. The 15 lines in the browser stay exactly the same.

## Let users stop it

A long answer to the wrong question is worth cancelling. Pass a signal to `fetch`:

``` js
const controller = new AbortController();
stopButton.addEventListener('click', () => controller.abort());

const response = await fetch('/api/ask', {
  method: 'POST',
  body: JSON.stringify({ question }),
  signal: controller.signal,
});
```

Aborting rejects the `fetch` or the pending `reader.read()` with an `AbortError`, depending on when the user clicks, so wrap both in `try/catch`. Aborting stops the answer in the browser right away. To also stop paying for tokens nobody will read, your server has to notice the closed connection and cancel its own request to the AI provider.

## When it still arrives all at once

If your code is right but the text still shows up in one piece, something between the server and the browser is buffering it. The usual suspect is a reverse proxy in front of your app. Check its buffering settings. On a very common setup, this response header turns buffering off for just this response:

```
headers: {
  'Content-Type': 'text/plain; charset=utf-8',
  'X-Accel-Buffering': 'no',
}
```

## Where this comes up

- **“Ask AI” boxes and chat widgets:** the obvious one.
- **“Explain this” or “Summarize” buttons:** anywhere a user waits for generated text.
- **Admin tools:** AI-written product descriptions, email drafts, alt text.
- **Anything slow that produces text:** streaming is not AI-specific. Large exports and log viewers work the same way.

## Why this matters

- **It feels fast without being faster:** the total time is the same, but reading starts almost immediately.
- **Nobody double-clicks:** visible progress stops the “is it broken?” retries that double your API bill.
- **It works everywhere:**`fetch` streaming and`TextDecoderStream` are supported in every modern browser. No library, no WebSockets.

Find the `await response.json()` behind your AI feature and swap it for the loop above.

Happy coding!

Marko
