# Streaming LLM Output to the Browser Through Your Own Backend: FastAPI, Server-Sent Events and the Buffering Traps

> Source: <https://dev.to/amankumar_apiclaw/streaming-llm-output-to-the-browser-through-your-own-backend-fastapi-server-sent-events-and-the-1g6l>
> Published: 2026-10-11 21:30:04+00:00

I'm Aman Kumar. I build an OpenAI-compatible gateway, and one of the most common support questions I see isn't about models at all. It's "streaming works in my terminal, but in the browser the whole answer shows up at once." The model is streaming fine. Something between your backend and the browser is holding the bytes.

This post is the setup I use for relaying a streamed chat completion from a Python backend to a web page, plus the specific places where buffering sneaks in.

Because the API key would ship to every visitor. Any key in front-end JavaScript is public. The normal pattern is:

`stream=True`, using a key stored in an environment variable.
Your backend also becomes the place to enforce per-user limits, log usage and pick the model, which you want anyway.

The official `openai` package works against any OpenAI-compatible endpoint as long as you set `base_url`. I use the async client so one slow stream doesn't block other requests.

``` python
import os, json
from fastapi import FastAPI, Request
from fastapi.responses import StreamingResponse
from openai import AsyncOpenAI

client = AsyncOpenAI(
    base_url=os.environ["LLM_BASE_URL"],   # any OpenAI-compatible /v1 URL
    api_key=os.environ["LLM_API_KEY"],
)
MODEL = os.environ["LLM_MODEL"]            # copy the exact ID from your provider's model list

app = FastAPI()

def sse(event: dict) -> str:
    return f"data: {json.dumps(event)}\n\n"

@app.post("/chat")
async def chat(request: Request):
    body = await request.json()
    messages = body["messages"]

    async def gen():
        try:
            stream = await client.chat.completions.create(
                model=MODEL, messages=messages, stream=True,
            )
            async for chunk in stream:
                if await request.is_disconnected():
                    await stream.close()   # stop paying for tokens nobody reads
                    return
                if not chunk.choices:
                    continue
                delta = chunk.choices[0].delta.content
                if delta:
                    yield sse({"type": "text", "text": delta})
            yield sse({"type": "done"})
        except Exception as e:
            yield sse({"type": "error", "message": str(e)[:300]})

    return StreamingResponse(
        gen(),
        media_type="text/event-stream",
        headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
    )
```

A few details in there matter more than they look:

`data:` line breaks the framing. Encoding each piece as JSON sidesteps that completely.`choices`.`[0]` on it throws.
`EventSource` only does GET requests and can't send a JSON body, so for a chat endpoint I read the stream with `fetch`:

``` js
async function send(messages, onText) {
  const res = await fetch("/chat", {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({ messages }),
  });
  const reader = res.body.getReader();
  const decoder = new TextDecoder();
  let buf = "";
  while (true) {
    const { value, done } = await reader.read();
    if (done) break;
    buf += decoder.decode(value, { stream: true });
    let i;
    while ((i = buf.indexOf("\n\n")) !== -1) {
      const line = buf.slice(0, i); buf = buf.slice(i + 2);
      if (!line.startsWith("data: ")) continue;
      const ev = JSON.parse(line.slice(6));
      if (ev.type === "text") onText(ev.text);
      if (ev.type === "error") throw new Error(ev.message);
    }
  }
}
```

The important part is the buffer. A network read can end in the middle of an event, or contain three events at once. Splitting on the blank line between events and keeping the leftover is what makes this reliable.

When streaming "doesn't work" in production but works locally, it's almost always one of these.

**Nginx proxy buffering.** By default Nginx buffers upstream responses. Either send the `X-Accel-Buffering: no` header (as above) or set `proxy_buffering off;` for that location. Also raise `proxy_read_timeout` so long answers aren't cut at 60 seconds.

**Compression middleware.** Gzip middleware often waits to fill a block before sending anything. Exclude `text/event-stream` from compression, or don't add `GZipMiddleware` to the streaming route.

**Serverless and some PaaS platforms.** Some hosting setups collect the full response before returning it. If your platform doesn't document streaming support for your runtime, test with a tiny endpoint that yields a timestamp every second before blaming your code.

**Cloudflare and other CDNs.** Streaming generally passes through, but caching rules or response transformations can interfere. `Cache-Control: no-cache` helps, and make sure the route isn't matched by a "cache everything" rule.

**Your own logging.** I've seen a middleware that read the full response body to log it, which turned the stream back into one blob. Anything that touches the body needs to be stream-aware.

A quick check from the command line tells you which side the problem is on:

```
curl -N -X POST https://your-app.example/chat \
  -H "Content-Type: application/json" \
  -d '{"messages":[{"role":"user","content":"Count to 20 slowly"}]}'
```

`-N` turns off curl's own buffering. If the text trickles in here but not in the browser, the problem is front-end parsing. If it arrives all at once here too, look at the server and proxies.

Appending raw text to a `<div>` on every token is fine. Re-rendering Markdown on every token gets slow for long answers. What I do is append plain text while streaming and render Markdown once on `done`, or re-render at most every 100 ms or so. Also, never insert model output as raw HTML; treat it as untrusted.

Since every request passes through your backend, this is where a simple limiter goes: a counter per user per minute in Redis or memory, checked before you open the upstream stream. It protects you from one user (or one buggy front-end loop) burning through your entire allowance, whether you pay per token or have a fixed monthly request budget.

`curl -N` before debugging the browser.
Disclosure: I build [APIClaw](https://apiclaw.biz), an OpenAI-compatible gateway with flat monthly pricing. The code above works the same with any provider that speaks the Chat Completions format; just set the base URL, key and model ID.
