{"slug": "streaming-llm-output-to-the-browser-through-your-own-backend-fastapi-server-sent", "title": "Streaming LLM Output to the Browser Through Your Own Backend: FastAPI, Server-Sent Events and the Buffering Traps", "summary": "Aman Kumar, who builds an OpenAI-compatible gateway, published a setup for relaying streamed chat completions from a Python FastAPI backend to a browser using Server-Sent Events, with code for an async OpenAI client and a fetch-based reader that splits events on blank lines. The writeup identifies the common buffering traps that make streaming appear to fail in production, including Nginx proxy buffering, which can be disabled with the X-Accel-Buffering: no header or proxy_buffering off.", "body_md": "I'm Aman Kumar. I build an OpenAI-compatible gateway, and one of the most common support questions I see isn't about models at all. It's \"streaming works in my terminal, but in the browser the whole answer shows up at once.\" The model is streaming fine. Something between your backend and the browser is holding the bytes.\n\nThis post is the setup I use for relaying a streamed chat completion from a Python backend to a web page, plus the specific places where buffering sneaks in.\n\nBecause the API key would ship to every visitor. Any key in front-end JavaScript is public. The normal pattern is:\n\n`stream=True`, using a key stored in an environment variable.\nYour backend also becomes the place to enforce per-user limits, log usage and pick the model, which you want anyway.\n\nThe official `openai` package works against any OpenAI-compatible endpoint as long as you set `base_url`. I use the async client so one slow stream doesn't block other requests.\n\n``` python\nimport os, json\nfrom fastapi import FastAPI, Request\nfrom fastapi.responses import StreamingResponse\nfrom openai import AsyncOpenAI\n\nclient = AsyncOpenAI(\n    base_url=os.environ[\"LLM_BASE_URL\"],   # any OpenAI-compatible /v1 URL\n    api_key=os.environ[\"LLM_API_KEY\"],\n)\nMODEL = os.environ[\"LLM_MODEL\"]            # copy the exact ID from your provider's model list\n\napp = FastAPI()\n\ndef sse(event: dict) -> str:\n    return f\"data: {json.dumps(event)}\\n\\n\"\n\n@app.post(\"/chat\")\nasync def chat(request: Request):\n    body = await request.json()\n    messages = body[\"messages\"]\n\n    async def gen():\n        try:\n            stream = await client.chat.completions.create(\n                model=MODEL, messages=messages, stream=True,\n            )\n            async for chunk in stream:\n                if await request.is_disconnected():\n                    await stream.close()   # stop paying for tokens nobody reads\n                    return\n                if not chunk.choices:\n                    continue\n                delta = chunk.choices[0].delta.content\n                if delta:\n                    yield sse({\"type\": \"text\", \"text\": delta})\n            yield sse({\"type\": \"done\"})\n        except Exception as e:\n            yield sse({\"type\": \"error\", \"message\": str(e)[:300]})\n\n    return StreamingResponse(\n        gen(),\n        media_type=\"text/event-stream\",\n        headers={\"Cache-Control\": \"no-cache\", \"X-Accel-Buffering\": \"no\"},\n    )\n```\n\nA few details in there matter more than they look:\n\n`data:` line breaks the framing. Encoding each piece as JSON sidesteps that completely.`choices`.`[0]` on it throws.\n`EventSource` only does GET requests and can't send a JSON body, so for a chat endpoint I read the stream with `fetch`:\n\n``` js\nasync function send(messages, onText) {\n  const res = await fetch(\"/chat\", {\n    method: \"POST\",\n    headers: { \"Content-Type\": \"application/json\" },\n    body: JSON.stringify({ messages }),\n  });\n  const reader = res.body.getReader();\n  const decoder = new TextDecoder();\n  let buf = \"\";\n  while (true) {\n    const { value, done } = await reader.read();\n    if (done) break;\n    buf += decoder.decode(value, { stream: true });\n    let i;\n    while ((i = buf.indexOf(\"\\n\\n\")) !== -1) {\n      const line = buf.slice(0, i); buf = buf.slice(i + 2);\n      if (!line.startsWith(\"data: \")) continue;\n      const ev = JSON.parse(line.slice(6));\n      if (ev.type === \"text\") onText(ev.text);\n      if (ev.type === \"error\") throw new Error(ev.message);\n    }\n  }\n}\n```\n\nThe important part is the buffer. A network read can end in the middle of an event, or contain three events at once. Splitting on the blank line between events and keeping the leftover is what makes this reliable.\n\nWhen streaming \"doesn't work\" in production but works locally, it's almost always one of these.\n\n**Nginx proxy buffering.** By default Nginx buffers upstream responses. Either send the `X-Accel-Buffering: no` header (as above) or set `proxy_buffering off;` for that location. Also raise `proxy_read_timeout` so long answers aren't cut at 60 seconds.\n\n**Compression middleware.** Gzip middleware often waits to fill a block before sending anything. Exclude `text/event-stream` from compression, or don't add `GZipMiddleware` to the streaming route.\n\n**Serverless and some PaaS platforms.** Some hosting setups collect the full response before returning it. If your platform doesn't document streaming support for your runtime, test with a tiny endpoint that yields a timestamp every second before blaming your code.\n\n**Cloudflare and other CDNs.** Streaming generally passes through, but caching rules or response transformations can interfere. `Cache-Control: no-cache` helps, and make sure the route isn't matched by a \"cache everything\" rule.\n\n**Your own logging.** I've seen a middleware that read the full response body to log it, which turned the stream back into one blob. Anything that touches the body needs to be stream-aware.\n\nA quick check from the command line tells you which side the problem is on:\n\n```\ncurl -N -X POST https://your-app.example/chat \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"messages\":[{\"role\":\"user\",\"content\":\"Count to 20 slowly\"}]}'\n```\n\n`-N` turns off curl's own buffering. If the text trickles in here but not in the browser, the problem is front-end parsing. If it arrives all at once here too, look at the server and proxies.\n\nAppending raw text to a `<div>` on every token is fine. Re-rendering Markdown on every token gets slow for long answers. What I do is append plain text while streaming and render Markdown once on `done`, or re-render at most every 100 ms or so. Also, never insert model output as raw HTML; treat it as untrusted.\n\nSince every request passes through your backend, this is where a simple limiter goes: a counter per user per minute in Redis or memory, checked before you open the upstream stream. It protects you from one user (or one buggy front-end loop) burning through your entire allowance, whether you pay per token or have a fixed monthly request budget.\n\n`curl -N` before debugging the browser.\nDisclosure: I build [APIClaw](https://apiclaw.biz), an OpenAI-compatible gateway with flat monthly pricing. The code above works the same with any provider that speaks the Chat Completions format; just set the base URL, key and model ID.", "url": "https://wpnews.pro/news/streaming-llm-output-to-the-browser-through-your-own-backend-fastapi-server-sent", "canonical_source": "https://dev.to/amankumar_apiclaw/streaming-llm-output-to-the-browser-through-your-own-backend-fastapi-server-sent-events-and-the-1g6l", "published_at": "2026-10-11 21:30:04+00:00", "updated_at": "2026-10-11 21:33:56.075100+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "developer-tools", "ai-tools"], "entities": ["Aman Kumar", "FastAPI", "OpenAI", "Nginx"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/streaming-llm-output-to-the-browser-through-your-own-backend-fastapi-server-sent", "markdown": "https://wpnews.pro/news/streaming-llm-output-to-the-browser-through-your-own-backend-fastapi-server-sent.md", "text": "https://wpnews.pro/news/streaming-llm-output-to-the-browser-through-your-own-backend-fastapi-server-sent.txt", "jsonld": "https://wpnews.pro/news/streaming-llm-output-to-the-browser-through-your-own-backend-fastapi-server-sent.jsonld"}}