cd /news/large-language-models/streaming-an-llm-response-in-next-js… · home › topics › large-language-models › article
[ARTICLE · art-145557] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Streaming an LLM response in Next.js 15 without the UI feeling broken

A developer published a Next.js 15 streaming pattern for LLM chat apps using the Vercel AI SDK, showing that `streamText` with `toDataStreamResponse()` on the server and the `useChat` hook on the client handles token streaming, optimistic user messages and loading state with minimal code. The writeup flags two common pitfalls: persisting the assistant message in the browser can truncate output if the client disconnects, so it should be saved server-side in `onFinish`, and usage credits must be checked and spent before the stream starts rather than after, otherwise a user who disconnects mid-stream gets a free generation. The developer also bundled the streaming layer with auth, a credit system and Stripe billing into a paid Next.js 15 starter.

by read2 min views35 publishedOct 5, 2026

Streaming is table stakes for an AI app now. People expect the answer to appear word by word, not after a 20-second spinner. Here's the setup I use in Next.js 15 with the Vercel AI SDK, plus the two things that quietly trip people up.

// app/api/chat/route.ts
import { openai } from '@ai-sdk/openai'
import { streamText, convertToCoreMessages } from 'ai'

export async function POST(req: Request) {
  const { messages } = await req.json()

  const result = streamText({
    model: openai('gpt-4o-mini'),
    messages: convertToCoreMessages(messages),
    onFinish: async ({ text }) => {
      await saveMessage({ role: 'assistant', content: text })
    },
  })

  return result.toDataStreamResponse()
}

The provider is one line. Swapping GPT for Claude is anthropic('claude-...') and nothing else in the handler changes.

'use client'
import { useChat } from '@ai-sdk/react'

export function Chat() {
  const { messages, input, handleInputChange, handleSubmit, status } = useChat()
  // render messages, wire the form to handleSubmit
}

useChat handles the streaming, the optimistic user message, and the state for you. You render messages and you're basically done.

Don't persist the assistant message from the browser. The stream can be cut off (the user closes the tab) and you end up with half a message, or none. Save it on the server in onFinish — that fires once, with the complete text, whether or not the client is still listening.

If you meter usage, check and spend the credit before the stream starts, not in onFinish. Charge after and someone who disconnects mid-stream got a free generation. Spend first (atomically, so two requests can't both pass the check — I wrote about the race-safe version here), and refund on a rare hard failure if you want to be generous.

Model output is markdown, so render it with react-markdown plus a syntax highlighter (rehype-highlight). Add a copy button on code blocks and a stop button wired to useChat's abort. Small touches, but they're the gap between "demo" and "product."

I bundled this streaming layer plus auth, the credit system, and Stripe billing into a Next.js 15 starter so I stop rebuilding it: live demo at https://ai-saas-starter-ashen.vercel.app, code at https://venturionai.gumroad.com/l/ai-saas-starter (40% off the first 10 with LAUNCH40). The patterns above work on their own though, starter or not.

── more in #large-language-models 4 stories · sorted by recency
── more on @next.js 15 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/streaming-an-llm-res…] indexed:0 read:2min 2026-10-05 · —