{"slug": "streaming-ai-replies-from-laravel-to-react", "title": "★ Streaming AI replies from Laravel to React", "summary": "Spatie developer Freek Van der Herten published a technical walkthrough of AgentStreamingService, a Laravel component that streams Laravel AI SDK agent responses to a React frontend as newline-delimited JSON events for the company's There There helpdesk, which is in private beta. The service wraps response()->stream() around a generator that yields typed events — 'delta' for TextDelta chunks, 'tool_call' for ToolCall, and 'done' with final HTML — while persisting the completed assistant message and its tool calls to the database in the same pass, and the frontend consumes the stream with plain fetch and a ReadableStream reader instead of SSE.", "body_md": "A couple of posts back I walked through [the agent and tools](https://freek.dev/3089-laravel-ai-sdk-in-practice-the-agent-behind-there-there) we built on top of the Laravel AI SDK for [There There](https://there-there.app), the helpdesk we're putting together at Spatie. I ended that post on a small cliffhanger. We hand the agent off to a streaming service that turns the SDK's response stream into events the frontend can consume. That's this post.\n\nA quick reminder on our angle. After two decades of running our own customer support, we wanted a helpdesk where AI makes support agents faster, not one that tries to replace them. The human reads, thinks, and directs. The model drafts and retrieves in real time while they watch. There There is in private beta right now, and you can apply for early access at [there-there.app](https://there-there.app).\n\nThe obvious way to hand an LLM response to the browser is to wait for it to finish, then render the whole thing. That's easy, but it makes the agent feel sluggish. What we actually want is for the reply to appear word by word, and for the agent's own reasoning (like which tool it's calling) to be narrated along the way.\n\nThe Laravel AI SDK exposes each step of a conversation as a typed event. We iterate over that stream and emit newline-delimited JSON for the frontend, one line per event.\n\nHere's the core of our `AgentStreamingService`.\n\n```\npublic function stream(Agent&HasTools $agent, string $message, AgentChat $chat): StreamedResponse\n{\n    return response()->stream(function () use ($agent, $message, $chat): Generator {\n        yield from $this->streamAgent($agent, $message, $chat);\n    }, 200, [\n        'Content-Type' => 'application/x-ndjson',\n        'Cache-Control' => 'no-cache',\n    ]);\n}\n\nprivate function streamAgent(Agent&HasTools $agent, string $message, AgentChat $chat): Generator\n{\n    set_time_limit(120);\n\n    $fullContent = '';\n    $toolCalls = [];\n\n    foreach ($agent->stream($message) as $event) {\n        if ($event instanceof ToolCall) {\n            yield json_encode(['type' => 'tool_call', 'tool_name' => $event->toolCall->name]).\"\\n\";\n        }\n\n        if ($event instanceof TextDelta) {\n            $fullContent .= $event->delta;\n            yield json_encode(['type' => 'delta', 'content' => $event->delta]).\"\\n\";\n        }\n\n        if ($event instanceof ToolResult) {\n            $toolCalls[] = [\n                'name' => $event->toolResult->name,\n                'arguments' => $event->toolResult->arguments,\n                'result' => $event->toolResult->result,\n            ];\n        }\n    }\n\n    $chat->messages()->create([\n        'role' => MessageRole::Assistant,\n        'content' => $fullContent,\n        'tool_calls' => $toolCalls ?: null,\n    ]);\n\n    yield json_encode([\n        'type' => 'done',\n        'html' => $this->parser->buildFinalHtml($fullContent),\n    ]).\"\\n\";\n}\n```\n\nA few things are worth calling out. `response()->stream()` accepts a generator and flushes each `yield` to the client immediately, which is the whole trick for server-pushed progress. We type the events (`delta`, `tool_call`, `done`) so the frontend can tell a text chunk apart from a tool invocation. And we persist the finished message to the database in the same pass. The stream itself is the source of truth, and the persisted copy is what we rehydrate the next time the user opens the chat.\n\nOn the frontend we consume the stream with plain `fetch`. No library needed, no SSE framing to worry about. Here's the hook we use, trimmed to the bits that matter.\n\n``` js\nconst response = await fetch(sendUrl, {\n    method: 'POST',\n    headers: { ...csrfHeaders(), 'Content-Type': 'application/json' },\n    body: JSON.stringify(buildBody(text, html)),\n    signal: controller.signal,\n});\n\nconst reader = response.body.getReader();\nconst decoder = new TextDecoder();\nlet buffer = '';\n\nwhile (true) {\n    const { done, value } = await reader.read();\n    if (done) break;\n\n    buffer += decoder.decode(value, { stream: true });\n    const lines = buffer.split('\\n');\n    buffer = lines.pop() ?? '';\n\n    for (const line of lines) {\n        processLine(line);\n    }\n}\n\nif (buffer.trim()) {\n    processLine(buffer);\n}\n```\n\nThe buffer pattern is the only thing you have to get right. A single `reader.read()` call can hand you a partial line, a whole line, or several lines at once. We split on `\\n`, keep the last (possibly incomplete) chunk in `buffer`, and flush it at the end. Every completed line is one of our events.\n\nDispatching each event to state is a short switch.\n\n``` js\nif (event.type === 'tool_call') {\n    setMessages((prev) => updateLastStreaming(prev, {\n        toolStatus: formatToolName(event.tool_name),\n    }));\n} else if (event.type === 'delta') {\n    setMessages((prev) => updateLastStreaming(prev, (last) => ({\n        content: last.content + event.content,\n        toolStatus: null,\n    })));\n} else if (event.type === 'done') {\n    setMessages((prev) => updateLastStreaming(prev, {\n        html: event.html,\n        isStreaming: false,\n    }));\n}\n```\n\nWhen a `tool_call` arrives, we show a small \"Looking up tickets\" label above the still-streaming reply. When deltas start landing, we clear the tool label and append to the message content. When `done` arrives, we swap in the server-rendered HTML and flip the streaming flag off. The whole round-trip feels like the agent is typing in front of you.\n\nTODO: short video of the Ask There There agent replying to a question, with a \"Looking up tickets\" label appearing briefly and then the tokens streaming in.\n\nThere are heavier ways to do this. Server-sent events, websockets, a full streaming library on each side. For a chat interface, none of it is necessary. A generator on the server, NDJSON on the wire, and a small buffer loop on the client is about a hundred lines end to end, and it behaves exactly like a proper streaming LLM chat should.\n\nYou can read more about the events the SDK emits in [the Laravel AI SDK repo](https://github.com/laravel/ai). And if you'd like to try There There yourself, we're in private beta right now and you can apply for early access at [there-there.app](https://there-there.app).", "url": "https://wpnews.pro/news/streaming-ai-replies-from-laravel-to-react", "canonical_source": "https://freek.dev/3092-streaming-ai-replies-from-laravel-to-react", "published_at": "2026-10-05 10:30:47+00:00", "updated_at": "2026-10-05 10:49:03.178561+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-tools", "developer-tools"], "entities": ["Spatie", "Laravel AI SDK", "There There", "Freek Van der Herten", "AgentStreamingService", "React", "Laravel"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/streaming-ai-replies-from-laravel-to-react", "markdown": "https://wpnews.pro/news/streaming-ai-replies-from-laravel-to-react.md", "text": "https://wpnews.pro/news/streaming-ai-replies-from-laravel-to-react.txt", "jsonld": "https://wpnews.pro/news/streaming-ai-replies-from-laravel-to-react.jsonld"}}