{"slug": "a-stream-can-start-finishing-is-another-matter", "title": "A Stream Can Start. Finishing Is Another Matter.", "summary": "A developer argues that the first chunk of a streamed LLM response proves only that generation started, not that it finished or that the output is usable, and that applications must treat transport closure, provider terminal events and stop reasons, and application-level validation as three separate checks. The writeup contrasts how OpenAI's Responses API exposes response.completed and response.incomplete events while Claude sends a final message_stop with the stop reason in message_delta, noting a stream can terminate cleanly while reporting max_tokens. It recommends accumulating partial tool-argument JSON and validating it against a schema rather than treating a closing brace or a disappearing spinner as proof of success.", "body_md": "The first chunk proves an LLM response has begun. It does not prove the answer finished, and neither does it prove that your application can use it.\n\nImagine a support assistant helping a customer reconnect an integration. The first words arrive quickly: “Let’s get this working again.” Then come numbered steps. Check the connection. Open the integration settings. Reauthorize the account.\n\nHalfway through the next instruction, the answer stops:\n\n“Before you reconnect, make sure you…”\n\nMake sure you what?\n\nThe UI has already displayed text that looks useful. In this example, the request began with HTTP 200. The customer sees something shaped like an answer, while the application has no confirmed completion.\n\nNow the app must decide what it received: an answer, an interrupted answer, or something that failed before becoming usable. That decision needs evidence. A spinner disappearing is a surprisingly weak definition of success.\n\nStreaming means the API sends pieces of an LLM response incrementally, while generation continues. With HTTP streaming over [server sent events (SSE)](https://developers.openai.com/api/docs/guides/streaming-responses?api-mode=responses), the client can process those events and display text without waiting for the entire response. That improves perceived responsiveness: the customer starts reading instead of staring at an empty chat bubble. It does not make the early pieces a completion signal.\n\nFor our support assistant, “Let’s get this working again” is evidence of progress. It says nothing about whether the final instruction will arrive. Keep two concepts separate: content becoming available and the response reaching a confirmed terminal state.\n\nAlso distinguish text deltas from other events. A stream may carry lifecycle events, tool argument fragments, usage information, pings, and errors. A parser that only extracts displayable text can miss the evidence needed to decide whether the request succeeded. [Claude’s streaming protocol](https://platform.claude.com/docs/en/build-with-claude/streaming) explicitly includes these different event types.\n\nThe support answer that stopped halfway could have several explanations.\n\nEach case needs its own fix; one generic “LLM failed” bucket hides the difference.\n\nProvider protocols also express completion differently. OpenAI’s Responses API exposes [response.completed and response.incomplete events](https://developers.openai.com/api/reference/resources/responses/streaming-events). Claude sends a final message_stop, with the stop reason supplied in message_delta. A terminal event therefore needs interpretation: Claude can finish its stream while [reporting that generation hit max_tokens](https://platform.claude.com/docs/en/build-with-claude/handling-stop-reasons). The protocol ended; the support instructions may not have.\n\n“Finished” hides three separate questions:\n\nTreat them as three separate checks. [Inferock’s methodology](https://inferock.ai/methodology/) likewise considers stream completion, finish reason, delivered content, and structural validation when assessing broken output.\n\nA clean connection close answers only the transport question. The provider’s terminal event and stop information answer the generation question. The application still owns the usability question.\n\nFor the support assistant, a declared contract might require a troubleshooting result containing instructions and a final verification step. If the model returns structured data, a missing closing JSON brace makes failure obvious. But a closing brace only proves syntactic closure: the object could still omit a required field.\n\nTool arguments make the distinction sharper. Claude [streams tool inputs as partial JSON strings](https://platform.claude.com/docs/en/build-with-claude/streaming), which must be accumulated and parsed. Parsing successfully still does not establish that the arguments [satisfy your tool’s schema](https://inferock.ai/methodology/).\n\nFor plain support prose, completeness is harder to prove. Define checks your application can actually enforce; do not pretend a period at the end establishes that every necessary instruction arrived.\n\nThe accidental promotion happens when rendering and success share the same handler:\n\n```\non_text_delta: append_to_ui(delta)\non_connection_close: mark_answer_complete()\n```\n\nThat second line assumes what it needs to establish.\n\nUse an explicit partial state instead. Our interrupted assistant answer can remain visible, labeled “Response interrupted,” without becoming the final answer in conversation history or a downstream workflow. Rendering is provisional; committing is a separate decision.\n\nKeep streaming. Just do not let the UI’s enthusiasm write your success criteria. Show progress promptly. Promote it deliberately.\n\nWhen the support answer stops, “it cut off” is a symptom report. A useful call record preserves enough evidence to distinguish possible failure classes without inventing a cause. [Inferock’s measurement approach](https://inferock.ai/methodology/) starts with observable request, response, timing, usage, and outcome evidence.\n\nFor your application’s own record, keep:\n\nThese are proposed application fields; whether a gateway exposes all of them varies. Inferock documents [streaming milestones “when available”](https://inferock.ai/methodology/) and explicitly notes that provider fields differ.\n\nKeep client cancellation separate from upstream failure when the evidence supports that distinction. If the customer presses Stop, record it. If your server reaches its own deadline, record that too. Otherwise, “provider interrupted” can become a convenient label for your own abort.\n\nFinally, record where observation occurred. Gateway receipt is not proof of browser delivery. That boundary matters when comparing the call record with what the customer actually saw.\n\n[Inferock’s gateway documentation](https://inferock.ai/docs/first-call/) states that successful calls, streaming or not, preserve the upstream HTTP status and provider response content. The gateway does not replace the answer with a measurement envelope of its own.\n\nIts [first call workflow](https://inferock.ai/docs/first-call/) has you supply a request ID, then find the corresponding measured call in Calls. [Call details](https://inferock.ai/docs/receipts-ledger/) can include provider, model, request ID, attempt and failure class, tokens, cost, timing, linked findings, and redacted payload evidence when those fields are available.\n\nFor our interrupted support answer, that gives engineers a concrete starting point: locate the request, inspect the recorded outcome, and compare it with the application’s interrupted state.\n\n[Inferock’s methodology for broken output](https://inferock.ai/methodology/) considers delivered content, finish reason, stream completion, output usage, JSON parsing, and schema validation. It describes checks for output that is malformed, truncated, empty, or invalid under a declared contract.\n\nThe boundary matters. A call record helps investigate an interruption; it does not expose every internal provider cause or discover every requirement your support workflow forgot to declare. Inferock explicitly [limits structural checks](https://inferock.ai/methodology/): they do not establish factual correctness, and tool argument validation needs a declared schema or objective contract.\n\nMake the application’s rule explicit: chunks are provisional; success requires an acceptable provider outcome and application validation. That follows the distinction between [provider stop information](https://platform.claude.com/docs/en/build-with-claude/handling-stop-reasons) and [checks against a declared response contract](https://inferock.ai/methodology/).\n\nA practical sequence is:\n\nBack in the support chat, “Before you reconnect, make sure you…” should not quietly become a completed answer. Keep it visible as an interruption, with a recovery path.\n\nThe assistant started helping. Your application still has to establish whether it finished.", "url": "https://wpnews.pro/news/a-stream-can-start-finishing-is-another-matter", "canonical_source": "https://dev.to/bharathkoneti/a-stream-can-start-finishing-is-another-matter-3e44", "published_at": "2026-10-05 12:40:06+00:00", "updated_at": "2026-10-05 12:48:59.651968+00:00", "lang": "en", "topics": ["large-language-models", "ai-agents", "ai-tools", "developer-tools"], "entities": ["OpenAI", "Claude", "Anthropic", "Inferock"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/a-stream-can-start-finishing-is-another-matter", "markdown": "https://wpnews.pro/news/a-stream-can-start-finishing-is-another-matter.md", "text": "https://wpnews.pro/news/a-stream-can-start-finishing-is-another-matter.txt", "jsonld": "https://wpnews.pro/news/a-stream-can-start-finishing-is-another-matter.jsonld"}}