{"slug": "aie-2-4-streaming-with-llms", "title": "aie_2.4: streaming with llms", "summary": "A developer's tutorial series on building LLM applications explains how to add streaming to a chat feature so users see tokens as they are generated rather than waiting roughly ten seconds for a complete response. Using a fictional shop's order-status assistant as the example, the writeup walks through the API's response object, token-by-token generation, and the latency problem that streaming solves.", "body_md": "I want to pick up where we left off. In aie_2.2 we met Kora Home, a small online shop that sells home goods, and a customer named Amara.\n\n**Amara:*** Hi, this is Amara. My order 4821 was due last week and I'm still waiting. Where is it?*\n\nWe said the app has three jobs to do with that message, and two of them are behind us. Structured outputs turned her message into a row in the support log, with her name, the order number and what she wants in three separate values. Then tool use let the model ask the app to look up order 4821 in the orders table, which is how we got an answer grounded in what Kora Home's database actually holds.\n\n```\nYour order 4821 is currently delayed at the courier. The updated\ndelivery date is September 24. Sorry for the wait!\n```\n\nThat answer is correct, and it is now in a variable in your code. Job three is getting it in front of Amara.\n\n### A Quick Refresher\n\nIn aie_1.0 we learned that an LLM is a prediction machine.\n\n*A **Large Language Model (LLM)** is a prediction machine that is trained on a large amount of text and is capable of generating coherent, contextually appropriate language by predicting likely continuations of any input given.*\n\nWhat matters most in this lesson is *how* it writes. The model produces its answer one small piece at a time, and at every step it picks from a range of possibilities ranked by how likely each one is. Each of those pieces is a ***token***.\n\n*A **token** is the smallest unit the model works with. Think of it as a chunk of text. “Hello” is one token, “Unbelievable” might be three.*\n\nThe delay this lesson is about comes from there.\n\nIn aie_2.0 we took the call apart. You send a messages array, which is the full transcript of the conversation, along with parameters like `model` and `max_tokens`, and what comes back is the response object.\n\n- `response.content` is a list of content blocks, where the model’s reply sits.`response.content[0].text` takes the first block and pulls out the words.\n- `response.usage` holds the token counts,`input_tokens` for what you sent and`output_tokens` for what the model wrote back.\n- `response.stop_reason` tells you why the model stopped writing, where`end_turn` means it finished its answer.\n\nTwo more turn up near the end of this lesson. In aie_2.2 we used a ***schema*** to fix the shape of an answer, which is a written description of the fields you expect back. In aie_2.3 we described a function to the model as a ***tool***, and the model asked for it by sending back a `tool_use` block, with the stop reason set to `tool_use` as well.\n\n### Where the Ten Seconds Go\n\nAmara’s answer is three lines, which is short as answers go. Plenty of others run longer, and a full walkthrough of the shop’s returns policy could come to several paragraphs.\n\nPut that together with how the model writes. Every token takes a small amount of time to produce, and in aie_2.0 we said a full page comes to roughly 400 tokens, which means a long answer takes eight or ten seconds from the first token to the last.\n\nEvery call we have made across this series behaves the same way. The API waits for the model to finish the entire answer, then sends it back in one piece, and your code sits there waiting for that piece to arrive.\n\nHere is what Amara sees while that happens.\n\n```\nTime     Amara's chat window\n-----    -------------------\n0.0s     empty\n2.0s     empty\n5.0s     empty\n9.8s     empty\n9.9s     the whole answer, all at once\n```\n\nNine and a half seconds of nothing, followed by a wall of text. The answer is correct but the flow still feels broken.\n\nThere is a word for the waiting itself, ***latency***.\n\n**Latency** is the delay between asking for something and getting it. In an LLM feature it is the time between your app sending the call and the answer being usable.\n\n*This is important because as an AI Engineer you are responsible for how a feature feels alongside whether it works. A correct answer that takes ten quiet seconds will be abandoned by the person who asked for it, and abandoned features get cut.*\n\n### Approach One: Show a Spinner\n\nThe first instinct is to fill the awkward silence window with something, which usually means a spinner or a row of animated dots.\n\n```\nTime     Amara's chat window\n-----    -------------------\n0.0s     ● ● ●\n2.0s     ● ● ●\n5.0s     ● ● ●\n9.8s     ● ● ●\n9.9s     the whole answer at once\n```\n\nThis helps a little. Amara can see that something is happening.\n\nShe still has to wait though and people who have been watching dots for as long as ten seconds tend to reload the page or close the tab. A spinner is good to keep to soften the user experience but it does not solve anything.\n\nSo the question is whether the delay can be made shorter.\n\n### Approach Two: Make the Answer Shorter\n\nThe delay we’re trying to tackle comes from the number of tokens, which means producing fewer of them cuts it down. You can do that from either end of the call.\n\n```\nresponse = client.messages.create(\n    model=\"claude-haiku-4-5\",\n    max_tokens=100,\n    messages=[\n        {\"role\": \"user\", \"content\": \"Answer in one short sentence. \" + amara_message}\n    ],\n)\n```\n\n`max_tokens` caps how long the answer can run, and the instruction asks the model to keep it brief. Both work fine, and Amara’s reply now arrives in about two seconds.\n\nThe problem with this shows up in the answer you get back. See the same reply under each setting.\n\n```\nFull length:\nYour order 4821 is currently delayed at the courier. The updated\ndelivery date is September 24. If it hasn't arrived by then, reply\nhere and we'll send a replacement or refund you in full.\n\nShortened:\nOrder 4821 is delayed, expected September 24.\n```\n\nThe short version answers the question and drops everything Amara might do next. For something like a returns policy explanation, the trimming gets worse, because some answers need the length they take.\n\nAnd there is a limit to how far this goes, you can keep trimming until the answer stops being useful, and even then you’d still have to wait, because the model still has to produce whatever tokens are left. A faster model shortens the wait too, and it runs into the same issue, since every answer needs a certain number of tokens to say what it says.\n\nBoth approaches so far accept the same thing, that nothing gets to the screen until the model is finished. The third one challenges that.\n\n### Approach Three: Send the Answer While It Is Written\n\nLet’s go back to the phone call from aie_2.0, where the person on the other end has amnesia and you read them the transcript every time.\n\nEvery call in this series has worked like leaving a voicemail. They record their full answer, the recording finishes, and only then do you get to hear any of it.\n\nStreaming turns it into a live call, where you hear each word at the moment they say it.\n\n**Streaming** is when the API sends each piece of the model’s answer as it is written, so the answer arrives piece by piece while the model is still working on it.\n\nHere is Amara’s window with streaming on, next to what she saw before.\n\n```\nTime     Without streaming          With streaming\n-----    -----------------          --------------\n0.0s     empty                      empty\n0.3s     empty                      Your order\n1.0s     empty                      Your order 4821 is currently\n3.5s     empty                      Your order 4821 is currently delayed\n                                      at the courier. The updated\n9.9s     the full answer           the full answer\n```\n\nSomething you’d notice is the bottom row is the same in both columns. Streaming leaves the model’s writing speed alone, so the complete answer is ready at 9.9 seconds either way.\n\nWhat changes is everything above that row. Without streaming, Amara sees nothing until 9.9 seconds, because the first word and the last word arrive together. With streaming, the first words appear at 0.3 seconds and the rest keep coming, so she spends those ten seconds reading instead of waiting.\n\nThat first measurement is something called ***Time to first token***.\n\n**Time to first token** is the delay between sending a request and receiving the very first token of the reply.\n\nSo streaming changes which number Amara experiences. Without it, the only number she feels is the ten seconds she spends staring at an empty window. With it, the number she feels is the 0.3 seconds before the first words show up, and the remaining time is spent reading an answer that keeps growing.\n\n#### What Arrives and When\n\nAn ordinary call sends one request and gets one response back, with the connection closing once the answer is delivered.\n\nStreaming keeps the connection open. The model writes a piece, that piece is sent, the model writes another, and this continues until the answer is complete.\n\nHere is what reaches your code for Amara’s answer.\n\n```\n\"Your order\"\n\" 4821 is\"\n\" currently delayed\"\n\" at the courier.\"\n...\nthen, at the very end:\nstop_reason: \"end_turn\", output_tokens: 38\n```\n\nEach piece of text is called a ***delta***.\n\n**Delta** means the change. In streaming it is whatever is new since the last piece, so a text delta holds the words the model just produced.\n\nYour code joins the deltas together in order, which gives you the full answer. The stop reason and the token count arrive after everything, because none of them exist until the model has finished writing.\n\n### What the Code Looks Like\n\nWe are going to build this up from a call you have already seen, changing one thing at a time.\n\nHere is the ordinary call that produced Amara’s answer in aie_2.3, with the tool list and the conversation so far.\n\n```\nresponse = client.messages.create(\n    model=\"claude-haiku-4-5\",\n    max_tokens=1024,\n    tools=tools,\n    messages=messages,\n)\n```\n\nAnthropic's Python library gives you a second method for streaming, and it takes the same parameters. The only change here is `create` becoming `stream`.\n\n```\nclient.messages.stream(\n    model=\"claude-haiku-4-5\",\n    max_tokens=1024,\n    tools=tools,\n    messages=messages,\n)\n```\n\nNow something has to change about how you handle it. An ordinary call hands you one finished answer, so assigning it to `response` makes sense. A streaming call hands you pieces over several seconds, which means the connection needs to stay open while they arrive, and close once they stop.\n\nPython has a term for that, `with`.\n\n```\nwith client.messages.stream(\n    model=\"claude-haiku-4-5\",\n    max_tokens=1024,\n    tools=tools,\n    messages=messages,\n) as stream:\n    # the pieces arrive in here\n```\n\nEverything indented under `with` happens while the connection is open. Once that indented part finishes, Python closes the connection, and it does this whether the code ran cleanly or hit an error partway through. `as stream` gives the open connection a name so you can reach it.\n\nSo now you need to do something with the pieces. `stream.text_stream` gives them to you one at a time, and a `for` loop runs once for each piece.\n\n```\n    for text in stream.text_stream:\n        print(text)\n```\n\nIf you run it as is, it’ll work fine but you’ll quickly see that the output looks wrong, it’ll look something like this:\n\n```\nYour order\n 4821 is\n currently delayed\n at the courier.\n```\n\nEach piece is on its own line because `print` adds a new line every time it runs. You fix this by telling it to add nothing instead, you pass end in and set it to an empty string, like so:\n\n```\n    for text in stream.text_stream:\n        print(text, end=\"\")\n```\n\nSo everything shows up on one line, more naturally.\n\n```\nYour order 4821 is currently delayed at the courier.\n```\n\nOne more thing about `print`. Python sometimes holds text back for a moment and writes it out in batches to save effort, which would undo the whole point of streaming. `flush=True` tells it to put each piece on the screen immediately.\n\n```\n    for text in stream.text_stream:\n        print(text, end=\"\", flush=True)\n```\n\nThe last piece is getting the finished response. The loop ends when the model stops writing, and at that point everything from aie_2.0 is available through `get_final_message()`.\n\n```\n    final = stream.get_final_message()\n\nprint(final.stop_reason)          # end_turn\nprint(final.usage.output_tokens)  # 38\n```\n\nHere is all of it together.\n\n```\nwith client.messages.stream(\n    model=\"claude-haiku-4-5\",\n    max_tokens=1024,\n    tools=tools,\n    messages=messages,\n) as stream:\n    for text in stream.text_stream:\n        print(text, end=\"\", flush=True)\n\n    final = stream.get_final_message()\n```\n\nSo you get both things. The loop put Amara’s answer on the screen piece by piece, and `final` holds the same complete response object you have worked with since aie_2.0.\n\nThis closes all the jobs we started out with in the earlier lessons. Amara's message came in, structured outputs filed it, tool use found order 4821, and now her answer starts appearing on screen about a third of a second after she hits send.\n\n### When to Stream\n\nStream when a person is reading the answer as it arrives, which covers chat features, assistants, anything with someone watching the screen.\n\nSkip it when your code is the only reader, like a background job that categorises yesterday’s support messages needs the complete answer before it can do anything with it, so the pieces arriving early don’t add much value, and `client.messages.create(...)` is simpler to write and simpler to debug.\n\nThe question to ask is whether a person is waiting on the other end, since that is the only situation where the time to first token means anything.\n\n### Where We Are\n\nThis closes the three jobs the Kora Home assistant had to do with one customer message. Structured outputs made the message something the app could file, tool use connected the model to order details it had no way of knowing, and streaming put the answer on Amara’s screen while it was still being written.\n\nNext is the build, where you put all three together into a working support assistant. Also, every call we have written has also included instructions for the model, and we have never looked at them properly. We’ll do that in aie_3.0, on the system prompt.", "url": "https://wpnews.pro/news/aie-2-4-streaming-with-llms", "canonical_source": "https://heymeraki.substack.com/p/aie_24-streaming-with-llms", "published_at": "2026-10-05 16:04:40+00:00", "updated_at": "2026-10-05 16:16:14.877054+00:00", "lang": "en", "topics": ["large-language-models", "ai-tools", "natural-language-processing"], "entities": [], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/aie-2-4-streaming-with-llms", "markdown": "https://wpnews.pro/news/aie-2-4-streaming-with-llms.md", "text": "https://wpnews.pro/news/aie-2-4-streaming-with-llms.txt", "jsonld": "https://wpnews.pro/news/aie-2-4-streaming-with-llms.jsonld"}}