aie_2.4: streaming with llms A developer's tutorial series on building LLM applications explains how to add streaming to a chat feature so users see tokens as they are generated rather than waiting roughly ten seconds for a complete response. Using a fictional shop's order-status assistant as the example, the writeup walks through the API's response object, token-by-token generation, and the latency problem that streaming solves. I want to pick up where we left off. In aie 2.2 we met Kora Home, a small online shop that sells home goods, and a customer named Amara. Amara: Hi, this is Amara. My order 4821 was due last week and I'm still waiting. Where is it? We said the app has three jobs to do with that message, and two of them are behind us. Structured outputs turned her message into a row in the support log, with her name, the order number and what she wants in three separate values. Then tool use let the model ask the app to look up order 4821 in the orders table, which is how we got an answer grounded in what Kora Home's database actually holds. Your order 4821 is currently delayed at the courier. The updated delivery date is September 24. Sorry for the wait That answer is correct, and it is now in a variable in your code. Job three is getting it in front of Amara. A Quick Refresher In aie 1.0 we learned that an LLM is a prediction machine. A Large Language Model LLM is a prediction machine that is trained on a large amount of text and is capable of generating coherent, contextually appropriate language by predicting likely continuations of any input given. What matters most in this lesson is how it writes. The model produces its answer one small piece at a time, and at every step it picks from a range of possibilities ranked by how likely each one is. Each of those pieces is a token . A token is the smallest unit the model works with. Think of it as a chunk of text. “Hello” is one token, “Unbelievable” might be three. The delay this lesson is about comes from there. In aie 2.0 we took the call apart. You send a messages array, which is the full transcript of the conversation, along with parameters like model and max tokens , and what comes back is the response object. - response.content is a list of content blocks, where the model’s reply sits. response.content 0 .text takes the first block and pulls out the words. - response.usage holds the token counts, input tokens for what you sent and output tokens for what the model wrote back. - response.stop reason tells you why the model stopped writing, where end turn means it finished its answer. Two more turn up near the end of this lesson. In aie 2.2 we used a schema to fix the shape of an answer, which is a written description of the fields you expect back. In aie 2.3 we described a function to the model as a tool , and the model asked for it by sending back a tool use block, with the stop reason set to tool use as well. Where the Ten Seconds Go Amara’s answer is three lines, which is short as answers go. Plenty of others run longer, and a full walkthrough of the shop’s returns policy could come to several paragraphs. Put that together with how the model writes. Every token takes a small amount of time to produce, and in aie 2.0 we said a full page comes to roughly 400 tokens, which means a long answer takes eight or ten seconds from the first token to the last. Every call we have made across this series behaves the same way. The API waits for the model to finish the entire answer, then sends it back in one piece, and your code sits there waiting for that piece to arrive. Here is what Amara sees while that happens. Time Amara's chat window ----- ------------------- 0.0s empty 2.0s empty 5.0s empty 9.8s empty 9.9s the whole answer, all at once Nine and a half seconds of nothing, followed by a wall of text. The answer is correct but the flow still feels broken. There is a word for the waiting itself, latency . Latency is the delay between asking for something and getting it. In an LLM feature it is the time between your app sending the call and the answer being usable. This is important because as an AI Engineer you are responsible for how a feature feels alongside whether it works. A correct answer that takes ten quiet seconds will be abandoned by the person who asked for it, and abandoned features get cut. Approach One: Show a Spinner The first instinct is to fill the awkward silence window with something, which usually means a spinner or a row of animated dots. Time Amara's chat window ----- ------------------- 0.0s ● ● ● 2.0s ● ● ● 5.0s ● ● ● 9.8s ● ● ● 9.9s the whole answer at once This helps a little. Amara can see that something is happening. She still has to wait though and people who have been watching dots for as long as ten seconds tend to reload the page or close the tab. A spinner is good to keep to soften the user experience but it does not solve anything. So the question is whether the delay can be made shorter. Approach Two: Make the Answer Shorter The delay we’re trying to tackle comes from the number of tokens, which means producing fewer of them cuts it down. You can do that from either end of the call. response = client.messages.create model="claude-haiku-4-5", max tokens=100, messages= {"role": "user", "content": "Answer in one short sentence. " + amara message} , max tokens caps how long the answer can run, and the instruction asks the model to keep it brief. Both work fine, and Amara’s reply now arrives in about two seconds. The problem with this shows up in the answer you get back. See the same reply under each setting. Full length: Your order 4821 is currently delayed at the courier. The updated delivery date is September 24. If it hasn't arrived by then, reply here and we'll send a replacement or refund you in full. Shortened: Order 4821 is delayed, expected September 24. The short version answers the question and drops everything Amara might do next. For something like a returns policy explanation, the trimming gets worse, because some answers need the length they take. And there is a limit to how far this goes, you can keep trimming until the answer stops being useful, and even then you’d still have to wait, because the model still has to produce whatever tokens are left. A faster model shortens the wait too, and it runs into the same issue, since every answer needs a certain number of tokens to say what it says. Both approaches so far accept the same thing, that nothing gets to the screen until the model is finished. The third one challenges that. Approach Three: Send the Answer While It Is Written Let’s go back to the phone call from aie 2.0, where the person on the other end has amnesia and you read them the transcript every time. Every call in this series has worked like leaving a voicemail. They record their full answer, the recording finishes, and only then do you get to hear any of it. Streaming turns it into a live call, where you hear each word at the moment they say it. Streaming is when the API sends each piece of the model’s answer as it is written, so the answer arrives piece by piece while the model is still working on it. Here is Amara’s window with streaming on, next to what she saw before. Time Without streaming With streaming ----- ----------------- -------------- 0.0s empty empty 0.3s empty Your order 1.0s empty Your order 4821 is currently 3.5s empty Your order 4821 is currently delayed at the courier. The updated 9.9s the full answer the full answer Something you’d notice is the bottom row is the same in both columns. Streaming leaves the model’s writing speed alone, so the complete answer is ready at 9.9 seconds either way. What changes is everything above that row. Without streaming, Amara sees nothing until 9.9 seconds, because the first word and the last word arrive together. With streaming, the first words appear at 0.3 seconds and the rest keep coming, so she spends those ten seconds reading instead of waiting. That first measurement is something called Time to first token . Time to first token is the delay between sending a request and receiving the very first token of the reply. So streaming changes which number Amara experiences. Without it, the only number she feels is the ten seconds she spends staring at an empty window. With it, the number she feels is the 0.3 seconds before the first words show up, and the remaining time is spent reading an answer that keeps growing. What Arrives and When An ordinary call sends one request and gets one response back, with the connection closing once the answer is delivered. Streaming keeps the connection open. The model writes a piece, that piece is sent, the model writes another, and this continues until the answer is complete. Here is what reaches your code for Amara’s answer. "Your order" " 4821 is" " currently delayed" " at the courier." ... then, at the very end: stop reason: "end turn", output tokens: 38 Each piece of text is called a delta . Delta means the change. In streaming it is whatever is new since the last piece, so a text delta holds the words the model just produced. Your code joins the deltas together in order, which gives you the full answer. The stop reason and the token count arrive after everything, because none of them exist until the model has finished writing. What the Code Looks Like We are going to build this up from a call you have already seen, changing one thing at a time. Here is the ordinary call that produced Amara’s answer in aie 2.3, with the tool list and the conversation so far. response = client.messages.create model="claude-haiku-4-5", max tokens=1024, tools=tools, messages=messages, Anthropic's Python library gives you a second method for streaming, and it takes the same parameters. The only change here is create becoming stream . client.messages.stream model="claude-haiku-4-5", max tokens=1024, tools=tools, messages=messages, Now something has to change about how you handle it. An ordinary call hands you one finished answer, so assigning it to response makes sense. A streaming call hands you pieces over several seconds, which means the connection needs to stay open while they arrive, and close once they stop. Python has a term for that, with . with client.messages.stream model="claude-haiku-4-5", max tokens=1024, tools=tools, messages=messages, as stream: the pieces arrive in here Everything indented under with happens while the connection is open. Once that indented part finishes, Python closes the connection, and it does this whether the code ran cleanly or hit an error partway through. as stream gives the open connection a name so you can reach it. So now you need to do something with the pieces. stream.text stream gives them to you one at a time, and a for loop runs once for each piece. for text in stream.text stream: print text If you run it as is, it’ll work fine but you’ll quickly see that the output looks wrong, it’ll look something like this: Your order 4821 is currently delayed at the courier. Each piece is on its own line because print adds a new line every time it runs. You fix this by telling it to add nothing instead, you pass end in and set it to an empty string, like so: for text in stream.text stream: print text, end="" So everything shows up on one line, more naturally. Your order 4821 is currently delayed at the courier. One more thing about print . Python sometimes holds text back for a moment and writes it out in batches to save effort, which would undo the whole point of streaming. flush=True tells it to put each piece on the screen immediately. for text in stream.text stream: print text, end="", flush=True The last piece is getting the finished response. The loop ends when the model stops writing, and at that point everything from aie 2.0 is available through get final message . final = stream.get final message print final.stop reason end turn print final.usage.output tokens 38 Here is all of it together. with client.messages.stream model="claude-haiku-4-5", max tokens=1024, tools=tools, messages=messages, as stream: for text in stream.text stream: print text, end="", flush=True final = stream.get final message So you get both things. The loop put Amara’s answer on the screen piece by piece, and final holds the same complete response object you have worked with since aie 2.0. This closes all the jobs we started out with in the earlier lessons. Amara's message came in, structured outputs filed it, tool use found order 4821, and now her answer starts appearing on screen about a third of a second after she hits send. When to Stream Stream when a person is reading the answer as it arrives, which covers chat features, assistants, anything with someone watching the screen. Skip it when your code is the only reader, like a background job that categorises yesterday’s support messages needs the complete answer before it can do anything with it, so the pieces arriving early don’t add much value, and client.messages.create ... is simpler to write and simpler to debug. The question to ask is whether a person is waiting on the other end, since that is the only situation where the time to first token means anything. Where We Are This closes the three jobs the Kora Home assistant had to do with one customer message. Structured outputs made the message something the app could file, tool use connected the model to order details it had no way of knowing, and streaming put the answer on Amara’s screen while it was still being written. Next is the build, where you put all three together into a working support assistant. Also, every call we have written has also included instructions for the model, and we have never looked at them properly. We’ll do that in aie 3.0, on the system prompt.