# aie_2.4: streaming with llms

> Source: <https://heymeraki.substack.com/p/aie_24-streaming-with-llms>
> Published: 2026-10-05 16:04:40+00:00

I want to pick up where we left off. In aie_2.2 we met Kora Home, a small online shop that sells home goods, and a customer named Amara.

**Amara:*** Hi, this is Amara. My order 4821 was due last week and I'm still waiting. Where is it?*

We said the app has three jobs to do with that message, and two of them are behind us. Structured outputs turned her message into a row in the support log, with her name, the order number and what she wants in three separate values. Then tool use let the model ask the app to look up order 4821 in the orders table, which is how we got an answer grounded in what Kora Home's database actually holds.

```
Your order 4821 is currently delayed at the courier. The updated
delivery date is September 24. Sorry for the wait!
```

That answer is correct, and it is now in a variable in your code. Job three is getting it in front of Amara.

### A Quick Refresher

In aie_1.0 we learned that an LLM is a prediction machine.

*A **Large Language Model (LLM)** is a prediction machine that is trained on a large amount of text and is capable of generating coherent, contextually appropriate language by predicting likely continuations of any input given.*

What matters most in this lesson is *how* it writes. The model produces its answer one small piece at a time, and at every step it picks from a range of possibilities ranked by how likely each one is. Each of those pieces is a ***token***.

*A **token** is the smallest unit the model works with. Think of it as a chunk of text. “Hello” is one token, “Unbelievable” might be three.*

The delay this lesson is about comes from there.

In aie_2.0 we took the call apart. You send a messages array, which is the full transcript of the conversation, along with parameters like `model` and `max_tokens`, and what comes back is the response object.

- `response.content` is a list of content blocks, where the model’s reply sits.`response.content[0].text` takes the first block and pulls out the words.
- `response.usage` holds the token counts,`input_tokens` for what you sent and`output_tokens` for what the model wrote back.
- `response.stop_reason` tells you why the model stopped writing, where`end_turn` means it finished its answer.

Two more turn up near the end of this lesson. In aie_2.2 we used a ***schema*** to fix the shape of an answer, which is a written description of the fields you expect back. In aie_2.3 we described a function to the model as a ***tool***, and the model asked for it by sending back a `tool_use` block, with the stop reason set to `tool_use` as well.

### Where the Ten Seconds Go

Amara’s answer is three lines, which is short as answers go. Plenty of others run longer, and a full walkthrough of the shop’s returns policy could come to several paragraphs.

Put that together with how the model writes. Every token takes a small amount of time to produce, and in aie_2.0 we said a full page comes to roughly 400 tokens, which means a long answer takes eight or ten seconds from the first token to the last.

Every call we have made across this series behaves the same way. The API waits for the model to finish the entire answer, then sends it back in one piece, and your code sits there waiting for that piece to arrive.

Here is what Amara sees while that happens.

```
Time     Amara's chat window
-----    -------------------
0.0s     empty
2.0s     empty
5.0s     empty
9.8s     empty
9.9s     the whole answer, all at once
```

Nine and a half seconds of nothing, followed by a wall of text. The answer is correct but the flow still feels broken.

There is a word for the waiting itself, ***latency***.

**Latency** is the delay between asking for something and getting it. In an LLM feature it is the time between your app sending the call and the answer being usable.

*This is important because as an AI Engineer you are responsible for how a feature feels alongside whether it works. A correct answer that takes ten quiet seconds will be abandoned by the person who asked for it, and abandoned features get cut.*

### Approach One: Show a Spinner

The first instinct is to fill the awkward silence window with something, which usually means a spinner or a row of animated dots.

```
Time     Amara's chat window
-----    -------------------
0.0s     ● ● ●
2.0s     ● ● ●
5.0s     ● ● ●
9.8s     ● ● ●
9.9s     the whole answer at once
```

This helps a little. Amara can see that something is happening.

She still has to wait though and people who have been watching dots for as long as ten seconds tend to reload the page or close the tab. A spinner is good to keep to soften the user experience but it does not solve anything.

So the question is whether the delay can be made shorter.

### Approach Two: Make the Answer Shorter

The delay we’re trying to tackle comes from the number of tokens, which means producing fewer of them cuts it down. You can do that from either end of the call.

```
response = client.messages.create(
    model="claude-haiku-4-5",
    max_tokens=100,
    messages=[
        {"role": "user", "content": "Answer in one short sentence. " + amara_message}
    ],
)
```

`max_tokens` caps how long the answer can run, and the instruction asks the model to keep it brief. Both work fine, and Amara’s reply now arrives in about two seconds.

The problem with this shows up in the answer you get back. See the same reply under each setting.

```
Full length:
Your order 4821 is currently delayed at the courier. The updated
delivery date is September 24. If it hasn't arrived by then, reply
here and we'll send a replacement or refund you in full.

Shortened:
Order 4821 is delayed, expected September 24.
```

The short version answers the question and drops everything Amara might do next. For something like a returns policy explanation, the trimming gets worse, because some answers need the length they take.

And there is a limit to how far this goes, you can keep trimming until the answer stops being useful, and even then you’d still have to wait, because the model still has to produce whatever tokens are left. A faster model shortens the wait too, and it runs into the same issue, since every answer needs a certain number of tokens to say what it says.

Both approaches so far accept the same thing, that nothing gets to the screen until the model is finished. The third one challenges that.

### Approach Three: Send the Answer While It Is Written

Let’s go back to the phone call from aie_2.0, where the person on the other end has amnesia and you read them the transcript every time.

Every call in this series has worked like leaving a voicemail. They record their full answer, the recording finishes, and only then do you get to hear any of it.

Streaming turns it into a live call, where you hear each word at the moment they say it.

**Streaming** is when the API sends each piece of the model’s answer as it is written, so the answer arrives piece by piece while the model is still working on it.

Here is Amara’s window with streaming on, next to what she saw before.

```
Time     Without streaming          With streaming
-----    -----------------          --------------
0.0s     empty                      empty
0.3s     empty                      Your order
1.0s     empty                      Your order 4821 is currently
3.5s     empty                      Your order 4821 is currently delayed
                                      at the courier. The updated
9.9s     the full answer           the full answer
```

Something you’d notice is the bottom row is the same in both columns. Streaming leaves the model’s writing speed alone, so the complete answer is ready at 9.9 seconds either way.

What changes is everything above that row. Without streaming, Amara sees nothing until 9.9 seconds, because the first word and the last word arrive together. With streaming, the first words appear at 0.3 seconds and the rest keep coming, so she spends those ten seconds reading instead of waiting.

That first measurement is something called ***Time to first token***.

**Time to first token** is the delay between sending a request and receiving the very first token of the reply.

So streaming changes which number Amara experiences. Without it, the only number she feels is the ten seconds she spends staring at an empty window. With it, the number she feels is the 0.3 seconds before the first words show up, and the remaining time is spent reading an answer that keeps growing.

#### What Arrives and When

An ordinary call sends one request and gets one response back, with the connection closing once the answer is delivered.

Streaming keeps the connection open. The model writes a piece, that piece is sent, the model writes another, and this continues until the answer is complete.

Here is what reaches your code for Amara’s answer.

```
"Your order"
" 4821 is"
" currently delayed"
" at the courier."
...
then, at the very end:
stop_reason: "end_turn", output_tokens: 38
```

Each piece of text is called a ***delta***.

**Delta** means the change. In streaming it is whatever is new since the last piece, so a text delta holds the words the model just produced.

Your code joins the deltas together in order, which gives you the full answer. The stop reason and the token count arrive after everything, because none of them exist until the model has finished writing.

### What the Code Looks Like

We are going to build this up from a call you have already seen, changing one thing at a time.

Here is the ordinary call that produced Amara’s answer in aie_2.3, with the tool list and the conversation so far.

```
response = client.messages.create(
    model="claude-haiku-4-5",
    max_tokens=1024,
    tools=tools,
    messages=messages,
)
```

Anthropic's Python library gives you a second method for streaming, and it takes the same parameters. The only change here is `create` becoming `stream`.

```
client.messages.stream(
    model="claude-haiku-4-5",
    max_tokens=1024,
    tools=tools,
    messages=messages,
)
```

Now something has to change about how you handle it. An ordinary call hands you one finished answer, so assigning it to `response` makes sense. A streaming call hands you pieces over several seconds, which means the connection needs to stay open while they arrive, and close once they stop.

Python has a term for that, `with`.

```
with client.messages.stream(
    model="claude-haiku-4-5",
    max_tokens=1024,
    tools=tools,
    messages=messages,
) as stream:
    # the pieces arrive in here
```

Everything indented under `with` happens while the connection is open. Once that indented part finishes, Python closes the connection, and it does this whether the code ran cleanly or hit an error partway through. `as stream` gives the open connection a name so you can reach it.

So now you need to do something with the pieces. `stream.text_stream` gives them to you one at a time, and a `for` loop runs once for each piece.

```
    for text in stream.text_stream:
        print(text)
```

If you run it as is, it’ll work fine but you’ll quickly see that the output looks wrong, it’ll look something like this:

```
Your order
 4821 is
 currently delayed
 at the courier.
```

Each piece is on its own line because `print` adds a new line every time it runs. You fix this by telling it to add nothing instead, you pass end in and set it to an empty string, like so:

```
    for text in stream.text_stream:
        print(text, end="")
```

So everything shows up on one line, more naturally.

```
Your order 4821 is currently delayed at the courier.
```

One more thing about `print`. Python sometimes holds text back for a moment and writes it out in batches to save effort, which would undo the whole point of streaming. `flush=True` tells it to put each piece on the screen immediately.

```
    for text in stream.text_stream:
        print(text, end="", flush=True)
```

The last piece is getting the finished response. The loop ends when the model stops writing, and at that point everything from aie_2.0 is available through `get_final_message()`.

```
    final = stream.get_final_message()

print(final.stop_reason)          # end_turn
print(final.usage.output_tokens)  # 38
```

Here is all of it together.

```
with client.messages.stream(
    model="claude-haiku-4-5",
    max_tokens=1024,
    tools=tools,
    messages=messages,
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)

    final = stream.get_final_message()
```

So you get both things. The loop put Amara’s answer on the screen piece by piece, and `final` holds the same complete response object you have worked with since aie_2.0.

This closes all the jobs we started out with in the earlier lessons. Amara's message came in, structured outputs filed it, tool use found order 4821, and now her answer starts appearing on screen about a third of a second after she hits send.

### When to Stream

Stream when a person is reading the answer as it arrives, which covers chat features, assistants, anything with someone watching the screen.

Skip it when your code is the only reader, like a background job that categorises yesterday’s support messages needs the complete answer before it can do anything with it, so the pieces arriving early don’t add much value, and `client.messages.create(...)` is simpler to write and simpler to debug.

The question to ask is whether a person is waiting on the other end, since that is the only situation where the time to first token means anything.

### Where We Are

This closes the three jobs the Kora Home assistant had to do with one customer message. Structured outputs made the message something the app could file, tool use connected the model to order details it had no way of knowing, and streaming put the answer on Amara’s screen while it was still being written.

Next is the build, where you put all three together into a working support assistant. Also, every call we have written has also included instructions for the model, and we have never looked at them properly. We’ll do that in aie_3.0, on the system prompt.
