Your LLM Is Fast in Testing. Why Does It Slow Down in Production? LLM inference slowdowns in production stem from serving mechanics rather than model intelligence, according to an analysis that separates inference into a compute-bound prefill phase and a memory-bandwidth-bound decode phase. The analysis advises engineers to stop reporting a single average latency figure and instead split it into TTFT, ITL/TPOT, throughput and goodput, noting that GPU memory, especially the KV cache, often limits concurrency before compute does. It identifies paged KV cache, continuous batching, chunked prefill, quantisation and speculative decoding as techniques that each save a specific resource, and recommends benchmarking with traffic that resembles real production load. An LLM can answer in two seconds during development and take twelve seconds to answer the same question in production. The model hasn’t changed. The prompt might be identical. The GPU might even be the same. So what happened? Usually, the answer has less to do with model intelligence and more to do with how inference is being served. In development, we test one request at a time: plenty of GPU memory, no queue, no competition. Production is different. One user sends a 500-token question. Another uploads a 20,000-token document. A third asks for a detailed answer that needs 2,000 generated tokens. Dozens more are waiting. Now the server must decide which request runs first, how much GPU memory each one gets, whether requests can share computation, whether a long prompt may delay users who are already receiving tokens, what happens when the KV cache fills up, and how to raise throughput without making individual requests painfully slow. These aren’t primarily prompt-engineering problems. They’re inference engineering problems, and understanding them changes how you think about deploying LLMs. Key takeaways: - Inference has two phases, prefill and decode, with different bottlenecks. A fix for one can do nothing for the other. - Never report only average latency. Split it into TTFT, ITL / TPOT, throughput and goodput. - GPU memory, especially the KV cache, often limits concurrency before compute does. - Each technique paged KV cache, continuous batching, chunked prefill, quantisation, speculative decoding saves a specific resource. - Benchmark with traffic that looks like yours, and ask “what exactly is slow?” before optimising anything. NOTE: My view is that most conversations about LLM performance start with the model and end with a bigger GPU. The more useful conversation starts with one question: what exactly is slow? This article is organised around that question. Examples are illustrative, and where I use numbers, they are arithmetic, not benchmark claims. LLM Inference Has Two Very Different Phases An autoregressive LLM doesn’t produce its response in one operation. It first prefills: it processes the whole input prompt in parallel and builds the attention key-value KV cache. Then it decodes: it generates output tokens one at a time, each depending on everything before it. Prefill involves large matrix computations with high arithmetic intensity, so on modern GPUs it can be compute-bound, particularly for big prompts. The longer the prompt, the more work before the first token. Decode at small batch sizes behaves differently: it spends much of its time moving model weights and cached attention data through GPU memory, which makes it frequently memory-bandwidth-bound. That is why optimising a server isn’t as simple as buying more compute. A faster compute engine doesn’t fix a memory-bandwidth bottleneck, and a technique that helps prefill may do little for decode. Stop Measuring Only Total Response Time A common mistake is reporting a single number: “average response time: 3.2 seconds.” That hides too much. Imagine two systems that both finish in about four seconds. System A makes users wait three seconds, then streams quickly. System B shows the first token after 0.3 seconds and streams steadily. The totals match; the experience doesn’t. A caution on definitions: benchmarking tools don’t all define TPOT identically some include the first token or TTFT in the average , so check the definition before comparing numbers across tools. With TPOT defined as above, a useful approximation is: For example, with TTFT = 400ms, TPOT = 30ms and 100 output tokens: 400 + 99 33 = 3,370 ms, about 3.4 seconds. If users complain the model takes too long to begin, investigate TTFT. If tokens arrive slowly after streaming starts, investigate ITL and TPOT. If the server degrades under concurrent traffic, look at queueing, scheduling, memory pressure and throughput. You cannot fix a latency problem properly until you know which part of latency is failing. Why GPU Memory Becomes the Real Bottleneck Take an 8-billion-parameter model. At FP16, each parameter takes two bytes, so the weights alone need roughly 8 billion × 2 bytes = 16 GB. That excludes the KV cache, runtime buffers, activations and other overhead. During decode, the model reads its weights again for every generated token. For small batches, the GPU may spend a large share of each step just moving those weights out of high-bandwidth memory. A simplified lower bound is: This isn’t a complete performance model it ignores attention-cache traffic, compute and kernel overhead , but it gives useful intuition: if every token requires moving a lot of data, reducing those bytes helps. That is one reason quantisation, batching and KV-cache management matter so much. The KV cache: what makes generation practical, and what eats your memory Without caching, the model would recompute keys and values for all earlier tokens at every step. The KV cache stores them so they can be reused. That saves computation, but the cache grows with both the number of tokens and the number of active sequences. A simplified formula for KV-cache memory per token is below, where L is the number of layers, Hₖᵥ the number of key-value heads, D the head dimension, and S the bytes per element. The leading 2 covers keys and values. With 80 layers, 8 KV heads, head dimension 128 and 2-byte precision, that is 327,680 bytes about 320 KiB per token, so a 32,768-token sequence needs roughly 10 GiB. That is one sequence. With many concurrent requests, KV-cache capacity can decide how many users a server handles. Weights fitting on the GPU does not mean the deployment has enough memory for production traffic. PagedAttention: Treat GPU Memory Like Virtual Memory Traditional KV-cache allocation wastes memory. If a request may generate up to 4,096 tokens, the server may reserve one contiguous region for that maximum. If the request finishes after 300 tokens, most of the reservation sits unused, and across hundreds of requests the fragmentation gets expensive. PagedAttention, introduced with vLLM 1 , organizes KV-cache storage into fixed-size blocks allocated on demand, mapping logical token positions to physical blocks that need not be adjacent. Better utilisation means more active sequences can share the GPU. The broader lesson: Efficient inference isn’t only about faster computation. Sometimes the biggest win comes from allocating memory more intelligently. Continuous Batching, and Why Throughput Isn’t Goodput Imagine three requests: A needs 500 output tokens, B needs 50, C needs 200. With simple static batching, they run as a fixed group, so shorter requests finish early, and their slots sit idle while the longest request continues. Continuous iteration-level batching, introduced by Orca 2 , updates the active batch at every decoding iteration. Finished requests leave, new ones enter, and the GPU stays busy. Real traffic never has identical prompts and output lengths, so this can significantly improve throughput and utilisation. There is a trade-off: a larger active batch can raise total throughput while increasing latency for individual requests. So the question that matters more than peak tokens per second is: How much throughput can we deliver while still meeting our latency targets? Suppose a server produces 10,000 output tokens per second but its p95 TTFT has grown to six seconds when the product needs responses to begin within one. The throughput looks great; the user experience is unacceptable. Goodput counts only the work completed within the service’s performance targets. These objectives are illustrative. A real-time coding assistant and an overnight document pipeline need very different values. The point is that throughput optimisation must be constrained by latency objectives. Why One Long Prompt Can Slow Down Everyone Else Picture a server with ten active conversations, mostly receiving tokens. A request arrives with a 50,000-token document. Its prefill is a lot of work. If the scheduler runs that entire prefill in a way that interrupts ongoing decoding, existing users see noticeable pauses between tokens. A new request has hurt users who were already being served. Chunked prefill 3 splits the prompt into smaller chunks that are scheduled alongside decoding work, which reduces long stalls in token streaming. Chunk size is another trade-off. Smaller chunks tend to stabilize inter-token latency but delay completion of the new request’s prefill. Larger chunks improve prefill efficiency but create more interference. No setting is universally optimal; it depends on your workload and service-level objectives. Quantization: Smaller Weights Don’t Mean Proportionally Faster Inference Quantization is often introduced as a way to shrink models, and that is true. For an 8B model, ignoring metadata and overhead, weights go from about 16 GB FP16/BF16 to 8 GB FP8/INT8 to 4 GB INT4 . Fewer bytes can also reduce bandwidth demand, which helps when decoding is limited by weight movement. But a fourfold reduction in weight bytes does not guarantee fourfold faster inference. Real execution also includes dequantization overhead, attention computation, KV-cache reads, kernel efficiency, hardware support, scheduling overhead and activation processing. Quantization can also change model behavior. A model may score well on generic benchmarks yet lose accuracy on specialized tasks, long-context retrieval, code generation or structured outputs. Evaluate both system performance latency, throughput, memory, cost and model behavior task correctness, output validity, regressions . A smaller model footprint is only a win if the application still performs acceptably Speculative Decoding: More Than One Token per Expensive Step Autoregressive decoding is sequential by nature. Speculative decoding 4, 5 reduces the number of expensive target-model steps: a cheaper draft mechanism proposes several future tokens, and the target model verifies those candidates in parallel. With a correct acceptance and rejection procedure, the target model’s output distribution is preserved. It is not a smaller approximate model replacing the big one: the draft proposes, the target verifies. It is not always faster, though. It works best when the draft is cheap and proposals are accepted often. At high batch sizes the target GPU may already be well utilized, and the extra draft and verification work can shrink or erase the gain. Evaluate it against your real traffic pattern instead of assuming it always helps. Choosing, Benchmarking and Monitoring a Serving Stack The serving engine is part of your architecture When self-hosting, the serving engine owns scheduling, batching, KV-cache management and precision handling, so it matters as much as the model. The right choice depends on model architecture, GPU hardware, prompt- and output-length distributions, concurrency, precision, structured-output needs, operational complexity and target latency. A leaderboard win for one model on one GPU doesn’t prove it will win for your workload, so benchmark under traffic that resembles yours. Benchmark the workload you actually have Running one prompt repeatedly isn’t a benchmark. Real workloads mix short prompts with long outputs and long prompts with short outputs, shared system prompts and unique contexts, steady load and bursts. Vary these dimensions: Then raise the request rate until latency begins rising sharply. That knee is where more incoming work stops producing proportionate useful throughput and just creates queueing. A benchmark should identify the sustainable operating region, not simply the best isolated throughput number. Monitor four areas, not just API latency These are starting points for investigation, not automatic fixes. You need enough observability to tell a compute bottleneck from a memory bottleneck, a scheduling bottleneck or a capacity problem. Otherwise optimisation is guesswork. A Production Scenario: When 50 Concurrent Requests Hit an LLM Server Hypothetical scenario. The numbers and curves in this section are illustrative and show the reasoning process, not measured results. The setup. A team self-hosts an 8B-parameter model behind an internal assistant. Most traffic is short chat. Then a “summarize this document” feature spreads, and at peak about 50 requests are in flight, several carrying very long prompts. Users report two things: “it hangs before it starts answering” and “the streaming stutters.” 6. Validate. Replay a recorded or synthetic traffic mix that includes the long-prompt bursts. Change one variable at a time and compare p95 TTFT, p95 ITL, goodput, queue depth, KV-cache utilisation, and preemptions. Run your quality evaluations if precision changed, then canary the change before full rollout. What the scenario shows. Nothing here required a new model, a new GPU, or a clever kernel. The work was classification, measurement and one targeted change, and it was verified against latency objectives rather than a peak-throughput number. The Real Production Decision: Self-Host or Use an API? After all these optimizations it’s tempting to conclude every serious AI product should self-host. That isn’t necessarily true. Self-hosting gives control, and also makes you own the operations. A hosted API transfers much of that burden to the provider. The right answer depends on the workload. Compare fully loaded cost, not just GPU rental price. Engineering time, idle capacity, redundancy, observability and maintenance all belong in the calculation. Final Thoughts: Ask “What Exactly Is Slow?” First When someone says “our LLM deployment is slow, we need to optimize it,” the first question shouldn’t be “should we quantize?” or “should we move to vLLM?” It should be what exactly is slow? Each symptom points to a different part of the system, and each technique solves a different problem. Applying an optimisation without identifying the bottleneck is like increasing a database connection pool when the real problem is an unindexed query. You change a number without fixing the system. Model quality gets most of the attention, but once an LLM is part of a product, other questions matter just as much. How many users can it serve? How quickly does the first token appear? How steadily does it stream? How much GPU memory does each request consume? What happens under load, and what does each million generated tokens cost? And what happens when an optimisation improves speed but damages output quality? A strong model can still power a poor product if the inference system around it is inefficient. The shift is to think in terms of three resources compute, memory bandwidth and memory capacity and to connect each optimisation to the one it actually saves: None of them is universally right. The correct choice depends on the workload, the hardware and the performance objective. The goal isn’t to make an LLM generate tokens as fast as possible in a benchmark. It’s to build a serving system that delivers correct, timely, economically sustainable responses under real traffic. That’s the difference between running a model and engineering an AI service. Your LLM Is Fast in Testing. Why Does It Slow Down in Production? https://pub.towardsai.net/your-llm-is-fast-in-testing-why-does-it-slow-down-in-production-835778d6974d was originally published in Towards AI https://pub.towardsai.net on Medium, where people are continuing the conversation by highlighting and responding to this story.