{"slug": "your-llm-is-fast-in-testing-why-does-it-slow-down-in-production", "title": "Your LLM Is Fast in Testing. Why Does It Slow Down in Production?", "summary": "LLM inference slowdowns in production stem from serving mechanics rather than model intelligence, according to an analysis that separates inference into a compute-bound prefill phase and a memory-bandwidth-bound decode phase. The analysis advises engineers to stop reporting a single average latency figure and instead split it into TTFT, ITL/TPOT, throughput and goodput, noting that GPU memory, especially the KV cache, often limits concurrency before compute does. It identifies paged KV cache, continuous batching, chunked prefill, quantisation and speculative decoding as techniques that each save a specific resource, and recommends benchmarking with traffic that resembles real production load.", "body_md": "An LLM can answer in two seconds during development and take twelve seconds to answer the same question in production.\n\nThe model hasn’t changed. The prompt might be identical. The GPU might even be the same. So what happened?\n\nUsually, the answer has less to do with model intelligence and more to do with how inference is being served.\n\nIn development, we test one request at a time: plenty of GPU memory, no queue, no competition. Production is different. One user sends a 500-token question. Another uploads a 20,000-token document. A third asks for a detailed answer that needs 2,000 generated tokens. Dozens more are waiting.\n\nNow the server must decide which request runs first, how much GPU memory each one gets, whether requests can share computation, whether a long prompt may delay users who are already receiving tokens, what happens when the KV cache fills up, and how to raise throughput without making individual requests painfully slow.\n\nThese aren’t primarily prompt-engineering problems. They’re inference engineering problems, and understanding them changes how you think about deploying LLMs.\n\n**Key takeaways:**\n\n- Inference has two phases, prefill and decode, with different bottlenecks. A fix for one can do nothing for the other. \n\n- Never report only average latency. Split it into TTFT, ITL / TPOT, throughput and goodput. \n\n- GPU memory, especially the KV cache, often limits concurrency before compute does. \n\n- Each technique (paged KV cache, continuous batching, chunked prefill, quantisation, speculative decoding) saves a specific resource. \n\n- Benchmark with traffic that looks like yours, and ask “what exactly is slow?” before optimising anything.\n\nNOTE: My view is that most conversations about LLM performance start with the model and end with a bigger GPU. The more useful conversation starts with one question: what exactly is slow? This article is organised around that question. Examples are illustrative, and where I use numbers, they are arithmetic, not benchmark claims.\n\n**LLM Inference Has Two Very Different Phases**\n\nAn autoregressive LLM doesn’t produce its response in one operation. It first prefills: it processes the whole input prompt in parallel and builds the attention key-value (KV) cache. Then it decodes: it generates output tokens one at a time, each depending on everything before it.\n\nPrefill involves large matrix computations with high arithmetic intensity, so on modern GPUs it can be compute-bound, particularly for big prompts. The longer the prompt, the more work before the first token. Decode at small batch sizes behaves differently: it spends much of its time moving model weights and cached attention data through GPU memory, which makes it frequently memory-bandwidth-bound.\n\nThat is why optimising a server isn’t as simple as buying more compute. A faster compute engine doesn’t fix a memory-bandwidth bottleneck, and a technique that helps prefill may do little for decode.\n\n**Stop Measuring Only Total Response Time**\n\nA common mistake is reporting a single number: “average response time: 3.2 seconds.” That hides too much. Imagine two systems that both finish in about four seconds. System A makes users wait three seconds, then streams quickly. System B shows the first token after 0.3 seconds and streams steadily. The totals match; the experience doesn’t.\n\nA caution on definitions: benchmarking tools don’t all define TPOT identically (some include the first token or TTFT in the average), so check the definition before comparing numbers across tools. With TPOT defined as above, a useful approximation is:\n\nFor example, with TTFT = 400ms, TPOT = 30ms and 100 output tokens: 400 + 99 * 33 = 3,370 ms, about 3.4 seconds.\n\nIf users complain the model takes too long to begin, investigate TTFT. If tokens arrive slowly after streaming starts, investigate ITL and TPOT. If the server degrades under concurrent traffic, look at queueing, scheduling, memory pressure and throughput.\n\nYou cannot fix a latency problem properly until you know which part of latency is failing.\n\n**Why GPU Memory Becomes the Real Bottleneck**\n\nTake an 8-billion-parameter model. At FP16, each parameter takes two bytes, so the weights alone need roughly 8 billion × 2 bytes = 16 GB. That excludes the KV cache, runtime buffers, activations and other overhead. \n\nDuring decode, the model reads its weights again for every generated token. For small batches, the GPU may spend a large share of each step just moving those weights out of high-bandwidth memory. A simplified lower bound is:\n\nThis isn’t a complete performance model (it ignores attention-cache traffic, compute and kernel overhead), but it gives useful intuition: if every token requires moving a lot of data, reducing those bytes helps. That is one reason quantisation, batching and KV-cache management matter so much.\n\n*The KV cache: what makes generation practical, and what eats your memory*\n\nWithout caching, the model would recompute keys and values for all earlier tokens at every step. The KV cache stores them so they can be reused. That saves computation, but the cache grows with both the number of tokens and the number of active sequences.\n\nA simplified formula for KV-cache memory per token is below, where L is the number of layers, Hₖᵥ the number of key-value heads, D the head dimension, and S the bytes per element. The leading 2 covers keys and values.\n\nWith 80 layers, 8 KV heads, head dimension 128 and 2-byte precision, that is 327,680 bytes (about 320 KiB) per token, so a 32,768-token sequence needs roughly 10 GiB. That is one sequence. With many concurrent requests, KV-cache capacity can decide how many users a server handles. Weights fitting on the GPU does not mean the deployment has enough memory for production traffic.\n\n**PagedAttention: Treat GPU Memory Like Virtual Memory**\n\nTraditional KV-cache allocation wastes memory. If a request may generate up to 4,096 tokens, the server may reserve one contiguous region for that maximum. If the request finishes after 300 tokens, most of the reservation sits unused, and across hundreds of requests the fragmentation gets expensive.\n\nPagedAttention, introduced with vLLM [1], organizes KV-cache storage into fixed-size blocks allocated on demand, mapping logical token positions to physical blocks that need not be adjacent.\n\nBetter utilisation means more active sequences can share the GPU. The broader lesson:\n\nEfficient inference isn’t only about faster computation. Sometimes the biggest win comes from allocating memory more intelligently.\n\n**Continuous Batching, and Why Throughput Isn’t Goodput**\n\nImagine three requests: A needs 500 output tokens, B needs 50, C needs 200. With simple static batching, they run as a fixed group, so shorter requests finish early, and their slots sit idle while the longest request continues.\n\nContinuous (iteration-level) batching, introduced by Orca [2], updates the active batch at every decoding iteration. Finished requests leave, new ones enter, and the GPU stays busy. Real traffic never has identical prompts and output lengths, so this can significantly improve throughput and utilisation.\n\nThere is a trade-off: a larger active batch can raise total throughput while increasing latency for individual requests. So the question that matters more than peak tokens per second is:\n\nHow much throughput can we deliver while still meeting our latency targets?\n\nSuppose a server produces 10,000 output tokens per second but its p95 TTFT has grown to six seconds when the product needs responses to begin within one. The throughput looks great; the user experience is unacceptable. Goodput counts only the work completed within the service’s performance targets.\n\nThese objectives are illustrative. A real-time coding assistant and an overnight document pipeline need very different values. The point is that throughput optimisation must be constrained by latency objectives.\n\n**Why One Long Prompt Can Slow Down Everyone Else**\n\nPicture a server with ten active conversations, mostly receiving tokens. A request arrives with a 50,000-token document. Its prefill is a lot of work. If the scheduler runs that entire prefill in a way that interrupts ongoing decoding, existing users see noticeable pauses between tokens. A new request has hurt users who were already being served.\n\nChunked prefill [3] splits the prompt into smaller chunks that are scheduled alongside decoding work, which reduces long stalls in token streaming.\n\nChunk size is another trade-off. Smaller chunks tend to stabilize inter-token latency but delay completion of the new request’s prefill. Larger chunks improve prefill efficiency but create more interference. No setting is universally optimal; it depends on your workload and service-level objectives.\n\n**Quantization: Smaller Weights Don’t Mean Proportionally Faster Inference**\n\nQuantization is often introduced as a way to shrink models, and that is true. For an 8B model, ignoring metadata and overhead, weights go from about 16 GB (FP16/BF16) to 8 GB (FP8/INT8) to 4 GB (INT4). Fewer bytes can also reduce bandwidth demand, which helps when decoding is limited by weight movement.\n\nBut a fourfold reduction in weight bytes does not guarantee fourfold faster inference. Real execution also includes dequantization overhead, attention computation, KV-cache reads, kernel efficiency, hardware support, scheduling overhead and activation processing.\n\nQuantization can also change model behavior. A model may score well on generic benchmarks yet lose accuracy on specialized tasks, long-context retrieval, code generation or structured outputs. Evaluate both system performance (latency, throughput, memory, cost) and model behavior (task correctness, output validity, regressions).\n\nA smaller model footprint is only a win if the application still performs acceptably\n\n**Speculative Decoding: More Than One Token per Expensive Step**\n\nAutoregressive decoding is sequential by nature. Speculative decoding [4, 5] reduces the number of expensive target-model steps: a cheaper draft mechanism proposes several future tokens, and the target model verifies those candidates in parallel.\n\nWith a correct acceptance and rejection procedure, the target model’s output distribution is preserved. It is not a smaller approximate model replacing the big one: the draft proposes, the target verifies.\n\nIt is not always faster, though. It works best when the draft is cheap and proposals are accepted often. At high batch sizes the target GPU may already be well utilized, and the extra draft and verification work can shrink or erase the gain. Evaluate it against your real traffic pattern instead of assuming it always helps.\n\n**Choosing, Benchmarking and Monitoring a Serving Stack**\n\n***The serving engine is part of your architecture*** \nWhen self-hosting, the serving engine owns scheduling, batching, KV-cache management and precision handling, so it matters as much as the model.\n\nThe right choice depends on model architecture, GPU hardware, prompt- and output-length distributions, concurrency, precision, structured-output needs, operational complexity and target latency. A leaderboard win for one model on one GPU doesn’t prove it will win for your workload, so benchmark under traffic that resembles yours.\n\n*Benchmark the workload you actually have*\n\nRunning one prompt repeatedly isn’t a benchmark. Real workloads mix short prompts with long outputs and long prompts with short outputs, shared system prompts and unique contexts, steady load and bursts. Vary these dimensions:\n\nThen raise the request rate until latency begins rising sharply. That knee is where more incoming work stops producing proportionate useful throughput and just creates queueing.\n\nA benchmark should identify the sustainable operating region, not simply the best isolated throughput number.\n\n*Monitor four areas, not just API latency*\n\nThese are starting points for investigation, not automatic fixes. You need enough observability to tell a compute bottleneck from a memory bottleneck, a scheduling bottleneck or a capacity problem. Otherwise optimisation is guesswork.\n\n**A Production Scenario: When 50 Concurrent Requests Hit an LLM Server**\n\nHypothetical scenario. The numbers and curves in this section are illustrative and show the reasoning process, not measured results.\n\nThe setup. A team self-hosts an 8B-parameter model behind an internal assistant. Most traffic is short chat. Then a “summarize this document” feature spreads, and at peak about 50 requests are in flight, several carrying very long prompts. Users report two things: “it hangs before it starts answering” and “the streaming stutters.”\n\n6. Validate. Replay a recorded or synthetic traffic mix that includes the long-prompt bursts. Change one variable at a time and compare p95 TTFT, p95 ITL, goodput, queue depth, KV-cache utilisation, and preemptions. Run your quality evaluations if precision changed, then canary the change before full rollout.\n\nWhat the scenario shows. Nothing here required a new model, a new GPU, or a clever kernel. The work was classification, measurement and one targeted change, and it was verified against latency objectives rather than a peak-throughput number.\n\n**The Real Production Decision: Self-Host or Use an API?**\n\nAfter all these optimizations it’s tempting to conclude every serious AI product should self-host. That isn’t necessarily true. Self-hosting gives control, and also makes you own the operations. A hosted API transfers much of that burden to the provider. The right answer depends on the workload.\n\nCompare fully loaded cost, not just GPU rental price. Engineering time, idle capacity, redundancy, observability and maintenance all belong in the calculation.\n\n**Final Thoughts: Ask “What Exactly Is Slow?” First**\n\nWhen someone says “our LLM deployment is slow, we need to optimize it,” the first question shouldn’t be “should we quantize?” or “should we move to vLLM?” It should be what exactly is slow? Each symptom points to a different part of the system, and each technique solves a different problem.\n\nApplying an optimisation without identifying the bottleneck is like increasing a database connection pool when the real problem is an unindexed query. You change a number without fixing the system.\n\nModel quality gets most of the attention, but once an LLM is part of a product, other questions matter just as much. How many users can it serve? How quickly does the first token appear? How steadily does it stream? How much GPU memory does each request consume? What happens under load, and what does each million generated tokens cost? And what happens when an optimisation improves speed but damages output quality?\n\nA strong model can still power a poor product if the inference system around it is inefficient. The shift is to think in terms of three resources (compute, memory bandwidth and memory capacity) and to connect each optimisation to the one it actually saves:\n\nNone of them is universally right. The correct choice depends on the workload, the hardware and the performance objective.\n\nThe goal isn’t to make an LLM generate tokens as fast as possible in a benchmark. It’s to build a serving system that delivers correct, timely, economically sustainable responses under real traffic. That’s the difference between running a model and engineering an AI service.\n\n[Your LLM Is Fast in Testing. Why Does It Slow Down in Production?](https://pub.towardsai.net/your-llm-is-fast-in-testing-why-does-it-slow-down-in-production-835778d6974d) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.", "url": "https://wpnews.pro/news/your-llm-is-fast-in-testing-why-does-it-slow-down-in-production", "canonical_source": "https://pub.towardsai.net/your-llm-is-fast-in-testing-why-does-it-slow-down-in-production-835778d6974d?source=rss----98111c9905da---4", "published_at": "2026-10-10 06:46:41+00:00", "updated_at": "2026-10-10 07:08:42.028895+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "mlops", "ai-research"], "entities": [], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/your-llm-is-fast-in-testing-why-does-it-slow-down-in-production", "markdown": "https://wpnews.pro/news/your-llm-is-fast-in-testing-why-does-it-slow-down-in-production.md", "text": "https://wpnews.pro/news/your-llm-is-fast-in-testing-why-does-it-slow-down-in-production.txt", "jsonld": "https://wpnews.pro/news/your-llm-is-fast-in-testing-why-does-it-slow-down-in-production.jsonld"}}