{"slug": "what-actually-happens-during-speculative-decoding-in-llms", "title": "What Actually Happens During Speculative Decoding in LLMs", "summary": "A developer's technical breakdown of speculative decoding explains that single-batch LLM generation is memory-bandwidth bound rather than compute bound: on an Nvidia H100, a 70B FP16 model spends roughly 41.8 ms loading 140 GB of weights from VRAM but only 0.07 ms on arithmetic per token. Because verifying five candidate tokens in parallel costs about 42.15 ms — nearly identical to generating one — a small draft model can propose K tokens that a larger target model verifies in a single forward pass, yielding 2-3x faster generation with no loss in output quality.", "body_md": "Autoregressive language model generation is notoriously slow. When you run a 70-billion parameter model on an enterprise GPU, you might get 20 to 30 tokens per second.\n\nThe intuitive assumption is that the GPU compute cores are sweating under massive matrix multiplications. But if you profile the GPU during single-batch generation, you find the exact opposite: **the GPU's compute cores are sitting 95% idle**.\n\nThe true bottleneck is memory bandwidth. Speculative decoding is an inference optimization that exploits this hardware reality to double or triple token generation speed with **zero loss in mathematical output quality**.\n\nHere is what actually happens inside the GPU, the transformer layers, and the sampling math during speculative decoding.\n\nTo understand why speculative decoding works, you first have to look at the GPU memory bus.\n\nDuring autoregressive generation (decoding phase), an LLM generates tokens sequentially: one token per forward pass. To generate a single token for a batch size of 1:\n\nLet us do the concrete math for a **70B parameter model in FP16** running on an **Nvidia H100 GPU**:\n\n$$\\text{Time}_{\\text{memory}} = \\frac{140\\text{ GB}}{3,350\\text{ GB/s}} \\approx 41.8\\text{ ms}$$\n\nGenerating 1 token requires roughly $2 \\times 70 \\times 10^9 = 1.4 \\times 10^{11}$ FLOPs.\n\n$$\\text{Time}_{\\text{compute}} = \\frac{1.4 \\times 10^{11}\\text{ FLOPs}}{1,979 \\times 10^{12}\\text{ FLOPs/s}} \\approx 0.07\\text{ ms}$$\n\nThe ratio is brutal: the GPU spends **41.8 ms loading weights from memory** and only **0.07 ms doing arithmetic**. The arithmetic intensity (FLOPs per byte transferred) is nearly 1, far below the GPU's saturation threshold (which is around 150 FLOPs/byte).\n\n```\nTraditional Autoregressive Decoding (Batch Size = 1):\n[ Load 140 GB Weights ] ---> ( Compute 1 Token: 0.07 ms ) ===> Takes 41.8 ms\n[ Load 140 GB Weights ] ---> ( Compute 1 Token: 0.07 ms ) ===> Takes 41.8 ms\n[ Load 140 GB Weights ] ---> ( Compute 1 Token: 0.07 ms ) ===> Takes 41.8 ms\nResult: 3 forward passes = ~125 ms for 3 tokens\n```\n\nWhat happens if, instead of 1 token, you feed **5 candidate tokens** into the 70B target model simultaneously?\n\nThe model still loads the 140 GB weights from VRAM exactly once (41.8 ms). But now the Tensor Cores multiply a $[5 \\times d]$ matrix against the weights instead of a $[1 \\times d]$ vector.\n\nThe compute time rises from 0.07 ms to 0.35 ms, but both numbers are dwarfed by the 41.8 ms memory transfer time.\n\n```\nEvaluating 1 Token:  41.8 ms (memory) + 0.07 ms (compute) = 41.87 ms\nEvaluating 5 Tokens: 41.8 ms (memory) + 0.35 ms (compute) = 42.15 ms\n```\n\n**Verifying 5 tokens in parallel takes virtually the exact same wall-clock time as generating 1 token.**\n\nSpeculative decoding exploits this asymmetry by splitting generation into two distinct roles:\n\nHere is the exact step-by-step pipeline executed on every speculative decoding iteration:\n\n```\n                  +-----------------------------------+\n                  |  Prompt / Current Prefix Sequence  |\n                  +-----------------+-----------------+\n                                    |\n                                    v\n                  +-----------------------------------+\n                  |      Draft Model (e.g. 1B)        |\n                  |  Generates K tokens sequentially  |\n                  |  Tokens: [x1, x2, x3, x4, x5]     |\n                  +-----------------+-----------------+\n                                    |\n                                    v\n                  +-----------------------------------+\n                  |      Target Model (e.g. 70B)      |\n                  |  Parallel Forward Pass on prefix  |\n                  |  + all K speculative tokens       |\n                  +-----------------+-----------------+\n                                    |\n                                    v\n                  +-----------------------------------+\n                  |    Speculative Rejection Sampler  |\n                  |  Accepts x1, x2, x3               |\n                  |  Rejects x4                       |\n                  |  Samples corrected replacement x4'|\n                  |  Discards x5                      |\n                  +-----------------+-----------------+\n                                    |\n                                    v\n                  +-----------------------------------+\n                  |  Append Accepted Tokens to Prefix |\n                  |  Total yield: 4 tokens in 1 step! |\n                  +-------------------+---------------+\n```\n\nA small model (such as a 1B parameter model paired with a 70B target) generates $K$ candidate tokens $(\\hat{x}_1, \\hat{x}_2, \\dots, \\hat{x}_K)$ autoregressively.\n\nBecause the draft model is tiny (only ~2 GB in FP16), streaming its weights takes under 0.6 ms per token. Generating 5 draft tokens takes roughly $5 \\times 0.6\\text{ ms} = 3.0\\text{ ms}$.\n\nFor each generated token, the draft model records its predicted probability distribution $q(x)$.\n\nAll $K$ draft tokens are appended to the input sequence and evaluated by the 70B target model in **one single forward pass**.\n\nBecause the attention mask allows each token at position $i$ to attend to all prior tokens, the target model produces the true output probability distributions $p_1(x), p_2(x), \\dots, p_K(x)$ for all $K$ positions simultaneously.\n\nWe iterate through the draft tokens from $i = 1$ to $K$:\n\nHere is a clean, dependency-free reference implementation of Leviathan & Chen's Speculative Sampling algorithm:\n\n``` python\nimport numpy as np\n\ndef speculative_sample_step(draft_probs_list, target_probs_list, draft_tokens):\n    \"\"\"\n    draft_probs_list: list of K numpy arrays, each shape (vocab_size,) representing q(x)\n    target_probs_list: list of K+1 numpy arrays, each shape (vocab_size,) representing p(x)\n    draft_tokens: list of K integers drafted by the small model\n\n    Returns: list of accepted and sampled tokens\n    \"\"\"\n    accepted_tokens = []\n    K = len(draft_tokens)\n\n    for i in range(K):\n        token = draft_tokens[i]\n        q_prob = draft_probs_list[i][token]\n        p_prob = target_probs_list[i][token]\n\n        # Calculate acceptance probability\n        acceptance_threshold = min(1.0, p_prob / (q_prob + 1e-10))\n        r = np.random.uniform(0.0, 1.0)\n\n        if r <= acceptance_threshold:\n            # Token accepted\n            accepted_tokens.append(token)\n        else:\n            # Token rejected: sample from modified residual distribution\n            residual = np.maximum(0.0, target_probs_list[i] - draft_probs_list[i])\n            residual_sum = np.sum(residual)\n\n            if residual_sum > 0:\n                p_prime = residual / residual_sum\n            else:\n                p_prime = target_probs_list[i]\n\n            replacement_token = int(np.random.choice(len(p_prime), p=p_prime))\n            accepted_tokens.append(replacement_token)\n\n            # Discard all remaining draft tokens\n            return accepted_tokens\n\n    # All K tokens were accepted! Sample bonus (K+1)-th token directly from target\n    bonus_token = int(np.random.choice(\n        len(target_probs_list[K]), \n        p=target_probs_list[K]\n    ))\n    accepted_tokens.append(bonus_token)\n    return accepted_tokens\n```\n\nA common misconception among developers is that speculative decoding is an approximation or lossy distillation (like quantization or pruning).\n\n**It is mathematically exact.** The probability of generating any sequence under speculative decoding is identical to generating directly from the large target model.\n\nHere is the simple algebraic proof:\n\nLet $x$ be the token candidate at position $i$. What is the total probability that token $x$ is emitted?\n\n**Probability $x$ was drafted and accepted:**\n\n$$P(\\text{drafted } x) \\times P(\\text{accepted} \\mid x) = q(x) \\times \\min\\left(1, \\frac{p(x)}{q(x)}\\right) = \\min(p(x), q(x))$$\n\n**Probability another token was rejected and $x$ was sampled from $p'(x)$:**\n\nThe total rejection probability across all possible tokens is:\n\n$$1 - \\alpha = 1 - \\sum_{y} \\min(p(y), q(y)) = \\sum_{y} (p(y) - \\min(p(y), q(y))) = \\sum_{y} \\max(0, p(y) - q(y))$$\n\nWhen rejection happens, the probability of choosing $x$ from the normalized residual is:\n\n   $$p'(x) = \\frac{\\max(0, p(x) - q(x))}{1 - \\alpha}$$\n\nTherefore, the joint probability of rejection and choosing $x$ is:\n\n   $$(1 - \\alpha) \\times \\frac{\\max(0, p(x) - q(x))}{1 - \\alpha} = \\max(0, p(x) - q(x)) = p(x) - \\min(p(x), q(x))$$\n\nThe draft distribution $q(x)$ completely cancels out. Whether the draft model has 50% accuracy or 90% accuracy affects only **speed**, never **correctness**.\n\nWhile standalone draft models (like Llama-3-8B drafting for Llama-3-70B) work well, the open-source community and inference engines (vLLM, SGLang, llama.cpp) have developed even faster variations:\n\nInstead of maintaining an entire second model in VRAM, Medusa attaches multiple lightweight feed-forward decoding heads to the target model's final hidden state. Head 1 predicts token $t+1$, Head 2 predicts token $t+2$, Head 3 predicts token $t+3$. This eliminates the draft model footprint entirely.\n\nStandard draft models operate on text tokens, losing the rich contextual embeddings computed by the target model. EAGLE passes the top-layer feature vectors of the target model into the draft head. This boosts token acceptance rates from ~60% to over 80-85%.\n\nLinear drafting guesses a single path $[x_1, x_2, x_3]$. If $x_1$ is rejected, everything else is wasted. Tree attention constructs a tree of multiple speculative branches (e.g. top-2 candidates at each depth). The target model verifies all candidate branches in parallel using a custom 2D attention mask, guaranteeing a higher average accepted token yield per step.\n\n``` php\nLinear Speculation (Fragile):\n[Root] -> x1 -> x2 -> x3 (If x1 fails, x2 and x3 are dead)\n\nTree Attention Speculation (Resilient):\n               +--> x1a --> x2a\n[Root] --------|\n               +--> x1b --> x2b\n(If x1a fails, x1b might still be accepted!)\n```\n\nSpeculative decoding is not a magic bullet for every deployment. Here are the 4 scenarios where it actually hurts performance:\n\nSpeculative decoding relies on spare GPU compute. When serving hundreds of concurrent requests (batch size 32, 64, or 128), your GPU is already 100% compute-bound. Adding draft forward passes creates compute contention and reduces overall throughput. Speculative decoding shines primarily at batch sizes 1 to 8 (low-latency streaming).\n\nWhen sampling with high temperature ($T > 1.0$) on highly creative tasks, the output entropy is high. The overlap $\\alpha = \\sum \\min(p, q)$ shrinks dramatically, causing frequent early rejections. Speculative decoding works best on structured output, code generation, and deterministic reasoning ($T \\le 0.7$).\n\nThe draft model and target model must share the exact same tokenizer vocabulary and special token mappings. If token IDs do not map to identical string representations, the logits cannot be compared directly.\n\nRunning a draft model requires reserving additional GPU VRAM for its weights and its own separate KV cache. On memory-constrained GPUs (e.g. 16 GB or 24 GB consumer cards), this extra memory might push the target model into slower quantized formats or offload buffers.", "url": "https://wpnews.pro/news/what-actually-happens-during-speculative-decoding-in-llms", "canonical_source": "https://dev.to/syed_anzar/what-actually-happens-during-speculative-decoding-in-llms-57a7", "published_at": "2026-10-04 13:09:36+00:00", "updated_at": "2026-10-04 13:12:44.563004+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-chips", "machine-learning"], "entities": ["Nvidia", "H100"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/what-actually-happens-during-speculative-decoding-in-llms", "markdown": "https://wpnews.pro/news/what-actually-happens-during-speculative-decoding-in-llms.md", "text": "https://wpnews.pro/news/what-actually-happens-during-speculative-decoding-in-llms.txt", "jsonld": "https://wpnews.pro/news/what-actually-happens-during-speculative-decoding-in-llms.jsonld"}}