{"slug": "dflash-2-keep-drafting-parallel", "title": "DFlash 2: Keep Drafting Parallel", "summary": "Inco AI released DFlash 2, a parallel speculative decoding technique that delivers over 20% more output from every verification pass with around 1% added cycle latency, achieving 2.7–3.4× throughput of autoregressive decoding at batch size 1 with the Qwen3.8-27B drafter in SGLang. The original DFlash has been downloaded more than 3.5 million times on Hugging Face and runs in SGLang, vLLM, TensorRT-LLM, and llama.cpp, with NVIDIA reporting up to 15× throughput on Blackwell GPUs and Google reporting 3× tokens per second on TPUs.", "body_md": "# DFlash 2: Keep Drafting Parallel\n\nInference is the bottleneck of the agent era. Agents read, plan, and call tools, often for hours or days. They consume tokens at a rate chat never approached. Every one of those tokens takes a full forward pass over the model. At Inco AI, we are building the inference stack scaled to the token economics of tomorrow. This post is a sneak peek.\n\nOur team released [DFlash](https://arxiv.org/abs/2602.06036) in January; it\nnow runs in SGLang, vLLM, TensorRT-LLM, and llama.cpp. NVIDIA measured\n[up to 15× throughput](https://developer.nvidia.com/blog/boost-inference-performance-up-to-15x-on-nvidia-blackwell-using-dflash-speculative-decoding/)\nwith it on Blackwell GPUs; Google reported\n[3× more tokens per second](https://developers.googleblog.com/supercharging-llm-inference-on-google-tpus-achieving-3x-speedups-with-diffusion-style-speculative-decoding/)\non TPUs; CoreWeave's production Kimi K2.7 Code endpoint, [the fastest for\nthat model on Artificial Analysis](https://www.coreweave.com/blog/kimi-k2-7-code-now-available-on-serverless-inference-with-leading-benchmark-price-performance),\nruns DFlash by default. The ecosystem now builds on it:\n[NVIDIA](https://huggingface.co/nvidia/Kimi-K2.6-DFlash),\n[Red Hat](https://huggingface.co/RedHatAI/gemma-4-31B-it-speculator.dflash), and\n[Modal](https://huggingface.co/modal-labs/Kimi-K3-DFlash) have all published\nDFlash drafters; Meta\n([Muse Glimmer](https://huggingface.co/meta-models/Muse-Glimmer-30B-assistant)),\nPoolside ([Laguna](https://huggingface.co/poolside/Laguna-S-2.1-DFlash)),\nXiaomi\n([MiMo-V2.5-Pro](https://huggingface.co/XiaomiMiMo/MiMo-V2.5-Pro-FP4-DFlash)),\nand NVIDIA\n([Nemotron 3.5 Lightning](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DFlash))\nship official drafters with their own models. On Hugging Face, DFlash models have\nbeen downloaded **more than 3.5 million times** (as of August 2026).\n\nSpeculative decoding is a core piece of the modern inference\nstack. 1 A small draft model guesses a block of tokens,\nand the target model verifies the whole block in one forward pass. Good guesses\nturn one pass into several tokens; bad ones just get thrown away. For years, though, the draft itself stayed\n\n*: one token at a time. DFlash made it one-pass too: the entire block, every position, predicted*\n\n**autoregressive***.*\n\n**in parallel** DFlash 2 pushes parallel drafting one step further: **over 20% more output from\nevery verification pass, for around 1% added cycle latency**, with the output\nprovably unchanged. Across benchmarks the gain runs 16–25%. With the\nQwen3.8-27B drafter released today, SGLang serves at **2.7–3.4× the\nthroughput of autoregressive decoding** at batch size 1. Predicting every\nposition independently leaves headroom in two places: choosing the right\ntokens and holding accuracy to the end of the block. DFlash 2 recovers\nboth without giving up the one-pass design.\n\n[Run It Now](#run-it-now)\n\nDFlash 2 already runs in the mainstream inference engines:\n\nDownload and install the [prebuilt oMLX with DFlash 2 support](https://github.com/z-lab/omlx-fork/releases/download/0.6.2-dflash2/oMLX-0.6.2-zlab-dflash2-arm64-signed.dmg).\n\nTo run Qwen3.8-27B with DFlash 2:\n\n-\nOpen the oMLX\n\n[Model Downloader](http://127.0.0.1:8891/admin/dashboard?tab=models&modelsTab=downloader)and download: -\nOpen the\n\n[Model Manager](http://127.0.0.1:8891/admin/dashboard?tab=models&modelsTab=manager)and edit`mlx-community/Qwen3.8-27B-4bit`\n\n. Configure DFlash with the following settings:**DFlash**: enabled** Draft model**:`incoai/Qwen3.8-27B-DFlash2`\n\n**Draft quantization**: enabled** Runtime block size**:`5`\n\n**Verify mode**:`dflash`\n\n-\nSave the settings and load the target model.\n\n[The Right Tokens Are Already There](#the-right-tokens-are-already-there)\n\nDFlash predicts every position independently, in parallel. Each pick is\nplausible on its own. Yet nothing makes them fit together, and an\nincoherent block is cut short at verification.\nRecent methods such as [Domino](https://arxiv.org/abs/2605.29707) and\n[DSpark](https://arxiv.org/abs/2607.05147) buy coherence with sequential\nMarkov heads that rewrite each position's full-vocabulary distribution.\nBut is that costly autoregressive correction really necessary?\n\nNo. The evidence is already in DFlash's own candidate lists. Take the first position: DFlash's top pick is right 85.4% of the time, but the right token is in its top 16 candidates 99.5% of the time. Even when the top pick is wrong, the right token is usually on the list.\n\n| Metric | 0 | 1 | 2 | 3 | 4 | 5 | 6 | Acceptance length |\n|---|---|---|---|---|---|---|---|---|\n| Recall@1 | 85.4% | 80.3% | 79.4% | 78.3% | 77.5% | 75.9% | 72.9% | 4.27 |\n| Recall@16 | 99.5% | 97.3% | 94.8% | 92.6% | 90.8% | 89.4% | 87.8% | 6.79 |\n\nAn oracle that always picks the right candidate from the top 16 would\nlift the acceptance length from 4.27 to 6.79. **That gap is pure selection\nheadroom.** We just need to select the right path through the candidates.\n\n[A Lightweight Path Selector](#a-lightweight-path-selector)\n\nCoherence is mostly local: a candidate's fit depends mainly on the token just before it, so scoring neighboring pairs should be enough. DFlash 2 keeps the top 16 candidates at each position and scores every adjacent pair: for predecessor and current candidate ,\n\nThe score has two parts. The first, , is DFlash's own logit: how much the drafter already liked on its own. The second asks how well follows : and give each token a compact 256-dimensional embedding, and the two embeddings are matched under a context gate that decides which parts of the match count. In essence, this is a low-rank bilinear attention over adjacent candidates.\n\nScoring stays fully parallel. Every adjacent pair at every position is scored in one shot, with no extra backbone or LM-head pass. The only sequential work is the final walk over precomputed scores: starting from the last verified token, greedy follows the best successor at each step, sampling draws from the same scores, and rejection sampling restores the exact target distribution.\n\n| Method | Params | Latency | T = 0 | T = 1 |\n|---|---|---|---|---|\n| DFlash | — | — | 4.27 | 3.78 |\n| + DSpark correction | +77.8M | +9.6% | 4.49 | 4.08 |\n| + path selection (ours) | +2.0M | +0.6% | 4.61 | 4.25 |\n\nThe selector improves DFlash by **0.34** tokens at and **0.47** at\n. It beats the DSpark correction in both settings with roughly 40×\nfewer parameters and 16× lower latency overhead. Choosing is cheaper than\npredicting. And there is still room: the oracle reaches 6.79. Pairwise\nscoring is the simplest selector we could think of, and we believe there\nis plenty to explore.\n\n[Suffix Decay Is a Local Problem](#suffix-decay-is-a-local-problem)\n\nWe also noticed\n[both recall rows above](#table-1) decline toward the end of the block.\nEven the oracle decays: with perfect selection, accuracy still falls from\n99.5% at the first position to 87.8% by the last. No selector can fix\nthat, because the candidates themselves are running out. We call this\n**suffix decay**, and it is a backbone problem.\n\nOne suspect is capacity: a five-layer backbone may be too small to preserve dependencies across the block. If that is right, depth should help most at later positions. And it does! 3-, 5-, and 15-layer DFlash models are almost identical at the first position, and fan apart down the block. But depth is indiscriminate: ten extra attention blocks add capacity everywhere, even at the early positions that had little left to gain, and erase much of the efficiency that makes DFlash attractive.\n\n| Draft position | 0 | 1 | 2 | 3 | 4 | 5 | 6 |\n|---|---|---|---|---|---|---|---|\n| DFlash 3L | 85.21% | 79.26% | 77.18% | 75.75% | 73.96% | 70.4% | 64.97% |\n| DFlash 5L | 85.39% | 80.31% | 79.39% | 78.27% | 77.39% | 76.03% | 72.86% |\n| DFlash 15L (3× more params) | 86.42% | 81.61% | 80.68% | 80.34% | 80.59% | 79.66% | 78.73% |\n| DFlash 5L + conv (+3% params) | 85.83% | 80.94% | 79.98% | 79.68% | 79.73% | 79.43% | 77.61% |\n\nWe want a targeted fix, and DFlash's attention shows where. It has two\njobs: read the context before the block, and model the dependencies\ninside. But it spends less and less on the second: the block's\nshare of attention falls from **30% in Layer 1 to 8% in Layer 5**, and\nwhat remains concentrates in [a shrinking handful of heads](#figure-3). So we split the\njobs: a dedicated module takes the within-block work, and attention keeps\nreading the context.\n\n| Attention head | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 |\n|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|\n| Layer 1 | 17.6% | 2.9% | 41.4% | 50.8% | 29.2% | 50.5% | 44.6% | 5.9% | 44.5% | 11.3% | 17.9% | 36.7% | 0.0% | 14.2% | 0.1% | 0.0% | 13.3% | 1.5% | 18.7% | 7.6% | 45.0% | 33.5% | 53.1% | 42.2% | 64.3% | 60.1% | 32.8% | 47.7% | 49.6% | 57.0% | 26.0% | 52.9% |\n| Layer 2 | 20.8% | 26.4% | 39.6% | 18.9% | 8.9% | 22.6% | 13.1% | 32.1% | 22.9% | 25.1% | 24.2% | 28.6% | 36.6% | 26.1% | 41.0% | 36.1% | 17.8% | 25.5% | 25.7% | 25.6% | 4.3% | 21.8% | 23.3% | 22.1% | 15.6% | 70.9% | 58.0% | 2.7% | 28.3% | 38.5% | 20.3% | 33.5% |\n| Layer 3 | 1.8% | 11.0% | 9.5% | 5.2% | 34.8% | 8.4% | 12.1% | 14.4% | 11.8% | 22.0% | 8.8% | 3.7% | 4.9% | 10.6% | 17.7% | 52.0% | 4.4% | 19.0% | 13.1% | 9.9% | 61.3% | 76.1% | 47.0% | 60.3% | 1.4% | 8.9% | 6.0% | 64.1% | 9.4% | 3.3% | 8.3% | 8.3% |\n| Layer 4 | 0.4% | 37.7% | 28.3% | 85.5% | 0.3% | 1.5% | 0.4% | 0.5% | 1.2% | 12.5% | 36.6% | 1.2% | 1.7% | 0.6% | 2.5% | 1.3% | 7.2% | 3.1% | 48.9% | 3.8% | 3.2% | 1.0% | 23.8% | 1.0% | 0.1% | 0.1% | 0.2% | 0.3% | 2.8% | 6.7% | 12.9% | 12.3% |\n| Layer 5 | 1.5% | 0.2% | 0.6% | 0.1% | 60.2% | 76.0% | 0.9% | 0.0% | 0.2% | 12.3% | 32.3% | 0.1% | 15.8% | 0.5% | 0.5% | 0.5% | 0.2% | 0.1% | 0.6% | 0.2% | 0.3% | 28.1% | 0.2% | 1.3% | 0.1% | 0.1% | 0.2% | 29.9% | 0.1% | 0.1% | 0.1% | 1.2% |\n\n[A Lightweight Local Convolution](#a-lightweight-local-convolution)\n\nThe within-block work is short-range to begin with: a block spans only\n4 to 16 tokens, and the tightest dependencies sit between neighbors. The\nnatural operator is a short convolution: two taps, one on the current\nposition and one reaching one position back, with weights that adapt to\nthe content. Following\n[Canon Layers](https://arxiv.org/abs/2512.17351),\n[Dynamic Short Convolutions](https://arxiv.org/abs/2606.03825), and\n[Convolution for Large Language Models](https://arxiv.org/abs/2607.18413),\nwe insert this two-tap dynamic depthwise convolution before and after each\nattention and feed-forward sublayer:\n\nEach coefficient combines a learned base kernel with a small correction computed from the current hidden state; every 16 channels share one correction. The first position reads the last verified token's representation, and every later position reads its predecessor's. Information crosses the block while all positions still compute in parallel.\n\nThe convolution is block-local and stateless, so it drops into DFlash without changing attention, the LM head, or verification.\n\nWith only **16.5M added parameters (3%)**, five-layer DFlash with\nconvolution [comes close to 15-layer DFlash](#figure-2), substantially\nreducing suffix decay. The convolutions add **0.7%** to draft–verify cycle\nlatency; ten more Transformer layers add 15.2%. Average within-block\nattention across Layers 4 and 5 also falls from **9.4% to 0.5%**,\nconsistent with the convolution absorbing the local work while attention\ngoes back to reading the context. A kernel reaching one position back\nrecovers most of what ten extra layers buy: suffix decay is mostly a\n*local* problem.\n\n[Putting It Together](#putting-it-together)\n\nSo far, the selector and the convolution have been measured separately;\n[the full comparison below](#table-3) puts them together. We trained the DFlash\nand DSpark drafters ourselves under matched setups, while MTP ships with\nthe model.\n\n| Dataset | MTP | DFlash | DSpark | DFlash 2 |\n|---|---|---|---|---|\n| GSM8K | 4.78 | 4.99 | 5.69 | 6.20 |\n| MATH-500 | 5.04 | 5.42 | 6.20 | 6.76 |\n| HumanEval | 4.84 | 5.43 | 5.80 | 6.28 |\n| MBPP | 4.16 | 4.49 | 4.96 | 5.41 |\n| MT-Bench | 3.90 | 4.26 | 4.77 | 5.20 |\n| Mean | 4.54 | 4.92 | 5.49 | 5.97 |\n\nDFlash 2 leads on every benchmark. Averaged across them, it gains\n**1.05 tokens over DFlash (21%)** and **0.48 over DSpark**. The upgrade\nstays cheap: the selector and the convolution together add only **1.3%**\nto the five-layer DFlash draft–verify cycle latency.\n\nOn MATH-500, [the gain is visible position by position](#figure-5):\nDFlash 2 holds steady near 86% to the last position, and every baseline\nends the block 6 to 9 points below it.\n\n| Draft position | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 |\n|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|\n| MTP | 84.57% | 80.23% | 79% | 78.42% | 78.63% | 78.17% | 77.36% | 77.74% | 77.91% | 76.96% | 78.06% | 77.4% | 77.49% | 77.48% | 77.85% |\n| DFlash | 88.35% | 77.7% | 77.8% | 79.45% | 80.3% | 81.12% | 81.22% | 81.07% | 81.29% | 80.28% | 80.64% | 80.29% | 79.56% | 78.77% | 77.48% |\n| DSpark | 87.24% | 84.59% | 83.79% | 83.63% | 83.6% | 83.27% | 82.97% | 82.54% | 82.21% | 82.39% | 81.58% | 80.7% | 81.35% | 80.57% | 79.86% |\n| DFlash 2 | 88.3% | 85.3% | 84.98% | 84.88% | 85.41% | 85.3% | 85.36% | 85.13% | 85.95% | 85.99% | 86.41% | 86.46% | 86.43% | 86.02% | 86.48% |\n\n[Two Drafters, Out Today](#two-drafters-out-today)\n\nWe are releasing two DFlash 2 drafters today:\n[one for Qwen3.8-27B](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2)\nand\n[one for Meta's Muse Glimmer](https://huggingface.co/incoai/Muse-Glimmer-30B-DFlash2).\nFor Qwen3.8-27B, we compare against the model's native MTP path and a\n[community DSpark drafter](https://huggingface.co/RadixArk/Qwen3.8-27B-DSpark).\n\n| Dataset | MTP | DSpark | DFlash 2 |\n|---|---|---|---|\n| GSM8K | 5.02 | 4.36 | 5.46 |\n| MATH-500 | 4.72 | 3.92 | 5.28 |\n| HumanEval | 3.91 | 3.30 | 4.39 |\n| MBPP | 3.99 | 3.51 | 4.79 |\n| MT-Bench | 3.74 | 3.01 | 4.10 |\n| Mean | 4.28 | 3.62 | 4.80 |\n\nFor Meta's Muse Glimmer, we compare against the official DFlash drafter\nshipped with the model and a\n[community DSpark drafter](https://huggingface.co/DaoCloud/Muse-Glimmer-30B-DSpark).\n\n| Dataset | DFlash | DSpark | DFlash 2 |\n|---|---|---|---|\n| GSM8K | 5.43 | 5.45 | 6.57 |\n| MATH-500 | 5.39 | 5.01 | 6.56 |\n| HumanEval | 4.11 | 4.33 | 5.66 |\n| MBPP | 3.74 | 4.02 | 5.30 |\n| MT-Bench | 3.52 | 3.59 | 4.42 |\n| Mean | 4.44 | 4.48 | 5.70 |\n\nThe margins are wide: on both models, DFlash 2 averages more than\n**a full token** ahead of DSpark. It also beats each model's official\ndrafter, MTP on Qwen3.8-27B and DFlash on Muse Glimmer. That translates\ninto **2.7–3.4×** the throughput of autoregressive decoding on\nQwen3.8-27B, and **3.1–4.6×** on Muse Glimmer. The\n[model cards](https://huggingface.co/collections/incoai/dflash-2-6a8432273c9998ce1685d4c5) break the speedups down\nby task and concurrency.\n\n[The Bottom Line](#the-bottom-line)\n\nAn agent writes in an afternoon what a chatbot writes in a month, and\ndecoding sits under every one of those tokens. DFlash 2 decodes at\n**close to 3× the speed of autoregressive decoding, about a third of the\ncompute per token**, with the same output.\n\nIn seven months, DFlash went from our paper to an industry standard, with more than 3.5 million downloads. Inside the same design, DFlash 2 decodes one more full token per pass, for free. That is only one component of the serving stack. Inference is nowhere near its floor.\n\nAt Inco AI, we are building an end-to-end serving stack to keep pushing that\nfloor lower. DFlash 2 is the first piece. Two drafters are out today\n[on Hugging Face](https://huggingface.co/collections/incoai/dflash-2-6a8432273c9998ce1685d4c5).\n\nIf you serve agents at scale and want to evaluate DFlash 2 in your stack,\nor want a drafter for a model you run, including your own fine-tunes, write\nto us: [contact@inco.ai](mailto:contact@inco.ai).\n\nWe are also hiring. If you want to help build this stack, reach out to us.\n\n**Connect the candidates. Keep drafting parallel.**\n\nGet updates\n\nOne email when we ship something new.\n\nWe will never share your email address.\n\n[Citation](#citation)\n\nPlease cite this post as:\n\n[Footnotes](#footnote-label)\n\n-\nModal's\n\n[\"Speculation Is All You Need\"](https://modal.com/blog/spec-is-all-u-need)points out that speculative decoding is the optimization that matters for low-latency serving. We are huge fans of their work and appreciate their support and discussions since DFlash's release.[↩](#user-content-fnref-modal)", "url": "https://wpnews.pro/news/dflash-2-keep-drafting-parallel", "canonical_source": "https://inco.ai/blog/dflash2/", "published_at": "2026-08-19 00:20:19+00:00", "updated_at": "2026-08-19 00:40:55.129938+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models"], "entities": ["Inco AI", "DFlash 2", "Qwen3.8-27B", "SGLang", "vLLM", "TensorRT-LLM", "llama.cpp", "NVIDIA"], "alternates": {"html": "https://wpnews.pro/news/dflash-2-keep-drafting-parallel", "markdown": "https://wpnews.pro/news/dflash-2-keep-drafting-parallel.md", "text": "https://wpnews.pro/news/dflash-2-keep-drafting-parallel.txt", "jsonld": "https://wpnews.pro/news/dflash-2-keep-drafting-parallel.jsonld"}}