{"slug": "margins-not-windows-training-free-per-step-lossy-speculative-decoding", "title": "Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding", "summary": "Researchers introduced AdaptiveSpec, a training-free per-step speculative decoding method that adapts both draft verification and tree shape from internal decoding signals, improving throughput over EAGLE-3 by up to 56% on SGLang while recovering 93% to fully lossless task accuracy across GSM8K, MATH-500, and HumanEval on three models: DeepSeek-R1-Distill-Llama-8B, Llama-3.1-8B-Instruct, and Qwen3-8B.", "body_md": "arXiv:2609.02897v1 Announce Type: new\nAbstract: Speculative decoding accelerates LLM inference by drafting candidate tokens and verifying them in parallel. Tree-attention drafters such as EAGLE-3 are widely adopted, yet typically hold two decisions fixed: (1) a strict token-match verification rule and (2) a static draft-tree shape. Prior work relaxes each in isolation under limiting assumptions: long draft chains for training-free lossy verification, and adaptive tree shaping under a fixed token budget. We introduce AdaptiveSpec, a training-free per-step speculative decoding method that adapts both decisions from internal signals already produced during decoding. A per-step margin rule promotes a mismatched draft-proposed token when the ratio of the target's probability on the drafted token to its top-1 probability exceeds a threshold with no dependence on draft length or underlying drafter architecture. A per-step tree policy adjusts the draft tree's depth, width, and node count directly from a fused signal of draft top-1 confidence and a rolling acceptance history capturing recent draft-target agreement, allowing the total draft count to vary rather than only be redistributed. The two adaptations operate on orthogonal axes and compound in effect. Implemented on the SGLang production-grade serving engine, AdaptiveSpec improves throughput over the state-of-the-art autoregressive speculative decoding method EAGLE-3 by up to 56%, recovering 93% to fully lossless task accuracy across GSM8K, MATH-500, and HumanEval on three target models (DeepSeek-R1-Distill-Llama-8B, Llama-3.1-8B-Instruct, Qwen3-8B).", "url": "https://wpnews.pro/news/margins-not-windows-training-free-per-step-lossy-speculative-decoding", "canonical_source": "https://arxiv.org/abs/2609.02897", "published_at": "2026-09-04 04:00:00+00:00", "updated_at": "2026-09-04 04:22:09.051723+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research", "ai-infrastructure"], "entities": ["AdaptiveSpec", "EAGLE-3", "SGLang", "DeepSeek-R1-Distill-Llama-8B", "Llama-3.1-8B-Instruct", "Qwen3-8B", "GSM8K", "MATH-500"], "alternates": {"html": "https://wpnews.pro/news/margins-not-windows-training-free-per-step-lossy-speculative-decoding", "markdown": "https://wpnews.pro/news/margins-not-windows-training-free-per-step-lossy-speculative-decoding.md", "text": "https://wpnews.pro/news/margins-not-windows-training-free-per-step-lossy-speculative-decoding.txt", "jsonld": "https://wpnews.pro/news/margins-not-windows-training-free-per-step-lossy-speculative-decoding.jsonld"}}