{"slug": "denoising-surface-modeling-and-predicting-inference-cost-for-diffusion-llm", "title": "Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving", "summary": "A new arXiv paper (2610.00499v1) proposes the Denoising Workload Surface (DWS), a two-dimensional block-step probability surface that predicts per-request inference cost for diffusion large language models (dLLMs). In real-world serving experiments, DWS reduced cost-prediction error by up to 2.50x over scalar-based predictors, and a DWS-guided shortest-job-first scheduler cut end-to-end latency by up to 1.92x for online chatbots. The predictor runs on a single CPU core and transfers across hardware configurations without retraining.", "body_md": "arXiv:2610.00499v1 Announce Type: new \nAbstract: As diffusion large language models (dLLMs) become more capable, they are moving from research settings to real-world \\textit{serving}, where request management (such as scheduling and resource allocation) relies on accurate estimation of per-request inference cost. However, common cost proxies fall short for dLLMs: output length ignores that one forward pass can unmask multiple tokens, and denoising-step count ignores the \\textit{heterogeneous} per-step costs. We observe that the block-autoregressive generation mechanism induces a two-dimensional execution structure over output blocks and within-block denoising steps, whereas these proxies collapse it into a scalar, discarding information essential for characterizing the cost. Motivated by this insight, we propose the Denoising Workload Surface (DWS), which preserves this two-dimensional block-step structure as a probability surface to weight the heterogeneous per-step costs. We then design a coarse-to-fine training scheme that enables a lightweight prompt-only predictor to accurately predict the complex DWS. This predictor runs efficiently even on a single CPU core, avoiding GPU contention with the serving model. Since DWS decouples request-dependent execution behavior from deployment-specific cost factors, the predictor transfers across hardware configurations without retraining. In \\textit{real-world} serving experiments, DWS reduces cost-prediction error by up to $2.50\\times$ over scalar-based predictors, while the DWS-guided shortest-job-first scheduler reduces end-to-end latency by up to $1.92\\times$ for online chatbots.", "url": "https://wpnews.pro/news/denoising-surface-modeling-and-predicting-inference-cost-for-diffusion-llm", "canonical_source": "https://www.machinebrief.com/news/denoising-surface-modeling-and-predicting-inference-cost-for-a2ry", "published_at": "2026-10-03 04:00:00+00:00", "updated_at": "2026-10-03 05:38:12.133421+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "mlops", "ai-research", "machine-learning"], "entities": ["Denoising Workload Surface", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/denoising-surface-modeling-and-predicting-inference-cost-for-diffusion-llm", "markdown": "https://wpnews.pro/news/denoising-surface-modeling-and-predicting-inference-cost-for-diffusion-llm.md", "text": "https://wpnews.pro/news/denoising-surface-modeling-and-predicting-inference-cost-for-diffusion-llm.txt", "jsonld": "https://wpnews.pro/news/denoising-surface-modeling-and-predicting-inference-cost-for-diffusion-llm.jsonld"}}