cd /news/large-language-models/denoising-surface-modeling-and-predi… · home › topics › large-language-models › article
[ARTICLE · art-144310] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving

A new arXiv paper (2610.00499v1) proposes the Denoising Workload Surface (DWS), a two-dimensional block-step probability surface that predicts per-request inference cost for diffusion large language models (dLLMs). In real-world serving experiments, DWS reduced cost-prediction error by up to 2.50x over scalar-based predictors, and a DWS-guided shortest-job-first scheduler cut end-to-end latency by up to 1.92x for online chatbots. The predictor runs on a single CPU core and transfers across hardware configurations without retraining.

by read1 min views1 publishedOct 3, 2026

arXiv:2610.00499v1 Announce Type: new Abstract: As diffusion large language models (dLLMs) become more capable, they are moving from research settings to real-world \textit{serving}, where request management (such as scheduling and resource allocation) relies on accurate estimation of per-request inference cost. However, common cost proxies fall short for dLLMs: output length ignores that one forward pass can unmask multiple tokens, and denoising-step count ignores the \textit{heterogeneous} per-step costs. We observe that the block-autoregressive generation mechanism induces a two-dimensional execution structure over output blocks and within-block denoising steps, whereas these proxies collapse it into a scalar, discarding information essential for characterizing the cost. Motivated by this insight, we propose the Denoising Workload Surface (DWS), which preserves this two-dimensional block-step structure as a probability surface to weight the heterogeneous per-step costs. We then design a coarse-to-fine training scheme that enables a lightweight prompt-only predictor to accurately predict the complex DWS. This predictor runs efficiently even on a single CPU core, avoiding GPU contention with the serving model. Since DWS decouples request-dependent execution behavior from deployment-specific cost factors, the predictor transfers across hardware configurations without retraining. In \textit{real-world} serving experiments, DWS reduces cost-prediction error by up to $2.50\times$ over scalar-based predictors, while the DWS-guided shortest-job-first scheduler reduces end-to-end latency by up to $1.92\times$ for online chatbots.

── more in #large-language-models 4 stories · sorted by recency
── more on @denoising workload surface 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/denoising-surface-mo…] indexed:0 read:1min 2026-10-03 · —