Why Deterministic PRM Guidance Underperforms in Discrete Diffusion Reasoning Process reward model (PRM) guidance underperforms in discrete diffusion language models (dLLMs) once denoising, PRM scoring, and outcome reward model (ORM) scoring are charged against the same budget of forward passes, according to research on dLLMs that expose a denoised solution at every step. The finding undercuts the assumption that PRM guidance is a straightforward way to spend additional test-time compute in discrete diffusion reasoning. Discrete diffusion language models dLLMs expose a denoised solution at every step, which makes process reward model PRM guidance look like a way to spend compute at test time. We show that once denoising, PRM scoring, and outcome reward model ORM scoring are charged in the same budget of forwa