World action models (WAMs) predict the future alongside actions during training. Due to the heavy computation cost of video denoising, whether the future must still be generated during inference is disputed: Explicit WAMs denoise it into clean frames along with every action chunk, whereas Latent WAM
Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration