WM-VLM: Probing Internal World Models for Interleaved Visual-Textual Reasoning Researchers introduced WM-VLM, a method that probes whether vision-language models can solve spatial problems by reasoning with both text and generated visual states rather than through language alone, according to the paper's abstract. The work contrasts conventional VLMs, which the authors say reason primarily through language, with an approach that interleaves visual and textual reasoning steps. Humans often solve spatial problems by mentally simulating visual transformations. In contrast, conventional vision-language models VLMs reason primarily through language. We investigate whether VLMs can solve spatial problems by reasoning with both text and generated visual states. To this end, w