Learning to Read the Contextual Tokens in Diffusion Transformers A new research paper introduces a framework for interpreting the dynamic contextual tokens that Multimodal Diffusion Transformers (MM-DiTs) form as they repeatedly update text tokens through multimodal attention during generation, addressing a function the authors say is not well understood. The work focuses on understanding how these models jointly process visual and textual representations throughout the generation process. Multimodal Diffusion Transformers MM-DiTs jointly process visual and textual representations throughout generation. These models repeatedly update the text tokens through multimodal attention, forming dynamic contextual tokens whose function is not well understood. In this work, we introduce a fra