ReMoMask-2: Latent Retrieval-Augmented Masked Motion Generation Researchers introduced ReMoMask-2, a latent retrieval-augmented masked motion generation model for text-to-motion (T2M) generation, which maps natural language to human joint movements for gaming, VR, and robotics. The work targets two challenges in existing Retrieval-Augmented Text-to-Motion (RAG-T2M) models, which condition generation on retrieved motion-text pairs to improve results on complex descriptions. Text-to-motion T2M generation maps natural language to human joint movements, aiding gaming, VR, and robotics. Retrieval-Augmented Text-to-Motion RAG-T2M improves generation on complex descriptions by conditioning on retrieved motion-text pairs. However, existing RAG-T2M models face two challeng