VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents VideoLoop, a new method for long-form video understanding, counters "semantic thrashing" in multimodal agents by looping working memory instead of appending to it, according to the paper's authors. The approach targets the collapse of attention to key evidence that occurs as append-only working memory grows over many reasoning steps, causing agents to lose access to previously gathered evidence. Long-form video understanding requires multimodal agents to iteratively gather evidence over many reasoning steps. However, most existing agentic methods suffer from semantic thrashing: as append-only working memory grows, attention to key evidence collapses, and the agent loses access to what it ha