04:00
2026-08-03
arxiv.org
artificial-intelligence
ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding
Researchers propose ViSAGE, a multimodal agentic memory framework that constructs self-correcting, entity-centric memories for long-form video understanding, outperforming the strongest baseline by 5.โฆ