{"slug": "visage-constructing-self-correcting-memories-for-long-form-video-understanding", "title": "ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding", "summary": "Researchers propose ViSAGE, a multimodal agentic memory framework that constructs self-correcting, entity-centric memories for long-form video understanding, outperforming the strongest baseline by 5.9% in accuracy. The framework anchors entity identity via cross-modal binding, applies bidirectional memory refinement, and uses multi-agent cross-verification to abstain from unsupported answers.", "body_md": "arXiv:2607.28678v1 Announce Type: new\nAbstract: Multimodal agents operating in long-horizon environments must build and continually update multimedia memories to support entity-consistent, temporally grounded reasoning. However, existing agentic memory approaches often discard fine-grained dentity cues under aggressive compression and segment-wise processing. They also rely heavily on vector similarity retrieval, which can surface semantically related yet identity-mismatched evidence, leading to entity confusion, error propagation, and hallucinated answers.\nWe propose ViSAGE, a multimodal agentic memory framework that constructs self-correcting, entity-centric memories. Specifically, ViSAGE anchors entity identity via cross-modal binding over long temporal ranges. It then applies bidirectional memory refinement to propagate delayed identity evidence, retroactively unifying historical records and improving future reasoning. We also introduce multi-agent cross-verification to assess retrieved evidence under an identity-evidence alignment onstraint, enabling abstention instead of unsupported answers when evidence is missing. Extensive results demonstrate that ViSAGE consistently outperforms the strongest baseline, achieving 5.9% higher accuracy.", "url": "https://wpnews.pro/news/visage-constructing-self-correcting-memories-for-long-form-video-understanding", "canonical_source": "https://arxiv.org/abs/2607.28678", "published_at": "2026-08-03 04:00:00+00:00", "updated_at": "2026-08-03 04:13:11.796389+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "computer-vision", "ai-agents"], "entities": ["ViSAGE"], "alternates": {"html": "https://wpnews.pro/news/visage-constructing-self-correcting-memories-for-long-form-video-understanding", "markdown": "https://wpnews.pro/news/visage-constructing-self-correcting-memories-for-long-form-video-understanding.md", "text": "https://wpnews.pro/news/visage-constructing-self-correcting-memories-for-long-form-video-understanding.txt", "jsonld": "https://wpnews.pro/news/visage-constructing-self-correcting-memories-for-long-form-video-understanding.jsonld"}}