{"slug": "finding-the-right-evidence-factor-guided-coarse-to-fine-reasoning-for-long", "title": "Finding the Right Evidence: Factor-Guided Coarse-to-Fine Reasoning for Long Videos", "summary": "Researchers at HKUST-KnowComp introduced PACE (Progressive Acquisition of Critical Evidence), a factor-guided framework for long-video question answering that improves evidence recovery and answer accuracy. On the MMR-V benchmark with the Qwen3-VL backbone, PACE achieved 42.6% accuracy, outperforming direct inference and prior agentic baselines such as Deep Video Discovery (DVD), and recovered 66.9% of annotated cues on a diagnostic subset. The framework also showed consistent gains over DVD on LVBench, Video-MME, EgoSchema, and LongVideoBench, with code available on GitHub.", "body_md": "arXiv:2608.26355v1 Announce Type: new\nAbstract: While LVLMs rapidly improve, long-video question answering still remains challenging: relevant evidence is sparse, and question-relevant context often fails to provide cues that discriminate the correct answer from plausible alternatives. Diagnostic analysis on a manually annotated subset of MMR-V shows that prior agentic systems substantially improve cue retrieval over direct VLM inference yet fail to achieve a corresponding gain in answer accuracy, indicating that the bottleneck lies in option-discriminative evidence rather than topical relevance alone. We propose PACE (Progressive Acquisition of Critical Evidence), a factor-guided framework for long-video evidence acquisition. PACE proceeds in two stages: it first indexes clip-level descriptions guided by question-derived factors without observing the candidate answers; it then uses the candidate answers to derive contrastive cues and queries the index for verification. On MMR-V with the open-source Qwen3-VL backbone, PACE achieves 42.6% accuracy, outperforming direct inference and prior agentic baselines including Deep Video Discovery (DVD). On the same diagnostic subset, PACE recovers 66.9% of the annotated cues, providing empirical evidence that its gains are associated with improved evidence recovery rather than stronger answer-side priors alone. Consistent gains over DVD on LVBench, Video-MME, EgoSchema, and LongVideoBench suggest that option-aware evidence acquisition transfers beyond MMR-V. Code is available at https://github.com/HKUST-KnowComp/PACE.", "url": "https://wpnews.pro/news/finding-the-right-evidence-factor-guided-coarse-to-fine-reasoning-for-long", "canonical_source": "https://arxiv.org/abs/2608.26355", "published_at": "2026-08-28 04:00:00+00:00", "updated_at": "2026-08-28 04:21:57.709152+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "computer-vision", "ai-research"], "entities": ["HKUST-KnowComp", "PACE", "MMR-V", "Qwen3-VL", "Deep Video Discovery (DVD)", "LVBench", "Video-MME", "EgoSchema"], "alternates": {"html": "https://wpnews.pro/news/finding-the-right-evidence-factor-guided-coarse-to-fine-reasoning-for-long", "markdown": "https://wpnews.pro/news/finding-the-right-evidence-factor-guided-coarse-to-fine-reasoning-for-long.md", "text": "https://wpnews.pro/news/finding-the-right-evidence-factor-guided-coarse-to-fine-reasoning-for-long.txt", "jsonld": "https://wpnews.pro/news/finding-the-right-evidence-factor-guided-coarse-to-fine-reasoning-for-long.jsonld"}}