cd /news/artificial-intelligence/racer-reflective-agent-coupling-quer… · home › topics › artificial-intelligence › article
[ARTICLE · art-147340] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

RACER: Reflective Agent Coupling Query Interpretation and Tool-Based Retrieval for Frame Selection in Long Video Understanding

Researchers proposed RACER, a training-free reflective agentic framework that decomposes long-video frame selection into query interpretation by a lightweight video large language model and evidence localization by an embedding model used as a retrieval tool. RACER addresses the Query Comprehension Gap in similarity-based methods and the Interpretation-Selection Gap in judgment-based methods by having the Vid-LLM reformulate complex queries into sub-queries, feeding retrieved frames back for sub-query refinement in a reflection loop. Experiments across multiple benchmarks show RACER consistently improves long video understanding and achieves effective frame selection even with limited-capability components.

by read1 min views1 publishedOct 8, 2026

arXiv:2610.08954v1 Announce Type: new Abstract: Video large language models (Vid-LLMs) excel at diverse video-language tasks by reasoning over selected frames. However, frame selection for long videos remains challenging, as it requires retrieving relevant frames distributed across segments from a large candidate pool given complex queries. This paper investigates dominant approaches to long-video frame selection from a task-decomposition perspective, identifying two key challenges: the Query Comprehension Gap in similarity-based methods and the Interpretation--Selection Gap in judgment-based methods. To address them, we propose RACER, a training-free reflective agentic framework that decomposes long-video frame selection into query interpretation driven by a lightweight Vid-LLM and evidence localization supported by an embedding model serving as a retrieval tool. Specifically, the Vid-LLM is responsible solely for reformulating the complex query into sub-queries that make implicit information requirements explicit, mitigating the Query Comprehension Gap. Meanwhile, the retrieval tool leverages these sub-queries to localize relevant evidence, relieving the Vid-LLM of direct frame selection and thus addressing the Interpretation--Selection Gap. Finally, the retrieved frames are fed back to the Vid-LLM for sub-query refinement, forming a reflection loop that iteratively improves query interpretation and frame selection. Experiments across multiple benchmarks show that RACER consistently improves long video understanding. Notably, RACER achieves effective frame selection even with limited-capability components, demonstrating that agentic integration enables these components to enhance more capable Vid-LLMs.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @racer 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/racer-reflective-age…] indexed:0 read:1min 2026-10-08 · —