04:00
2026-08-24
arxiv.org
artificial-intelligence
Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs
Researchers introduced OraRL, a reinforcement learning method for video multimodal large language models that uses annotations as oracle rollouts to improve sample efficiency and scalability. OraRL reβ¦