{"slug": "verifine-scaling-verification-for-self-improvement-in-embodied-reasoning", "title": "VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning", "summary": "Researchers submitted VeriFine, an agent harness framework for scaling verification in embodied reasoning, to arXiv on 6 October 2026. VeriFine co-evolves the policy, training curriculum, and judge through a Policy Improvement Loop that uses a rubric judge to diagnose recurring failures and an Judge Improvement Loop that queries human guidance on informative failure cases and refines the judge via coactive calibration. Experiments on driving and robot navigation tasks showed continuous self-improvement in both policy and judge capability across reinforcement and supervised fine-tuning.", "body_md": "# Computer Science > Artificial Intelligence\n\n  [Submitted on 6 Oct 2026]\n\n# Title:VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning\n\n[View PDF](http://arxiv.org/pdf/2610.08761v1)\n\n[HTML (experimental)](https://arxiv.org/html/2610.08761v1)\n\nAbstract:Self-improving policies continually expose new failure patterns, changing what their judges must be able to verify. However, current fixed judges constrain both optimization feedback and the discovery of useful training examples, limiting further self-improvement. This challenge is even more acute in embodied reasoning, where reliable evaluation must account for spatial grounding, causal reasoning, and safety-aware decision-making. We introduce VeriFine, an agent harness framework that scales verification through the co-evolution of the policy, training curriculum, and judge. The Policy Improvement Loop uses a rubric judge to diagnose recurring failures, construct an adaptive curriculum, and optimize the policy. When progress plateaus and verification becomes a bottleneck, the Judge Improvement Loop selectively queries human guidance on informative failure cases and refines the judge through coactive calibration, in which humans and agents resolve disagreements and converge toward the objective rubric of physical reasoning. The revised judge then guides the next stage of data selection and policy optimization. Experiments on driving and robot navigation tasks demonstrate continuous self-improvement in both policy and judge capability across reinforcement and supervised fine-tuning. These results show how scaling verification supports continuous self-improvement as policy failure patterns evolve.\n    \n\n### Additional Features\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/verifine-scaling-verification-for-self-improvement-in-embodied-reasoning", "canonical_source": "http://arxiv.org/abs/2610.08761v1", "published_at": "2026-10-07 14:07:50+00:00", "updated_at": "2026-10-07 14:17:36.104347+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-agents", "robotics", "autonomous-vehicles"], "entities": ["VeriFine", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/verifine-scaling-verification-for-self-improvement-in-embodied-reasoning", "markdown": "https://wpnews.pro/news/verifine-scaling-verification-for-self-improvement-in-embodied-reasoning.md", "text": "https://wpnews.pro/news/verifine-scaling-verification-for-self-improvement-in-embodied-reasoning.txt", "jsonld": "https://wpnews.pro/news/verifine-scaling-verification-for-self-improvement-in-embodied-reasoning.jsonld"}}