{"slug": "encp-episode-normalized-conformal-prediction-for-vision-and-language-navigation", "title": "ENCP: Episode-Normalized Conformal Prediction for Vision-and-Language Navigation", "summary": "Researchers submitted a paper to arXiv on 15 Sep 2026 proposing Episode-Normalized Conformal Prediction (ENCP), a method that rescales a nonconformity score by a policy's residual confidence and calibrates one maximum score per episode to provide step-coverage guarantees for Vision-and-Language-Navigation (VLN) agents. Across four VLN policies and three nonconformity scores on the R2R and REVERIE datasets, ENCP met all reported empirical step-coverage targets on the seen-to-unseen evaluation, covering ground truth at every step with probability at least 1 - α. The authors state the model-agnostic uncertainty estimates could help determine when a VLN agent should defer to a more capable predictor, including human assistance.", "body_md": "# Computer Science > Machine Learning\n\n  [Submitted on 15 Sep 2026]\n\n# Title:ENCP: Episode-Normalized Conformal Prediction for Vision-and-Language Navigation\n\n[View PDF](http://arxiv.org/pdf/2609.17499v1)\n\n[HTML (experimental)](https://arxiv.org/html/2609.17499v1)\n\nAbstract:Uncertainty estimation for Vision-Language-Navigation (VLN) models is a critical task since it can help identify ambiguous and unreliable predictions, enabling agents to make safer navigation decisions. As one of the most advanced uncertainty estimation frameworks, conformal prediction (CP) offers a promising approach for uncertainty estimation in VLN. However, given that VLN agent requires a sequence of steps, standard calibration in conformal prediction fails to provide coverage guarantee it promises over a dependent, variable-length VLN episode. To this end, we propose Episode-Normalized Conformal Prediction (ENCP), which rescales a nonconformity score by the policy's residual confidence and calibrates one maximum score per episode. Under exchangeable calibration and test episodes, this construction covers the ground truth at every step with probability at least $1 - \\alpha$, while allowing dependence among steps within an episode. Across four VLN policies and three nonconformity scores on R2R and REVERIE dataset, ENCP meets all reported empirical step-coverage targets on the seen-to-unseen evaluation. These results demonstrate that ENCP can provide model-agnostic uncertainty estimates, which might be useful for determining when a VLN agent should defer to a more capable predictor, including human assistance.\n    \n\n### Current browse context:\n\ncs.LG\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\nIArxiv Recommender\n\n*(*[What is IArxiv?](https://iarxiv.org/about))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/encp-episode-normalized-conformal-prediction-for-vision-and-language-navigation", "canonical_source": "http://arxiv.org/abs/2609.17499v1", "published_at": "2026-09-16 14:21:45+00:00", "updated_at": "2026-09-16 14:42:28.046300+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "autonomous-vehicles", "ai-safety"], "entities": ["arXiv", "ENCP", "Vision-and-Language-Navigation", "R2R", "REVERIE"], "alternates": {"html": "https://wpnews.pro/news/encp-episode-normalized-conformal-prediction-for-vision-and-language-navigation", "markdown": "https://wpnews.pro/news/encp-episode-normalized-conformal-prediction-for-vision-and-language-navigation.md", "text": "https://wpnews.pro/news/encp-episode-normalized-conformal-prediction-for-vision-and-language-navigation.txt", "jsonld": "https://wpnews.pro/news/encp-episode-normalized-conformal-prediction-for-vision-and-language-navigation.jsonld"}}