ENCP: Episode-Normalized Conformal Prediction for Vision-and-Language Navigation Researchers submitted a paper to arXiv on 15 Sep 2026 proposing Episode-Normalized Conformal Prediction (ENCP), a method that rescales a nonconformity score by a policy's residual confidence and calibrates one maximum score per episode to provide step-coverage guarantees for Vision-and-Language-Navigation (VLN) agents. Across four VLN policies and three nonconformity scores on the R2R and REVERIE datasets, ENCP met all reported empirical step-coverage targets on the seen-to-unseen evaluation, covering ground truth at every step with probability at least 1 - α. The authors state the model-agnostic uncertainty estimates could help determine when a VLN agent should defer to a more capable predictor, including human assistance. Computer Science Machine Learning Submitted on 15 Sep 2026 Title:ENCP: Episode-Normalized Conformal Prediction for Vision-and-Language Navigation View PDF http://arxiv.org/pdf/2609.17499v1 HTML experimental https://arxiv.org/html/2609.17499v1 Abstract:Uncertainty estimation for Vision-Language-Navigation VLN models is a critical task since it can help identify ambiguous and unreliable predictions, enabling agents to make safer navigation decisions. As one of the most advanced uncertainty estimation frameworks, conformal prediction CP offers a promising approach for uncertainty estimation in VLN. However, given that VLN agent requires a sequence of steps, standard calibration in conformal prediction fails to provide coverage guarantee it promises over a dependent, variable-length VLN episode. To this end, we propose Episode-Normalized Conformal Prediction ENCP , which rescales a nonconformity score by the policy's residual confidence and calibrates one maximum score per episode. Under exchangeable calibration and test episodes, this construction covers the ground truth at every step with probability at least $1 - \alpha$, while allowing dependence among steps within an episode. Across four VLN policies and three nonconformity scores on R2R and REVERIE dataset, ENCP meets all reported empirical step-coverage targets on the seen-to-unseen evaluation. These results demonstrate that ENCP can provide model-agnostic uncertainty estimates, which might be useful for determining when a VLN agent should defer to a more capable predictor, including human assistance. Current browse context: cs.LG References & Citations Loading... Bibliographic and Citation Tools Bibliographic Explorer What is the Explorer? https://info.arxiv.org/labs/showcase.html arxiv-bibliographic-explorer Connected Papers What is Connected Papers? https://www.connectedpapers.com/about Litmaps What is Litmaps? https://www.litmaps.co/ scite Smart Citations What are Smart Citations? https://www.scite.ai/ Code, Data and Media Associated with this Article alphaXiv What is alphaXiv? https://alphaxiv.org/ CatalyzeX Code Finder for Papers What is CatalyzeX? https://www.catalyzex.com DagsHub What is DagsHub? https://dagshub.com/ Gotit.pub What is GotitPub? http://gotit.pub/faq Hugging Face What is Huggingface? https://huggingface.co/huggingface ScienceCast What is ScienceCast? https://sciencecast.org/welcome Demos Recommenders and Search Tools Influence Flower What are Influence Flowers? https://influencemap.cmlab.dev/ CORE Recommender What is CORE? https://core.ac.uk/services/recommender IArxiv Recommender What is IArxiv? https://iarxiv.org/about arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs https://info.arxiv.org/labs/index.html .