{"slug": "pro-step-step-level-process-reward-optimization-for-retrieval-augmented", "title": "PRO-Step: Step-level Process Reward Optimization for Retrieval-Augmented Generation", "summary": "Researchers introduced PRO-STEP, a step-level process reward optimization method for retrieval-augmented generation that trains a generative process reward model to evaluate both logical validity and evidential grounding at each step, using PRM-guided value tree search and step-level direct preference optimization. On single and multi-hop question-answering benchmarks, PRO-STEP achieved the best average exact match and F1 scores across five datasets, addressing error propagation and spurious successes in multi-hop reasoning.", "body_md": "arXiv:2609.01658v1 Announce Type: new\nAbstract: Retrieval-Augmented Generation enhances Large Language Models by grounding responses in external knowledge, but multi-hop reasoning remains vulnerable to error propagation, where early retrieval failures confound subsequent steps. Standard outcome-based optimization only rewards the final answer, leaving intermediate retrieval and reasoning errors undetected. While existing process-based methods introduce step-level signals, they still score each step against the final answer, rewarding spurious successes where flawed retrieval coincidentally produces the correct answer. Step-level supervision in RAG requires evaluating both logical validity and evidential grounding at each step. We introduce PRO-STEP: we train a generative PRM that evaluates both dimensions, employ PRM-guided value tree search to construct preference pairs contrasting valid steps against flawed ones, and optimize the policy via step-level Direct Preference Optimization. Experiments on single and multi-hop QA datasets demonstrate that PRO-STEP achieves the best average EM and F1 across five benchmarks. Code, models, and training data are publicly available at https://github.com/keemminnke/PRO-Step.", "url": "https://wpnews.pro/news/pro-step-step-level-process-reward-optimization-for-retrieval-augmented", "canonical_source": "https://arxiv.org/abs/2609.01658", "published_at": "2026-09-03 04:00:00+00:00", "updated_at": "2026-09-03 04:25:23.267783+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research", "ai-tools"], "entities": ["PRO-STEP", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/pro-step-step-level-process-reward-optimization-for-retrieval-augmented", "markdown": "https://wpnews.pro/news/pro-step-step-level-process-reward-optimization-for-retrieval-augmented.md", "text": "https://wpnews.pro/news/pro-step-step-level-process-reward-optimization-for-retrieval-augmented.txt", "jsonld": "https://wpnews.pro/news/pro-step-step-level-process-reward-optimization-for-retrieval-augmented.jsonld"}}