{"slug": "near-optimal-sample-complexity-for-recursive-entropic-risk-reinforcement-with-a", "title": "Near-Optimal Sample Complexity for Recursive Entropic Risk Reinforcement Learning with a Generative Model", "summary": "A new arXiv paper (2610.06931v1) derives (ε,δ)-PAC guarantees for model-based risk-sensitive Q-value iteration (MB-RS-QVI) in finite discounted Markov decision processes under recursive entropic risk preferences with risk parameter β≠0, given access to a generative model. The analysis matches existing lower bounds in their exponential dependence on |β|/(1-γ) and in S, A, ε, and |β| up to logarithmic factors, removing the exponential gap between prior upper and lower bounds and leaving only a polynomial gap in the effective horizon 1/(1-γ).", "body_md": "arXiv:2610.06931v1 Announce Type: new \nAbstract: In this paper, we study the sample complexities of value and policy learning in finite discounted Markov decision processes (MDPs) under recursive entropic risk preferences with risk parameter $\\beta\\neq 0$, assuming access to a generative model of the MDP. We provide a refined analysis of model-based risk-sensitive Q-value iteration (MB-RS-QVI), a plug-in model-based method introduced in prior work, and derive $(\\varepsilon,\\delta)$-PAC guarantees for both learning the optimal $Q$-value function and an $\\varepsilon$-optimal policy. Our bounds improve the exponential dependence on the effective horizon $1/(1-\\gamma)$ compared with the best existing guarantees for this setting. In particular, they match the existing lower bounds in their exponential dependence on $|\\beta|/(1-\\gamma)$, as well as in $S$, $A$, $\\varepsilon$, and $|\\beta|$, up to logarithmic factors. Consequently, our analysis removes the exponential gap between the previously known upper and lower bounds, leaving only a polynomial gap in the effective horizon.", "url": "https://wpnews.pro/news/near-optimal-sample-complexity-for-recursive-entropic-risk-reinforcement-with-a", "canonical_source": "https://arxiv.org/abs/2610.06931", "published_at": "2026-10-07 04:00:00+00:00", "updated_at": "2026-10-07 04:16:52.280866+00:00", "lang": "en", "topics": ["machine-learning", "ai-research"], "entities": ["arXiv", "MB-RS-QVI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/near-optimal-sample-complexity-for-recursive-entropic-risk-reinforcement-with-a", "markdown": "https://wpnews.pro/news/near-optimal-sample-complexity-for-recursive-entropic-risk-reinforcement-with-a.md", "text": "https://wpnews.pro/news/near-optimal-sample-complexity-for-recursive-entropic-risk-reinforcement-with-a.txt", "jsonld": "https://wpnews.pro/news/near-optimal-sample-complexity-for-recursive-entropic-risk-reinforcement-with-a.jsonld"}}