# Near-Optimal Sample Complexity for Recursive Entropic Risk Reinforcement Learning with a Generative Model

> Source: <https://arxiv.org/abs/2610.06931>
> Published: 2026-10-07 04:00:00+00:00

arXiv:2610.06931v1 Announce Type: new 
Abstract: In this paper, we study the sample complexities of value and policy learning in finite discounted Markov decision processes (MDPs) under recursive entropic risk preferences with risk parameter $\beta\neq 0$, assuming access to a generative model of the MDP. We provide a refined analysis of model-based risk-sensitive Q-value iteration (MB-RS-QVI), a plug-in model-based method introduced in prior work, and derive $(\varepsilon,\delta)$-PAC guarantees for both learning the optimal $Q$-value function and an $\varepsilon$-optimal policy. Our bounds improve the exponential dependence on the effective horizon $1/(1-\gamma)$ compared with the best existing guarantees for this setting. In particular, they match the existing lower bounds in their exponential dependence on $|\beta|/(1-\gamma)$, as well as in $S$, $A$, $\varepsilon$, and $|\beta|$, up to logarithmic factors. Consequently, our analysis removes the exponential gap between the previously known upper and lower bounds, leaving only a polynomial gap in the effective horizon.
