{"slug": "robust-average-reward-markov-decision-processes-minimax-optimal-learning-via-in", "title": "Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions", "summary": "Researchers at arXiv introduced new minimax-optimal sample complexity bounds for learning epsilon-optimal robust policies in average-reward Markov decision processes under total-variation uncertainty. The study, published as arXiv:2608.06545v1, shows the total sample complexity scales as NSA ~ (SA/epsilon^2) times min{H0, H_sigma} in the high-tolerance regime and plus sigma H_sigma^2 in the low-tolerance regime, where H0 and H_sigma are nominal and robust optimal bias spans. The authors propose plug-in reduction procedures that achieve these rates, including a span-informed and a span-agnostic variant.", "body_md": "arXiv:2608.06545v1 Announce Type: new\nAbstract: Distributionally robust Markov decision processes provide a principled framework for sequential decision making under model uncertainty. We study how many samples are necessary and sufficient to learn an $\\varepsilon$-optimal robust policy under the average-reward criterion. A generative model provides samples from the nominal transition kernel, whereas policy performance is evaluated over $(s,a)$-rectangular total-variation uncertainty sets of radius at most $\\sigma$.\nLet $H_0$ and $H_\\sigma$ denote the nominal and robust optimal bias spans, respectively. We identify $\\sigma H_0$ as the perturbation scale separating high- and low-tolerance regimes. Our matching upper and lower bounds show that, up to logarithmic factors, the minimax total sample complexity is $$ NSA \\asymp \\frac{SA}{\\varepsilon^2}\\begin{cases} \\min\\{H_0,H_\\sigma\\}, & \\varepsilon\\gtrsim\\sigma H_0,\\\\ \\min\\{H_0,H_\\sigma\\}+\\sigma H_\\sigma^2, & \\varepsilon\\lesssim\\sigma H_0. \\end{cases} $$ Here $S$ and $A$ are the numbers of states and actions, and $N$ is the number of samples per state-action pair. The sample complexity consists of a linear-span term that resembles the nominal AMDP results and a robustness-specific term that appears only in the low-tolerance regime. We attain these rates using reduction-based plug-in procedures that select the reduction---nominal or robust---and its discount factor: a span-informed procedure that makes these choices using known span parameters, and a span-agnostic procedure that calibrates both choices from data.", "url": "https://wpnews.pro/news/robust-average-reward-markov-decision-processes-minimax-optimal-learning-via-in", "canonical_source": "https://arxiv.org/abs/2608.06545", "published_at": "2026-08-10 04:00:00+00:00", "updated_at": "2026-08-10 04:13:05.413101+00:00", "lang": "en", "topics": ["machine-learning", "artificial-intelligence"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/robust-average-reward-markov-decision-processes-minimax-optimal-learning-via-in", "markdown": "https://wpnews.pro/news/robust-average-reward-markov-decision-processes-minimax-optimal-learning-via-in.md", "text": "https://wpnews.pro/news/robust-average-reward-markov-decision-processes-minimax-optimal-learning-via-in.txt", "jsonld": "https://wpnews.pro/news/robust-average-reward-markov-decision-processes-minimax-optimal-learning-via-in.jsonld"}}