Preprint
Aug 2026
Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions
This work studies how many samples are necessary and sufficient to learn an $\varepsilon$-optimal robust policy under the average-reward criterion and achieves these rates using reduction-based plug-in procedures that select the reduction---nominal or robust---and its discount factor.
Yue-Peng Yang, Yuxin Chen, Yuejie Chi
· 0 citations