Distributionally Robust Average-Reward Reinforcement Learning: Finite-Sample Guarantees under Weak Communication
We study distributionally robust reinforcement learning (DR-RL) in the average-reward setting under weak communication. Our main result provides finite-sample guarantees for estimating the robust optimal average reward and learning a near-optimal policy, covering both SA-rectangular and S-rectangular structures with di...