Preprint
Aug 2026
Finite-Time Analysis of Discounted Exponential-Utility Reinforcement Learning
This work establishes finite-time rates of $\tilde{O} (1/\sqrt{n})$ for the aforementioned two algorithms under asynchronous Markovian sampling, where $n$ is the iteration index and $\tilde{O}$ hides logarithmic expressions.
Ankur Naskar, A. VivekT, Aditya Kumar et al.
· 0 citations