Preprint
Jul 2026
Model-Free Q-Learning for Infinite-Horizon Stochastic Linear Quadratic Problems with Regime Switching
On-policy and off-policy Q-learning algorithms that learn the optimal controller solely from online state trajectory data are developed, specifically developing on-policy and off-policy Q-learning algorithms that learn the optimal controller solely from online state trajectory data.
Xinyu Zhang, Na Li, Xun Li et al.
· 0 citations