Zeroth-Order Nonsmooth Nonconvex Optimization with Convex Liftings and Its Application to State-Feedback $H_\infty$ Policy Optimization
This work proposes a zeroth-order proximal point algorithm and verifies that the assumptions underlying the analysis hold for discrete-time state-feedback state-feedback policy optimization, yielding an oracle complexity of $\widetilde{O}\left(n_u n_x\epsilon^{-3}\right)$ for attaining a prescribed objective value gap.