Preprint
Aug 2026
Sharper Regret Bounds for Time-Varying Gaussian Process Bandits with Constant Exploration
Bayesian optimization in a time-varying environment where the unknown reward function evolves according to a Gaussian process drift model is studied, and GP-UCB can be run with a constant exploration parameter and obtained an expected-regret bound whose coefficient depends on the drift rate.
Matthias Mandl, Hanne Kekkonen
· 0 citations