Toward the Optimal Regret-Instability Trade-off in Multi-Armed Bandits
The results resolve the open question raised in the literature concerning the sharp arm-dependent regret--instability frontier and develop a new offline top-prefix representation that removes path dependence from online decisions.
Kaifei Wang, Y. Ye, Han Zhong
· 0 citations