Bandit algorithms generate data for downstream inference, but adaptive sampling biases post-bandit sample means. We analyze this bias for stable index algorithms, including UCB1 and its generalizations, and derive sharp leading-order expressions for the sample-mean bias and expected $Z$-statistic. Our characterization reveals the algorithmic origin of bias through a key index-function-dependent quantity, which we term effective exploration rate. For example, under UCB1, the effective exploration rate is of order $\sqrt{\log T}$, and the standardized bias of any arm (that is not uniquely optimal) decays at the extremely slow rate $1/\sqrt{\log T}$. We also show how the choice of the index function affects both regret and bias, which reveals a regret-bias trade-off: more exploratory algorithm reduces bias but increases regret. Our sharp characterization for bias uses a novel empirical fluid approximation of the algorithm's sampling dynamics, which may be of independent interest.
This work develops a novel expansion representation for the OS value and proves that truncating this expansion yields a simulation-based algorithm that implements an optimal stopping policy with computational complexity scaling polynomially in the time horizon and the underlying dimension.
Yilun Chen, David A. Goldberg· Stochastic Systems· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.