Nonlinear Bandit
An algorithm EHM is proposed that extends the adaptive Huber loss method with one-pass update with one-pass update and achieves an almost optimal regret of $\widetilde{\mathcal{O}(1)$ computational complexity with respect to current round $t$ and the time horizon $T$), which simultaneously achieves an almost optimal regret.