We study online convex optimization with dueling (pairwise comparison) feedback, where the learner observes only a binary preference between two queried points. While dueling feedback is well understood in discrete or stochastic settings, the adversarial convex setting has remained unexplored. We propose a simple reduction that converts dueling feedback into approximate gradients, enabling the use of standard first-order methods. We show that regret guarantees transfer under this reduction, yielding the first results for this setting, including $\mathcal{O}(T^{3/4})$ static, adaptive, and dynamic regret. Under additional structure, we obtain improved rates of $\mathcal{O}(T^{2/3})$ for smooth objectives and $\mathcal{O}(\sqrt{T \log T})$ for strongly convex functions.
Yiyang Lu, Hareshkumar Jadav, M. Pedramfar et al.· 1 citation
This work proposes Decentralized Barrier Follow-the-Regularized-Leader (Dec-BFTRL), and evaluates each agent's played action against the average of all local objectives, with applications to online continuous diminishing-return (DR) submodular maximization.
Yiyang Lu, M. Pedramfar, Vaneet Aggarwal· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.