#machine learning
May 2026
Online Learning-to-Defer with Varying Experts
An online multiclass L2D algorithm that combines queried-action bandit feedback with a dynamically varying pool of experts is introduced that achieves expected true-deferral regret under a concentrated-score condition.
Duy Hoang Dang, Yannis Montreuil, Maxime Meyer et al.
· arXiv.org · 5 citations