Bounded unfiltered teacher continuations at learner-induced contexts improve over pure behavioral cloning at matched budgets and suggest that a few teacher steps, placed at learner-induced contexts, can be a more cost-efficient supervision allocation than longer or more heavily curated teacher completions.
Junze Ye, Jiayi Cheng, Miao Lu et al.· arXiv.org· 2 citations
This work proposes a statistically consistent and scalable estimator for score differences based on Sobolev regularization and demonstrates its effectiveness on real-world tasks, including transfer learning for ECG signal generation, where it substantially outperforms non-regularized score difference estimators in downstream classification performance.
Chenghan Xie, Jose H. Blanchet, Renyuan Xu· 0 citations
This work uncover and identify the notion of ``robust item-wise coverage''as the minimal data requirement to enable sample-efficient robust assortment learning and bridges the gap between robustness and statistical efficiency in assortment learning.
Miao Lu, Yuxuan Han, Han Zhong et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.