MineDraft, a batch parallel speculative decoding framework designed to effectively hide drafting latency by overlapping it with verification, is proposed, and the theoretical analysis shows that PSD is substantially more efficient than standard SD.
Zhenwei Tang, A. Verma, Zijian Zhou et al.· arXiv.org· 2 citations
This paper proposes algorithms that model online fair division as a contextual bandit problem and achieve provable sublinear regret and proposes algorithms that model utility is an unknown function of item-agent features.
A. Verma, Indrajit Saha, Makoto Yokoo et al.· 4 citations
This work describes the near-optimal region, the set of allocations within a specified tolerance of peak performance, which is wide even for small tolerances, widens with model scale, and transfers reliably from small proxy models to large target models.
Jingtan Wang, A. Verma, Xiaoqiang Lin et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.