ADAS is proposed, a training-free reranking rule that leaves the base sampler's stopping rule unchanged and greedily discounts each token-wise confidence score according to its attention to already selected positions, weighted by their prediction uncertainty.
Y. Şahin, Ahmed R. Saikia, V. Cevher et al.· arXiv.org· 1 citation
Interpolating between these models, Raven is introduced, a linear-time sequence model that maintains a fixed set of memory slots and, at each step, decays and updates only a selected subset via learned, input-dependent routing, thereby preserving long-range content much more effectively.
Arshia Afzal, Aviv Bick, Eric P. Xing et al.· arXiv.org· 5 citations
This work introduces OVI, an interactive on-policy IL algorithm that is statistically efficient whenever the learner can represent the expert's value function and computationally efficient given access to a linear maximization oracle, and introduces a negative result showing that interaction is necessary.
Luca Viano, Antoine Moulin, Audrey Huang et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.