Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models
ADAS is proposed, a training-free reranking rule that leaves the base sampler's stopping rule unchanged and greedily discounts each token-wise confidence score according to its attention to already selected positions, weighted by their prediction uncertainty.