DAVET: Denoising-Aware Visual Evidence Trajectory Allocation for Diffusion Vision-Language Models
Denoising-Aware Visual Evidence Trajectory Allocation (DAVET), a training-free framework that allocates visual evidence according to the evolving generation state, achieves an average speedup of 1.55 times with an average relative performance drop of 1.86%, showing that denoising-aware visual evidence allocation can re...