Skip to content
Preprint

DAVET: Denoising-Aware Visual Evidence Trajectory Allocation for Diffusion Vision-Language Models

Aug 2026 · 0 citations · 28 references
Computer Science

TL;DR

Denoising-Aware Visual Evidence Trajectory Allocation (DAVET), a training-free framework that allocates visual evidence according to the evolving generation state, achieves an average speedup of 1.55 times with an average relative performance drop of 1.86%, showing that denoising-aware visual evidence allocation can reduce visual conditioning cost while largely preserving generation quality.

Abstract

Diffusion vision-language models (dVLMs) iteratively denoise masked responses while conditioning each denoising step on visual evidence, making visual conditioning a substantial recurring inference cost. Unlike autoregressive decoding, diffusion generation repeatedly revisits the entire response as uncertainty evolves. Our analysis reveals that visual evidence demand is strongly step-dependent, motivating adaptive allocation across denoising steps. Existing inference acceleration methods operate through decoding-side strategies or visual token compression via pruning and merging, but do not explicitly treat visual evidence as a resource whose demand evolves across the diffusion process. Therefore, we present Denoising-Aware Visual Evidence Trajectory Allocation (DAVET), a training-free framework that allocates visual evidence according to the evolving generation state. Starting from a phase-conditioned evidence trajectory, the proposed allocation policy uses operation demand to set an evidence reserve whose allocation at each denoising step is modulated by trajectory risk. DAVET realizes the resulting budgets through a hierarchy of evidence views constructed from a single visual encoding, separating when and how much evidence is needed from how the evidence views are constructed. Evaluated on two representative dVLMs, LLaDA-V and LaViDa, across multiple visual-understanding benchmarks, DAVET achieves an average speedup of 1.55$\times$ with an average relative performance drop of 1.86\%, showing that denoising-aware visual evidence allocation can reduce visual conditioning cost while largely preserving generation quality.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

CAST: Causal Advantage-Structured Training with Spatially Grounded Compositional Rewards for Diffusion Models

Online reinforcement learning has been extended to flow matching for diffusion model (DM) image generation. However, this paradigm faces three limitations: (1) Window selection. Existing methods manually set the stochastic differential equation (SDE) sampling window, i.e., the denoising steps where exploration noise is...

Yu Shu, Chao-Chao Lu · 0 citations
#natural language process... Preprint Sep 2026

Informed Masking: Structure-Aware Perturbation for Reinforcement Learning in Diffusion Large Language Models

Diffusion Large Language Models (dLLMs) have emerged as an efficient alternative to autoregressive models, yet aligning them via Reinforcement Learning (RL) requires likelihood surrogates estimated from masked reconstruction subproblems under a small Monte Carlo budget per rollout. Existing methods construct these subp...

Xiao-Yi Yu, E. Sangineto, Pei Fu et al. · 0 citations
Preprint Aug 2026

Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models

The Visual Information-Guided Sampler (VIG-Sampler), which prioritizes tokens based on their attention to image tokens, and imposes a constraint that penalizes candidate tokens whose image-attention distributions are similar to those of previously selected tokens, thereby increasing the information gain of the decoded...

Insu Lee, Woo-Soon Park, Wonseok Shin et al. · 1 citation
Preprint Sep 2026

JEPA Guided Diffusion: Predictive Vision-Language Conditioning for Generative Traffic Forecasting

Accurate traffic forecasting requires both understanding scene dynamics and synthesizing realistic future observations. Recent diffusion-based video generation models produce visually plausible predictions but require expensive end-to-end training and often entangle scene understanding with image synthesis. In this wor...

T. Nguyen, Thanh-Nhan Vo, Bui Hoai Thuong Nguyen et al. · 1 citation
Conference Sep 2026

Raise One and Infer Three: Toward Reasoning- and Memory-Augmented Diffusion Policy Generalization

This work proposes "raise one and infer three" diffusion policy (ROITDP), a novel approach that introduces two complementary mechanisms, including a reasoning mechanism built upon the Chain-of-Skill Noise Watermark, which enables temporally coherent multi-step reasoning throughout the diffusion process under distributi...

Yi-Hang Zhu, Yuxuan Wang, Tong Li et al. · 0 citations
Preprint Sep 2026

Learning via Self-Consistency for Diffusion-based Video Reasoning

Video generation models have demonstrated emerging zero-shot capabilities for visual reasoning, perception, and other vision tasks. However, diffusion-based video generation is inherently stochastic, while many downstream vision tasks are deterministic. Motivated by the effectiveness of self-consistency in chain-of-tho...

Zheng-Hao Ni, Wei-Min Qiu, Meng Tang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.