Skip to content

Routing by Reasoning Need: Trajectory-Aware Decoding Control for Diffusion Vision-Language Models

Sep 2026 · 0 citations · 27 references
Computer Science

TL;DR

Across answer-focused, mixed-reasoning, and CoT-sensitive benchmarks, routed control improves robustness over fixed long decoding, pure short decoding, and single-rule interventions.

Abstract

Diffusion vision-language models generate answers through iterative refinement, exposing intermediate answer trajectories that can be inspected and controlled at inference time. However, this controllability creates a reasoning-need mismatch, where a universal generation length is applied to questions with different reasoning demands. Visually closed questions may be harmed by continued refinement after a stable answer has formed, whereas reasoning-sensitive questions may be harmed by premature commitment. We formulate this problem as reasoning-budget mismatch and study it in LLaDA-V. Rather than choosing a universal generation length, our training-free controller routes each example to early commitment, baseline preservation, or reasoning-supportive decoding using trajectory signals from answer closure, commitment evidence, and representation revision pressure, without using ground-truth answers. Across answer-focused, mixed-reasoning, and CoT-sensitive benchmarks, routed control improves robustness over fixed long decoding, pure short decoding, and single-rule interventions. The gains are not explained by shorter outputs alone. Answer-closed examples often benefit from commitment, whereas CoT-sensitive examples require preserving or supporting intermediate reasoning. Taken together, these results suggest diffusion VLM decoding should route inference-time control by the state suggested by the observed trajectory instead of relying on a universal decoding length.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

AutoCRAT: Within-trajectory Joint Control of Stochasticity and Compute for LLM Reasoning

This work instantiates AutoCRAT, a decoder-side controller for frozen backbones that operates over a discrete action space and updates control decisions only at semantic boundaries, improving stability while remaining responsive to the evolving reasoning process.

Han-Jun Luo, Qiu-Shi Liu, Jing-Yang Zhang et al. · 0 citations
#machine learning Preprint Sep 2026

Learning to Explain While Planning: Rule-Aligned Diffusion Planning for Autonomous Driving

Diffusion planners exhibit strong capabilities in generating multimodal trajectories. However, existing methods primarily rely on expert demonstrations to fit trajectory distributions, learning statistical correlations among scenes, behaviors, and trajectories without explicitly modeling driving rules. In long-tail sce...

Jia-Xi Ye, Chun-Ji Lv, Guo-Ren Wang et al. · 0 citations
Preprint Aug 2026

Selective Regenerative Decoding: Trajectory-Level Intervention for Inference-Time Reasoning

Inference-time decoding methods improve LLM reasoning by exploring multiple candidate trajectories, yet treat each trajectory as atomic: either retaining it whole or discarding it irreversibly. This wastes computation on partially promising candidates whose high-quality prefixes are abandoned alongside degraded suffixe...

Sophia Xiao Pu, Yumo Xu, Sailik Sengupta et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Answer-Distribution Trajectories: A Stochastic-Dynamics View of LLM Reasoning

This work introduces answer-distribution trajectories, a stochastic-dynamics-inspired representation that tracks the model's full predictive distribution over answers as reasoning unfolds, and suggests that answer-distribution trajectories provide a rich framework for analysing and evaluating the dynamics of LLM reason...

M. Catala, Haitz Sáez de Ocáriz Borde, D. Murari et al. · 0 citations
#natural language process... Preprint Sep 2026

Towards Efficient Reasoning: Learning Causal Shortcuts for Diffusion Language Models

This work introduces a step-by-step token extraction procedure to extract causal shortcuts from data, and applies parallel prioritized masking on these tokens during training to enable efficient and accurate convergence to correct answers via causal shortcuts.

Dian Jin, Kai-Rong Han, Bao-Hong Li et al. · 0 citations
#machine learning Preprint Sep 2026

Sampling via Decision-Flow: Training-Free Extraction of Improved Latent Reasoning Paths in Large Language Models

A central question in LLM reasoning is whether reinforcement learning (RL) instills genuinely new capabilities or merely reshapes how existing knowledge is expressed during inference. Building on the distribution-sharpening hypothesis, which holds that RL reallocates probability mass toward high-reward trajectories alr...

Zhen-Dong Mi, Shao-Yi Huang · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.