Skip to content

TTRSD: Test-Time Reinforcement Learning with Self-Distillation for Vision-Language Models

Sep 2026 · 0 citations · 42 references
Computer Science

TL;DR

TTRSD separates update direction from update position, determined by group-relative advantages, from update position, determined by visual sensitivity, without requiring ground-truth labels, external verifiers, or a separate teacher, demonstrating cross-dataset generalization while preserving inherent reasoning integrity.

Abstract

Test-time reinforcement learning enables vision-language models (VLMs) to adapt using unlabeled inputs. However, repeated sampling under fixed visual conditions can reinforce shared perceptual errors, while sequence-level rewards fail to isolate visual perception the foundational bottleneck that anchors multimodal reasoning risking the degradation of pre-trained reasoning capabilities. We propose TTRSD, a test-time reinforcement learning framework combining multi-view answer-level self-distillation with visual contrastive token selection. A shared policy aggregates teacher predictions across original, cropped, and downsampled views into an answer distribution. Student trajectories generated from the original image receive rewards based on the support for their final answers in this distribution. To allocate this feedback precisely toward perceptual bottlenecks, we compare the log-probabilities of the same sampled tokens under original and visually ablated inputs while holding their textual prefixes fixed, selecting visually sensitive positions for policy-gradient updates. TTRSD separates update direction, determined by group-relative advantages, from update position, determined by visual sensitivity, without requiring ground-truth labels, external verifiers, or a separate teacher. With only 20 unlabeled adaptation samples, TTRSD improves performance across seven benchmarks and three VLMs, raising InternVL3-2B's MMMU accuracy from 35.79% to 49.32%(+13.53%), demonstrating cross-dataset generalization while preserving inherent reasoning integrity.

View source

Similar papers

#machine learning Preprint Sep 2026

Harnessing Image Question Dependence for Better VLM Test-time Reinforcement Learning

Test-time reinforcement learning can adapt vision-language models (VLMs) to unlabeled target data, but its effectiveness is fundamentally limited by the reliability of self-generated learning signals. To assess the reliability of consensus-based learning signals, we analyze VLM test-time reinforcement learning across d...

Xinrui He, Ting-Wei Li, Jun-Ting Wang et al. · 0 citations
#small language model Preprint Sep 2026

SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models

Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference expensive. Training-free token pruning reduces this cost, but aggressive compression can sharply degrade performance, often attributed to irreversible loss of task-relevant visual information. We show...

Ahmadreza Jeddi, En-Ming Zhang, Jasper Gerigk et al. · 0 citations
Preprint Aug 2026

Perturb the Thought, Not the Pixels: Latent-Space Rollout Diversification for Reinforcement Learning of Vision-Language Models

Noise-Contrastive GRPO is introduced, which injects scale-calibrated Gaussian noise into the last hidden layer of the prompt-encoding pass for half of each rollout group, branching those rollouts from a displaced departure state and integrates into a standard RLVR pipeline as a ~50-line change to the inference engine.

Michael M. Jerge, Joseph Pelczar, J. Downes · 1 citation
#artificial intelligence Preprint Sep 2026

Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces

Test-time reinforcement learning (TTRL) enables models to improve their reasoning without relying on labeled training data, but existing approaches typically optimize a large fraction of the model parameters. This raises a natural question: can effective test-time adaptation emerge when both the reward signal and the o...

Naveen Vakada, Ming-Yuan Li, Shao-Xiong Ji · 0 citations
#machine learning Preprint Sep 2026

Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards

Gradient-Aligned Reward (GAR), which operates in the policy's own gradient space: truncated backpropagation through the output projection layer extracts a compact gradient vector for each rollout, and cosine similarity with an expert-anchor gradient yields a dense, reasoning-aware reward with less than 9% wall-clock ov...

Le-Qi Zheng, Jin-Bo Su, Fang Niu et al. · 2 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.