Skip to content
Preprint

LLaVAFlow: Preserving Latent Alignment Flow for Parameter-Efficient Multimodal Fine-Tuning

Aug 2026 · 0 citations · 48 references
Computer Science

TL;DR

This work argues that cross-modal alignment is implicitly captured in the information-compression trajectory, and proposes LLaVAFlow, an information-theoretic distillation framework that preserves alignment flow and enhances both downstream performance and generalization.

Abstract

While Multimodal Large Language Models (MLLMs) exhibit strong generalization, visual instruction tuning for downstream tasks inevitably causes catastrophic forgetting, impairing overall generalization. While existing methods regulate weight updates to reduce forgetting, they overlook the fundamental cross-modal alignment in MLLMs. Based on prior work and our observations, we argue that cross-modal alignment is implicitly captured in the information-compression trajectory. To preserve the alignment flow embedded in the trajectory, we propose LLaVAFlow, an information-theoretic distillation framework. First, we compress the mutual information between the extracted relations and MLLM embeddings, encouraging a learnable module to produce a refined alignment flow that benefits downstream tasks. Second, we maximize the mutual information between the extracted alignment flows of the pretrained and fine-tuned MLLMs, enabling the transfer of compact alignment information. Extensive experiments show that LLaVAFlow is an effective plug-and-play framework that preserves alignment flow and enhances both downstream performance and generalization.

View source

Similar papers

Aug 2026

GLA-LoRA: Parameter-efficient LLM fine-tuning with global-local knowledge alignment.

GLA-LoRA establishes a unified learning strategy that synergistically integrates multi-granular contrastive learning with knowledge distillation and establishes that explicit global-local knowledge alignment is essential for achieving high-fidelity, parameter-efficient fine-tuning across diverse language tasks.

Hao Wu, Jianqi Gao, Xiangfeng Luo · 0 citations
Preprint Jul 2026

Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models

The results suggest that competitive MLLM can emerge from alignment alone, reducing multimodal extension to a lightweight projector-training problem that generalizes across modalities and adapts rapidly to each new LLM release.

Xuanru Zhou, Yiwen Shao, Jiahong Li et al. · 1 citation
Preprint Jul 2026

LP-SFT: Local-Preserving Supervised Fine-Tuning via Multimodal Entropy Structure

LP-SFT, a Local-Preserving Supervised Fine-Tuning objective designed to explicitly protect this inherent entropy structure, improves overall performance over vanilla SFT and recent SFT-enhancement baselines, suggesting that local preservation helps mitigate capability degradation without collapsing sampling-accessible diversity.

Yueyang Wang, Baolong Bi, Shuo Lu et al. · 0 citations

CoDA: Co-Adaptive Dual-Path Alignment for Vision-Language Models

In CoDA, a new adaptation framework that explicitly disentangles and coordinates cross-modal semantic alignment and intra-modal structural consistency is proposed, and it is shown that CoDA outperforms state-of-the-art parameter-efficient methods, particularly under few-shot learning and distribution-shift scenarios.

Yi Zhang, Rui Zhu, Chan-Ni Li et al. · 0 citations

Edge-Efficient Compositional Recognition via Disentangled Prompt Tuning of Frozen Vision-Language Models

This work proposes Progressively Disentangled and Recurrent Prompt Tuning (PDRPT), an edge-efficient framework that decouples object and state updates before joint refinement, suppresses traction force from highly-entangled prompts, and preserves alignment with the natural language space of CLIP.

Xiaocheng Lu, Chuan He, Ziming Liu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

REIGN: Refurbished Embeddings with Integrated Guidance Networks for Efficient Context-Length Scaling

This work proposes REIGN (Refurbished Embeddings with Integrated Guidance Networks), a contrastively trained bi-encoder that operates on sequences of contextualised chunk embeddings from a frozen Guidance Network rather than on raw tokens for document-to-document retrieval.

Devrim Cavusoglu, Emre Akbas · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.