Skip to content

Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering

Jul 2026 · arXiv.org · Vol abs/2607.26411 · 0 citations · 38 references
Computer Science

TL;DR

It is shown that steering vectors learned from the understanding branch can transfer to generation, enabling controllable image synthesis and improved semantic faithfulness, and establish cross-branch steering as a practical tool for probing multimodal representations.

Abstract

Unified multimodal models (UMMs) aim to integrate understanding and generation within a single architecture, yet it remains unclear whether these capabilities share a unified and transferable semantic space. This question is fundamentally challenging, as the two branches operate over heterogeneous representations (text tokens vs.\ visual latents) and distinct training objectives, making direct comparison difficult. To address this, we introduce \emph{cross-branch semantic steering}, an intervention-based framework that extracts semantic directions from one branch and applies them to the other. We show that steering vectors learned from the understanding branch can transfer to generation, enabling controllable image synthesis and improved semantic faithfulness. In contrast, the reverse direction consistently shows limited effectiveness. Our analysis suggests that this asymmetry may be related to a practical representational mismatch: understanding-derived vectors capture transferable, object-centric semantics, while generation-derived vectors primarily encode low-level appearance features. Our results reveal that architectural unification does not guarantee semantic alignment, and establish cross-branch steering as a practical tool for probing multimodal representations.

View source

Similar papers

Preprint Jul 2026

Transferability Between Understanding and Generation in Unified Multimodal Models

This work empirically finds that transferability depends on architecture-models with fully shared transformer backbone and a unified visual encoder exhibit consistent cross-task transfer, while loosely coupled designs show little or none.

Jiwon Kang, Heeji Yoon, Jaewoo Jung et al. · 1 citation
Jul 2026

Twins: Learn to Predict Unified Representations with Focal Loss

Twin, a unified continuous token space formed by channel-wise concatenating ViT and VAE features on the same token grid, so the sequence length is unchanged and attention cost does not increase is proposed, which performs competitively on multimodal understanding benchmarks and improves reconstruction fidelity.

Kaixiong Gong, Xin Cai, Bin Lin et al. · 0 citations
Preprint Aug 2026

When Does Visual Generation Help Visual Understanding in Unified Multimodal Models?

VGAU-Diag is introduced, a fine-grained evaluation framework for vision generation-assisted understanding that stratifies samples by difficulty, enables unified evaluation of multiple reasoning paradigms, and uses Oracle-Ass Reference Protocols.

Yubo Zhu, Zhehan Kan, Jing-Yi Yang et al. · 1 citation
Preprint Aug 2026

A Model-Internal Protocol for Assessing Multimodal Models as Integrated Systems

As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge. Current evaluation protocols predominantly treat generative and discriminative capabilities as separate tasks, leaving a gap in system-level evaluation for unified multimodal models (UMMs). In this work, we propose Self-Generative-Understanding (SGU), a novel, annotation-free evaluation framework that probes the integrated capabilities of unified models through a semantic closed-loop challenge. Without requiring new annotations, SGU leverages the dual understanding-and-generation abilities of UMMs by asking them to first perceive an image and produce a textual description, subsequently reconstruct a visual context based on that description, and finally perform reasoning over the self-generated output. This pipeline provides a zero-cost testbed that yields an integrated performance score specifically tailored for evaluating UMMs as unified systems. Extensive experiments show that even high-performing UMMs often struggle to reason over their own generated contexts, revealing limitations that are not captured by separate evaluations of understanding or generation alone. Our work provides a complementary holistic evaluation framework and offers a foundation for benchmarking the development of next-generation unified multimodal models.

Hao Zhang, Jiaxin Qi, Zhijiang Tang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.