Skip to content
Review

A Survey of Multimodal Models on Language and Vision: A Unified Modeling Perspective

· 2 citations · 265 references

TL;DR

This survey investigates the current research landscape of multimodality modeling from three perspectives: the first group of multimodal models adopts a heterogeneous architecture to bridge different modality data, the second leverages LLM for multimodality modeling via a unified language modeling objective, and the third represents multimodal data entirely within a single visual representation.

View source

Similar papers

Review Open access Jul 2026

Multimodal Video Understanding: A Capability-Based Survey of Alignment, Expression, and Reasoning

A structured, comprehensive survey of the latest MVU progress is presented, establishing a novel three-tier taxonomy that categorizes existing studies into cross-modal alignment, multi-granularity semantic expression and multimodal reasoning.

Rongyong Zhao, Da Pu, Cuiling Li et al. · 0 citations
Preprint Aug 2026

A Model-Internal Protocol for Assessing Multimodal Models as Integrated Systems

As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge. Current evaluation protocols predominantly treat generative and discriminative capabilities as separate tasks, leaving a gap in system-level evaluation for unified multimodal models (UMMs). In this work, we propose Self-Generative-Understanding (SGU), a novel, annotation-free evaluation framework that probes the integrated capabilities of unified models through a semantic closed-loop challenge. Without requiring new annotations, SGU leverages the dual understanding-and-generation abilities of UMMs by asking them to first perceive an image and produce a textual description, subsequently reconstruct a visual context based on that description, and finally perform reasoning over the self-generated output. This pipeline provides a zero-cost testbed that yields an integrated performance score specifically tailored for evaluating UMMs as unified systems. Extensive experiments show that even high-performing UMMs often struggle to reason over their own generated contexts, revealing limitations that are not captured by separate evaluations of understanding or generation alone. Our work provides a complementary holistic evaluation framework and offers a foundation for benchmarking the development of next-generation unified multimodal models.

Hao Zhang, Jiaxin Qi, Zhijiang Tang et al. · 0 citations
Preprint Jul 2026

Vision as Unified Multimodal Generation

Experiments show that a single unified model can match leading task-specialized systems across structured visual understanding, dense geometric prediction, segmentation, and multi-view visual geometry.

Xiaoyang Han, Jianhua Li, Kewang Deng et al. · 1 citation
Review Open access Jul 2026

A Review of Multimodal Large Language Models: Fusion Mechanisms and Capability Evolution

This paper reviews the principal technical paradigms of multimodal fusion, including early fusion, intermediate fusion, late fusion, and hybrid fusion, and compares the structural characteristics and applicable scenarios of different fusion approaches and explores the development of MLLMs from the perspectives of vision-language understanding, multimodal content generation, multimodal interaction and agent-oriented tasks, as well as domain-specific applications.

Zibo Xu, Shuqiang Gao · 0 citations
Jul 2026

Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering

It is shown that steering vectors learned from the understanding branch can transfer to generation, enabling controllable image synthesis and improved semantic faithfulness, and establish cross-branch steering as a practical tool for probing multimodal representations.

Yu Wang, Sharon Li · 0 citations
Jul 2026

MIRROR: Learning from the Other View for Multi-Modal Reasoning

Modality-Informed Reciprocal Reasoning Optimization (MIRROR), a reinforcement learning approach for improving multimodal reasoning via self supervision, is developed and improves over standard RL and yields more accurate and consistent behavior across modalities.

Wen Ye, Yuxiao Qu, Aviral Kumar et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.