Skip to content
Review Open access

Multimodal Document Classification Across Domains and Methods: A Systematic Review

Aug 2026 · Engineering, Technology & Applied Science Research · Vol 16, pp. 37853-37861 · 0 citations · 49 references

TL;DR

This systematic review synthesizes existing literature on MDC, focusing on methods, architectures, cross-domain adaptability, evaluation metrics, and current challenges, and provides a comprehensive resource for researchers and practitioners, aiming to consolidate knowledge and guide future advancements in MDC.

Abstract

Multimodal Document Classification (MDC) is a significant area of research, enabling the integration of textual, visual, and structural features to enhance document understanding across diverse application domains. While substantial progress has been made in developing methods that leverage deep learning, graph-based representations, and cross-modal fusion techniques, the field remains fragmented accross datasets, evaluation protocols, and methodological frameworks. This systematic review synthesizes existing literature on MDC, focusing on methods, architectures, cross-domain adaptability, evaluation metrics, and current challenges. Moreover, following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, the current work surveyed peer-reviewed studies published between 2021 and 2025, identifying trends in feature extraction, fusion strategies, and benchmark datasets. The analysis conducted highlights three key insights: i) the shift from handcrafted features to transformer-based and multimodal pre-trained models; ii) the growing importance of domain-specific adaptations in legal, healthcare, and scientific documents; and iii) persistent challenges related to scalability, interpretability, and generalizability across domains. Overall, this review provides a comprehensive resource for researchers and practitioners, aiming to consolidate knowledge and guide future advancements in MDC.

Read PDF

Similar papers

#computer vision Preprint Sep 2026

Evolution of Multimodal Question Answering: From Modality-Adaptive Extraction to Unified Language Representation

The rapid growth of multimodal data has intensified the need for question answering (QA) systems capable of reasoning across heterogeneous sources such as text, tables, and images. In this paper, we present a comprehensive methodological comparison of three influential frameworks, namely Multimodal Adaptive Extraction (MAE), Solar, and UniMMQA, tracing the evolution of multimodal question answering from modality-adaptive pipelines to fully unified architectures. We examine how each approach models cross-modal interactions, transforms heterogeneous inputs, and performs reasoning, highlighting key design differences in modality representation, reasoning, and answer generation. Our analysis demonstrates a clear shift from explicit modality-specific processing toward unified text-centric formulations enabled by pre-trained language models (PLMs). Empirical comparisons across benchmark datasets show that this transition leads to substantial improvements in both Exact Match (EM) and F1-Scores, with UniMMQA achieving the most consistent and scalable performance. Despite these advances, we identify persistent challenges, including information loss during modality transformation, error propagation in multi-stage pipelines, and limitations in capturing fine-grained cross-modal dependencies. Overall, this study provides a deeper understanding of current design trends and offers insights into the future direction of unified multimodal reasoning systems.

Abdullah Al Shafi · 0 citations
Review Open access Jul 2026

Multimodal Video Understanding: A Capability-Based Survey of Alignment, Expression, and Reasoning

A structured, comprehensive survey of the latest MVU progress is presented, establishing a novel three-tier taxonomy that categorizes existing studies into cross-modal alignment, multi-granularity semantic expression and multimodal reasoning.

Rongyong Zhao, Da Pu, Cuiling Li et al. · 0 citations
Open access 2026

Enhancing Biomedical Multi-Label Text Classification via Topic-Based Text Representation

: Biomedical texts naturally contain multiple biological and medical concepts within a document, resulting in a semantically rich and complex structure. Consequently, multi-label text classification (MLTC) has become a suitable framework for comprehensively modeling biomedical texts, including clinical reports, laboratory records, and scientific abstracts. However, relying solely on contextual language representations may be insufficient to explicitly reflect the broader scientific focus and conceptual orientation of a document. In this study, the MLTC problem in the biomedical domain is investigated using the Hallmarks of Cancer (HoC) dataset. Topic probability distributions obtained from CombinedTM are incorporated as an additional representational signal into pre-trained language models (PLMs), enabling the joint exploitation of contextual semantic information and document-level topical characteristics. Both general-purpose and biomedical language models were fine-tuned and evaluated within this unified framework. The experimental results demonstrate that incorporating topic information substantially improves the domain robustness of the models. In particular, the Macro-F1 score of the general-purpose BERT model increases from 0.7604 to 0.8458, indicating that it becomes competitive with biomedical language models on tasks that require biomedical domain knowledge. Similarly, biomedical models such as BioMed-RoBERTa also benefit from topic modeling, with Macro-F1 improving from 0.8665 to 0.8819; meanwhile, the highest overall performance is achieved by the PubMedBERT + CombinedTM configuration, with a Macro-F1 score of 0.8900. These findings indicate that integrating topic modeling– based thematic information with PLMs provides an effective, generalizable, and practical solution for biomedical MLTC tasks. Moreover, enriching general-purpose language models with topic information offers a promising alternative that reduces reliance on costly, data-intensive domain adaptation.

Unknown authors · 0 citations
Jul 2026

MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware

Multimodal long documents are core carriers of professional knowledge, where critical evidence is sparsely distributed across paragraphs and modalities. This easily causes key information omission and cross-modal hallucinations in summarization by multimodal LLMs. These issues stem from attention drift in long-range dependency modeling and gaps in inter-modal alignment. To address this, we introduce MMLDSum-Bench, a high-quality benchmark for multimodal long-document summarization, covering multiple domains, context-length scales, and visual-textual modality distributions. We further propose MMLDSum-LLM, a reproducible two-stage training framework that combines supervised fine-tuning with visual-alignment weighted loss and keyword-aware weighted loss, followed by GRPO with a multi-objective reward (keyword coverage, image-text alignment, ROUGE, and length control). Extensive experiments on MMLDSum-Bench, comparing against leading closed-source and open-source multimodal models under a unified evaluation protocol - including LLM-as-a-judge scoring, atomic-claim precision/recall, image-text alignment (ITA), and ROUGE - demonstrate that our approach significantly improves key-information coverage and cross-modal consistency.

Xianpeng Zhang, Jiahuai Yang, Dong Chen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation

The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pathology report generation remains limited by the scarcity of large-scale WSI--report datasets and the complexity of mapping spatially distributed visual patterns to structured clinical text. To address this, we introduce a clinically curated Pan-Asia WSI--report dataset of approximately 10,500 pairs from five institutions and establish the REG 2025 benchmark through a MICCAI challenge for systematic evaluation of multimodal models. We analyze submitted methods spanning pretrained VLMs, multiple-instance learning frameworks, hierarchical expert models, retrieval-augmented generation, and cross-modal Transformers. Rather than indicating that VLM use alone was sufficient for superior performance, the results suggest that top-performing methods benefited from structured report representations, hierarchical diagnostic decomposition, and effective multimodal grounding. We identify key limitations, including instability in quantitative attribute estimation (e.g., numeric hallucination) and a tendency toward diagnostic overspecification, with some errors resembling known diagnostic pitfalls in routine pathology. These findings establish REG 2025 as a benchmark for evaluating WSI-based structured report generation and vision-language understanding in computational pathology, providing insights for the design of clinically grounded multimodal pathology models.

Yu-Mi Lee, Harim Oh, Hyo-yun Kim et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.