This survey reviews automated multimodal machine learning (MMAutoML), distinguishing core systems that automate both representation and fusion decisions from near-core systems, partial multimodal AutoML tools, and adjacent multimodal ML infrastructure.
Abstract
Automated machine learning (AutoML) reduces the cost of developing high-performing models by automating pipeline design, model selection, and hyperparameter optimization. As data sets increasingly combine heterogeneous sources (e.g., text, images, audio, and structured/tabular records), AutoML is extending to multimodal settings, where performance depends on representation learning and cross-modal fusion as well as model search. This survey reviews automated multimodal machine learning (MMAutoML), distinguishing core systems that automate both representation and fusion decisions from near-core systems, partial multimodal AutoML tools, and adjacent multimodal ML infrastructure. We organize multimodal learning around early, late, and hybrid fusion paradigms and discuss trade-offs in robustness, interpretability, and deployment. We then curate and compare representative open-source and commercial MMAutoML frameworks, summarize supported modalities, and characterize the ecosystem via time-evolution and similarity-based groupings. Finally, we overview applications in healthcare, autonomous systems, finance, and e-commerce, and highlight open challenges in modality alignment, missing or degraded inputs, efficiency and scalability, reproducibility, and governance (privacy, bias, and monitoring). We conclude with practical guidance and research directions toward reliable, end-to-end MMAutoML.
This review aims to offer a conceptual framework and a solid reference for building intelligent multimodal material selection systems and presents a forward-looking research agenda covering self-supervised learning, knowledge-enhanced models, and interpretable human–AI collaboration.
Yuwei Zhang, C. Tan, Bei-Chen Wang et al.· IEEE Access· 0 citations
Large language models (LLMs) are being applied across diverse fields due to their capability to derive various insights from complex data. In biotechnology, where complex multimodal data including images is rapidly expanding, LLMs offer powerful capabilities for data analysis. However, unlocking the full potential of these models depends critically on prompt engineering, which is often a labor-intensive process that requires specialized expertise and lacks reproducibility. To address these challenges, we developed novel automatic prompt engineering (APE) approaches tailored for multimodal tasks in the bioscience and bioengineering domains. This study introduces two types of approaches: the Batch APE method, an efficient method for optimizing prompts for powerful black-box models, and a fine-tuning method with supervised fine-tuning (SFT) and direct preference optimization (DPO) on a local vision-language model (VLM). These methods were systematically evaluated across four diverse scientific datasets of microscopic images of protein crystals, human cell images, molecular structure images, and medical radiography images of Chest X-ray. Experimental results demonstrated that Gemini with Batch APE and SFT-based local LLM alignment generally outperformed baseline APE techniques, though DPO alone showed inconsistent results across datasets. Furthermore, qualitative analysis of the generated prompts revealed key characteristics of prompts that enhance image classification performance in multimodal LLMs. These findings highlight the potential of advanced APE to improve the utility of both local and black-box multimodal models for specialized scientific applications, particularly in domains where fine-tuning was restricted or infeasible.
Keisuke Mizutani, Rintaro Yashiro, Kento Tokuyama· Journal of Bioscience and Bi...· 0 citations
Recent advances in large language models (LLMs) and vision transformers have enabled multimodal systems that integrate clinical text with medical imaging for diagnostic decision-making. While these systems show promising results on benchmark datasets in well-resourced research settings, their applicability in low-resource healthcare environments where diagnostic disparities are most severe remains limited and poorly understood. This mini review synthesizes key developments in LLM–vision fusion architectures from 2018 to 2026, with a focus on radiology-oriented visual question answering (VQA) and report generation systems viewed from a deployment perspective. Rather than comprehensively cataloguing multimodal medical AI, we synthesize the evolution of LLM–vision fusion architectures and discuss complementary deployment-enabling strategies, including parameter-efficient adaptation, post-training quantization, federated learning, and multilingual support, where they directly improve the feasibility of radiology AI in resource-constrained healthcare settings. Rather than focusing solely on performance benchmarks, we examine these approaches through a deployment-oriented lens, highlighting trade-offs between representational capacity, computational efficiency, interpretability, and memory footprint. We argue that current progress remains substantially shaped by model scaling and benchmark optimization, which often do not address the memory, connectivity, and annotation constraints of low-resource healthcare systems. While cross-modal transformer architectures provide strong representational alignment, their computational demands and reliance on large curated datasets limit real-world deployment. In contrast, emerging directions including parameter-efficient fine-tuning, post-training quantization, federated learning, and modular agent-based systems offer more tractable pathways toward clinical integration under hardware and data constraints. To bridge the gap between benchmark performance and clinical utility, we identify concrete challenges in data scarcity, multilingual coverage, and calibration, and propose a shift toward lightweight, interpretable, and hardware-aware multimodal AI. This perspective highlights the need to move beyond scaling-centric design toward models that can run on 4–8 GB VRAM, operate offline, and generalize across languages and imaging equipment.
Kahakashan Ashraf, Md.Hamid Hosen, N. Farah et al.· Frontiers in Digital Health· 0 citations
This Review examines volumetric foundation models, language alignment and compression strategies, and agentic systems that extend MLLMs through planning, tools, memory, and workflow interaction, and introduces a Claim-Design-Validation framework to assess whether technical, workflow, and clinical claims are matched by appropriate design and validation.
Zanting Ye, Shengyuan Liu, Xin Liu et al.· 0 citations
PURPOSE
Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias.
METHODS
Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative.
RESULTS
Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain.
CONCLUSION
Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.
Md. Mazharul Islam, Abrar Mohammed Tanzim Alam, Md Sadikur Rahman Rony et al.· International Journal of Med...· 0 citations
Abstract. Functional near-infrared spectroscopy (fNIRS) and diffuse optical 1 tomography (DOT) are rapidly evolving toward wearable, multimodal, data-driven, and artificial-intelligence-supported neuroimaging in the everyday world. However, current analytical tools are fragmented across platforms, limiting reproducibility, interoperability, and integration with modern machine learning (ML) workflows. Cedalion is a Python-based open-source framework designed to unify advanced model-based and data-driven analysis of multimodal fNIRS and DOT data within a reproducible, extensible, and community-driven environment. Cedalion integrates forward modeling, photogrammetric optode coregistration, signal processing, general linear model (GLM) analysis, DOT image reconstruction, and ML-based data-driven methods within a single standardized architecture based on the Python ecosystem. It adheres to SNIRF and BIDS standards, supports cloud-executable Jupyter notebooks, and provides containerized workflows for scalable, fully reproducible analysis pipelines that can be provided alongside original research publications. Cedalion connects established optical-neuroimaging pipelines with ML frameworks such as scikit-learn and PyTorch, enabling seamless multimodal fusion with electroencephalography (EEG), magnetoencephalography (MEG), and physiological data. It implements validated algorithms for signal quality assessment, motion correction, GLM modeling, and DOT reconstruction, complemented by modules for simulation, data augmentation, and multimodal physiology analysis. Automated documentation links each method to its source publication, and continuous-integration testing ensures robustness. This tutorial paper provides seven fully executable notebooks that demonstrate core features. Cedalion offers an open, transparent, and community-extensible foundation that supports reproducible, scalable, and cloud- and ML-ready fNIRS/ DOT workflows for laboratory-based and real-world neuroimaging.
E. Middell, Laura B. Carlton, Shakiba Moradi et al.· Neurophotonics· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.