Skip to content

From Compound Figures to Medical Multi-image Reasoning: Scaling Multimodal Large Language Models with Biomedical Literature

Zhen Chen Yihang Fu Rong Zhou Serina Applebaum Min Kyu Kim Aidan Gilson Morten Lee Salahudeen Mirza Gabriel Madera Mauro Giuffre Yuanting Pan Roy Jiang Hyunjae Kim Hua Xu Qingyu Chen
Oct 2026
Artificial Intelligence Natural Language Processing Computer Vision

Abstract

Multimodal large language models (MLLMs) are increasingly capable in medical imaging, yet most focus on single-image settings. Clinical interpretation often requires integrating evidence across multiple images, such as different modalities, views, or time points. However, large-scale medical multi-image data and training strategies for such reasoning remain limited. We construct PMC-MI, a large-scale resource comprising 234,956 instruction instances derived from biomedical compound figures, with 10,555 multi-subimage instances structured for reinforcement learning and assessed by medical reviewers. We also introduce PMC-MI-Bench, a manually reviewed benchmark separated at the source-article level. We further propose a three-stage training framework, instantiated as M3LLM, combining supervised fine-tuning on multi-image instructions, selection-aware reinforcement learning for question-conditioned visual-evidence selection, and supervised consolidation on a broader instruction mixture. Evaluating M3LLM on PMC-MI-Bench, two public medical benchmarks, longitudinal chest radiographs, and a retrospective dermatology cohort reveals it achieves the highest performance among representative MLLMs. On PMC-MI-Bench, M3LLM scored 81.6, 84.6, and 85.2 on single-subimage, multi-subimage, and relative-position tasks, respectively, outperforming the strongest baselines (75.6, 79.1, and 77.6). It also set new records on the provenance-screened OmniMedVQA (88.6%) and MMMU-Med (68.0%), alongside top performance in two clinical settings after task-specific adaptation. These findings support using literature-derived multi-image supervision and targeted post-training for medical multi-image MLLMs. Data and code are publicly available.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Diffusion models as plug-and-play priors

The possibility of inferring high-dimensional data inference in a model that consists of a prior and an auxiliary differentiable constraint given some additional information is considered, thereby allowing a range of potential applications in adapting models to new domains and tasks.

Alexandros Graikos, Esmeralda S. Whitammer, N. Jojic et al. · 316 citations · ⚡15

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.