Sep 2026· IEEE Transactions on Image Processing· Vol 35, pp. 9744-9756· 0 citations· 92 references
Medicine
Abstract
Medical foundation models (FMs) typically rely on contrastive vision-language pretraining over large-scale datasets, yet such datasets often exhibit substantial heterogeneity in data quality and demand extensive computational resources. Recent studies suggest that data quality matters more than data quantity, but how to identify a compact subset of pretraining data that best supports generalization remains unclear. In this paper, we propose an active Medical data Curation framework for efficient vision-language pretraining (named MedCure), which jointly prioritizes two key properties of training data: difficulty and diversity. Specifically, during pretraining, MedCure establishes undirected data graphs to characterize sample importance, dynamically updating the difficulty score of each sample by transferring and aggregating information from its neighbors within the graph. These refined scores are then leveraged by a backward difficulty transfer mechanism to adaptively curate subsets that comprehensively capture both difficult and heterogeneous regions of the data space. Extensive experiments on a chest X-ray database of over 1.2M image-text pairs illustrate that MedCure achieves comparable results to full-data training using only 25% of the pretraining data and 32.7% of the full training cost, as evaluated on a variety of zero-shot generalization and downstream-adapted tasks (e.g., disease classification, cross-modal retrieval, lesion and anatomy segmentation, and radiology report generation), highlighting its effectiveness for data-efficient medical FMs. Our code and pretrained models will be available at https://github.com/sendyma/MedCure
Recent medical multimodal models have benefited from larger corpora, broader modality coverage, and stronger reasoning-oriented training, yet effective data design across continued pretraining (CPT) and post-training remains challenging. Medical sources vary substantially in structure, granularity, and information dens...
Guang-Hao Zhu, Ze-Yu Liu, Zhitian Hou et al.· 0 citations
Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing considerable potential in medicine. However, their application in medical settings remains limited by the scarcity of visual question answering (VQA) datasets that capture clinical reasoning and explicit image-text alignm...
Ling-Xuan Hou, Yu-Hua Xie, Yue Hu et al.· 0 citations
This work proposes MedRecord-CLIP, a knowledge-enhanced foundation model featuring a diagnosis-guided cross-attention mechanism to adaptively extract and fuse salient patient history with diagnostic representations that highlights the critical value of integrating personalized clinical context to enhance the generaliza...
Lei Shi, Wenbin Zhai, Lei Yu et al.· Health Information Science a...· 0 citations
Satellite imagery is proposed as a novel pretraining domain for MedVFM development and benchmarking, motivated by its closer visual alignment with medical data and its freedom from the privacy constraints that limit medical datasets.
Lovre Antonio Budimir, Ming Gong, Alyssa Foong Quinney et al.· 0 citations
The rapid expansion of large-scale medical datasets and computational resources has driven significant progress in medical foundation models. Given the inherent heterogeneity of medical imaging modalities, current research mainly follows two paths: specialized models optimized for specific modalities, and generalist mo...
Chu Zhang, Hao-Yu Jiang, Hong-Yuan Zhang et al.· 0 citations
Medical time series (MedTS) underpin many clinical classification tasks, yet existing methods usually represent them only as numerical sequences and underuse the morphology that is explicit in waveform inspection. To bridge this gap, we introduce Vision-Informed Retrieval (ViRe), which uses a frozen VLM-derived wavefor...