Sep 2026· AIiH· pp. 122-135· 0 citations· 31 references
Computer Science
TL;DR
A unified, patch-based framework that processes raw multimodal imaging through training-free registration, automated localization, and mask-filtered patch extraction that culminates in a hierarchical strategy that aggregates patch-level insights into subject-level diagnostics.
Abstract
Adapting natural-image foundation models like DINOv3 to multi-modal medical imaging is challenging due to the significant domain gap between natural color images and multi-channel medical scans. We present a unified, patch-based framework that processes raw multimodal imaging through training-free registration, automated localization, and mask-filtered patch extraction. This architecture culminates in a hierarchical strategy that aggregates patch-level insights into subject-level diagnostics. Using liver fibrosis staging as a case study, we evaluate four patch-level feature representations: handcrafted Radiomics features, learned ResNet features, pre-trained foundation model SAM-Med2D features, and frozen DINOv3 features. To ensure a controlled comparison, all models utilize the same lightweight MLP head and are evaluated across both rigid and deformable registration settings. Our training protocol focuses on mild fibrosis (S1) and cirrhosis (S4) classes only, enabling a single classifier to address both substantial fibrosis detection and cirrhosis staging. Evaluated via 10 random train (90%)/ test (10%) splits on 360 subjects from the CARE 2025 Liver Track 4 cohort, our DINOv3-based framework significantly outperforms all baselines, achieving the best classification accuracy of 78.4% for S1 and 75.8% for S4.
Accurate staging of liver fibrosis is critical for guiding antifibrotic therapy and predicting disease prognosis. Liver biopsy, the current reference standard, is invasive, costly, and subject to significant sampling error. Although T2-weighted MRI and B-mode ultrasound offer established noninvasive alternatives, exist...
Raufu Umme Rhidi, Ferdaus Anam Jibon, Farhan Shahriar et al.· Journal of Umm Al-Qura Unive...· 0 citations
Comparisons across all datasets confirm that the proposed framework exhibits strong robustness and generalization capability when processing multiple medical imaging modalities, thereby providing reliable technical support for computer-aided medical diagnosis systems.
Ya-Chao Si, Yi Zhang, Ming-Zhan Zhao· Scientific Reports· 0 citations
Medical image classification is fundamental to computer-aided diagnosis. Limited labeled samples, subtle inter-class differences, and heterogeneous lesion morphology make it difficult for a single representation to capture all relevant cues. This study investigates MedFuse, a simple dual-stream framework that combi...
Ya-Jing Ren, Hai Ling, Zheng Gu et al.· Frontiers in Medicine· 0 citations
MRD-UNet provides a practical balance between segmentation accuracy and computational efficiency and outperforms baseline CNNs and performs comparably to heavier transformer-based models while using significantly fewer parameters.
Musa Doğan, I. Ozkan· BMC Medical Imaging· 0 citations
GLNet adopts a dual-branch encoder that combines a CNN-based Local Detail Perception Branch with a Mamba-based Global Context Modeling Branch, enabling the joint extraction of fine-grained local features and long-range semantic representations.
This work proposes a cross-domain universal medical IQA method termed CDMIQA, which integrates efficient feature extractors with Hierarchical Perceptual Encoding Modules to capture and refine multi-level perceptual features while mitigating interference from noise and artifacts.
Lei-Lei Huang, Yue Sun, Ming-Xiang Wu et al.· Proceedings of the Thirty-Fi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.