Skip to content

DINO-Med: A Unified Patch-Based Adaptation Framework for Multi-Modal Medical Image Analysis Applied to Liver Fibrosis Staging

Sep 2026 · AIiH · pp. 122-135 · 0 citations · 31 references
Computer Science

TL;DR

A unified, patch-based framework that processes raw multimodal imaging through training-free registration, automated localization, and mask-filtered patch extraction that culminates in a hierarchical strategy that aggregates patch-level insights into subject-level diagnostics.

Abstract

Adapting natural-image foundation models like DINOv3 to multi-modal medical imaging is challenging due to the significant domain gap between natural color images and multi-channel medical scans. We present a unified, patch-based framework that processes raw multimodal imaging through training-free registration, automated localization, and mask-filtered patch extraction. This architecture culminates in a hierarchical strategy that aggregates patch-level insights into subject-level diagnostics. Using liver fibrosis staging as a case study, we evaluate four patch-level feature representations: handcrafted Radiomics features, learned ResNet features, pre-trained foundation model SAM-Med2D features, and frozen DINOv3 features. To ensure a controlled comparison, all models utilize the same lightweight MLP head and are evaluated across both rigid and deformable registration settings. Our training protocol focuses on mild fibrosis (S1) and cirrhosis (S4) classes only, enabling a single classifier to address both substantial fibrosis detection and cirrhosis staging. Evaluated via 10 random train (90%)/ test (10%) splits on 360 subjects from the CARE 2025 Liver Track 4 cohort, our DINOv3-based framework significantly outperforms all baselines, achieving the best classification accuracy of 78.4% for S1 and 75.8% for S4.

View source

Similar papers

Open access Sep 2026

HybridDenseViT: a unified explainable hybrid deep learning framework for multi-stage liver fibrosis staging across T2-weighted MRI and B-mode ultrasound

Accurate staging of liver fibrosis is critical for guiding antifibrotic therapy and predicting disease prognosis. Liver biopsy, the current reference standard, is invasive, costly, and subject to significant sampling error. Although T2-weighted MRI and B-mode ultrasound offer established noninvasive alternatives, exist...

Raufu Umme Rhidi, Ferdaus Anam Jibon, Farhan Shahriar et al. · 0 citations
Open access Aug 2026

A medical image classification algorithm based on a hierarchical and complementary attention-enhanced Swin Transformer model

Comparisons across all datasets confirm that the proposed framework exhibits strong robustness and generalization capability when processing multiple medical imaging modalities, thereby providing reliable technical support for computer-aided medical diagnosis systems.

Ya-Chao Si, Yi Zhang, Ming-Zhan Zhao · 0 citations
Open access Sep 2026

MedFuse: dual-stream fusion of convolutional and vision transformer-based features for enhanced medical image classification

Medical image classification is fundamental to computer-aided diagnosis. Limited labeled samples, subtle inter-class differences, and heterogeneous lesion morphology make it difficult for a single representation to capture all relevant cues. This study investigates MedFuse, a simple dual-stream framework that combi...

Ya-Jing Ren, Hai Ling, Zheng Gu et al. · 0 citations
Aug 2026

GLNet: global-to-local aware hybrid framework for medical image segmentation

GLNet adopts a dual-branch encoder that combines a CNN-based Local Detail Perception Branch with a Mamba-based Global Context Modeling Branch, enabling the joint extraction of fine-grained local features and long-range semantic representations.

Dengdi Sun, Longlong Liu, Xiao-Wei Zhao et al. · 0 citations
Conference Open access Sep 2026

CDMIQA: A Cross-Domain Perceptual Method and Benchmark Dataset for Medical Image Quality Assessment

This work proposes a cross-domain universal medical IQA method termed CDMIQA, which integrates efficient feature extractors with Hierarchical Perceptual Encoding Modules to capture and refine multi-level perceptual features while mitigating interference from noise and artifacts.

Lei-Lei Huang, Yue Sun, Ming-Xiang Wu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.