Skip to content
Open access

Dual-branch cross-modal architecture with global-to-local feedback for radiographic image–text retrieval in chest X-rays

Sep 2026 · Scientific Reports · Vol 16 · 0 citations · 48 references

TL;DR

Through hierarchical gated co-attention, this approach dynamically aligns image and text representations, addressing the limitations of static fusion and providing a foundation for interpretable, real-time retrieval systems that can accelerate and improve clinical decision support.

Abstract

Radiology reports are vital for accurate diagnosis and treatment planning, yet their manual generation is time-consuming and dependent on radiologist expertise, leading to delays and inconsistent clinical decisions. Medical image–text retrieval offers a scalable solution by enabling the retrieval of relevant prior cases and reports. However, existing models often rely on static embeddings, which limits their ability to capture fine-grained, clinically meaningful cross-modal relationships. This limitation is particularly critical in chest X-ray interpretation, where nuanced textual descriptions must align with subtle visual cues. We introduce a novel dual-branch retrieval framework that distinguishes between shared semantics and complementary features through a Synergy branch and a Difference branch, respectively. These branches are stabilized through orthogonal regularization, ensuring minimal redundancy while keeping complementary diagnostic cues. A Global-to-Local Feedback mechanism guides fine-grained local attention using global context, enhancing interpretability and clinical relevance. Evaluated on two benchmark datasets, our model achieves state-of-the-art retrieval performance across both image\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\rightarrow$$\end{document}text and text\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\rightarrow$$\end{document}image tasks. Through hierarchical gated co-attention, our approach dynamically aligns image and text representations, addressing the limitations of static fusion and providing a foundation for interpretable, real-time retrieval systems that can accelerate and improve clinical decision support.

Read PDF

Similar papers

Sep 2026

Deep learning-based diagnostic report generation for low-resolution functional medical images via cross-modal visual and textual alignment

A unified framework for automatic report generation from SPECT bone scintigrams that integrates domain-adaptive representation learning, fine-grained image–text alignment, and anatomy-guided supervision is proposed, offering a valuable pathway to achieving trustworthy and intelligent diagnostic support within nuclear m...

Tao Song, Qiang Lin, Tong-Tong Li et al. · 0 citations
Sep 2026

Comorbidity-Aware Radiology Report Generation.

Chest X-ray report generation systems are valuable for assisting disease diagnosis and improving healthcare efficiency. However, existing methods still face two key challenges. First, multiple diseases often co-occur, leading to a combinatorial explosion of label combinations and sparse supervision for learning a gener...

Hong-Ze Zhu, Hong Liu, Ya-Wen Huang et al. · 0 citations
Preprint Aug 2026

SeVeR: Selective Visual Exposure and Retrieval for 3D Medical Image Question Answering

SeVeR is proposed, a selective visual exposure framework that compresses dense volumes into modality-wise prototypes and retrieves complementary multi-level evidence with change-aware gated attention during decoding, trained with a marginal-utility self-consistency objective that suppresses unhelpful retrieval.

Yao-Jun Hu, Danyang Tu, Yang Liu et al. · 0 citations
Preprint Sep 2026

Integrating Local Detail and Global Context: A Dual-Input Multi-Task Learning Framework for Bone Tumor Diagnosis

Primary bone tumors are rare but clinically aggressive neoplasms whose diagnosis from radiographs is challenged by heterogeneous morphology, subtle lesion margins, and overlapping bone structures. To address the limitations of existing single-view models, we present a dual-input, multi-task learning framework that, to...

S. M. Nasif Uddin, Rusab Sarmun, Muhammad E. H. Chowdhury et al. · 0 citations
Conference Open access Aug 2026

DisenTextSeg: Disentangling Textual Information from Clinical Reports for Multi-Modal Medical Image Segmentation

Accurate segmentation of multi-modal magnetic resonance imaging (MRI) is central to clinical diagnosis, treatment planning, and prognostic evaluation of tumors. Existing methods predominantly rely on imaging data alone, neglecting the rich semantic information embedded in clinical reports. Although recent vision-langua...

Ding-Jie Suo, Han-Di Zhu, Tian-Tian Liu et al. · 0 citations
Preprint Sep 2026

PLRS-IC: A Dual-Calibration Framework for Chest X-Ray Vision-Language Alignment

Fine-grained vision-language alignment in chest radiography enables zero-shot classification, grounding, and segmentation without task-specific annotations. However, this alignment is fundamentally hindered by two intertwined sources of ambiguity: projection-induced visual mismatch and patient-agnostic semantic overlap...

Qi-Xing Zhao, Jin-Peng Li · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.