Through hierarchical gated co-attention, this approach dynamically aligns image and text representations, addressing the limitations of static fusion and providing a foundation for interpretable, real-time retrieval systems that can accelerate and improve clinical decision support.
Abstract
Radiology reports are vital for accurate diagnosis and treatment planning, yet their manual generation is time-consuming and dependent on radiologist expertise, leading to delays and inconsistent clinical decisions. Medical image–text retrieval offers a scalable solution by enabling the retrieval of relevant prior cases and reports. However, existing models often rely on static embeddings, which limits their ability to capture fine-grained, clinically meaningful cross-modal relationships. This limitation is particularly critical in chest X-ray interpretation, where nuanced textual descriptions must align with subtle visual cues. We introduce a novel dual-branch retrieval framework that distinguishes between shared semantics and complementary features through a Synergy branch and a Difference branch, respectively. These branches are stabilized through orthogonal regularization, ensuring minimal redundancy while keeping complementary diagnostic cues. A Global-to-Local Feedback mechanism guides fine-grained local attention using global context, enhancing interpretability and clinical relevance. Evaluated on two benchmark datasets, our model achieves state-of-the-art retrieval performance across both image\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\rightarrow$$\end{document}text and text\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\rightarrow$$\end{document}image tasks. Through hierarchical gated co-attention, our approach dynamically aligns image and text representations, addressing the limitations of static fusion and providing a foundation for interpretable, real-time retrieval systems that can accelerate and improve clinical decision support.
A unified framework for automatic report generation from SPECT bone scintigrams that integrates domain-adaptive representation learning, fine-grained image–text alignment, and anatomy-guided supervision is proposed, offering a valuable pathway to achieving trustworthy and intelligent diagnostic support within nuclear m...
Tao Song, Qiang Lin, Tong-Tong Li et al.· Applied intelligence (Boston...· 0 citations
Chest X-ray report generation systems are valuable for assisting disease diagnosis and improving healthcare efficiency. However, existing methods still face two key challenges. First, multiple diseases often co-occur, leading to a combinatorial explosion of label combinations and sparse supervision for learning a gener...
Hong-Ze Zhu, Hong Liu, Ya-Wen Huang et al.· IEEE Transactions on Medical...· 0 citations
SeVeR is proposed, a selective visual exposure framework that compresses dense volumes into modality-wise prototypes and retrieves complementary multi-level evidence with change-aware gated attention during decoding, trained with a marginal-utility self-consistency objective that suppresses unhelpful retrieval.
Yao-Jun Hu, Danyang Tu, Yang Liu et al.· 0 citations
Primary bone tumors are rare but clinically aggressive neoplasms whose diagnosis from radiographs is challenged by heterogeneous morphology, subtle lesion margins, and overlapping bone structures. To address the limitations of existing single-view models, we present a dual-input, multi-task learning framework that, to...
S. M. Nasif Uddin, Rusab Sarmun, Muhammad E. H. Chowdhury et al.· 0 citations
Accurate segmentation of multi-modal magnetic resonance imaging (MRI) is central to clinical diagnosis, treatment planning, and prognostic evaluation of tumors. Existing methods predominantly rely on imaging data alone, neglecting the rich semantic information embedded in clinical reports. Although recent vision-langua...
Ding-Jie Suo, Han-Di Zhu, Tian-Tian Liu et al.· Journal of Physics, Conferen...· 0 citations
Fine-grained vision-language alignment in chest radiography enables zero-shot classification, grounding, and segmentation without task-specific annotations. However, this alignment is fundamentally hindered by two intertwined sources of ambiguity: projection-induced visual mismatch and patient-agnostic semantic overlap...
Qi-Xing Zhao, Jin-Peng Li· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.