Temporal vision-language models (TVLMs) offer a reusable, prompt-based interface for surgical video understanding, yet, their robustness under clinically realistic acquisition artifacts in endoscopy remains insufficiently characterized. In practice, degradations such as defocus, haze, motion blur, noise, cautery smoke,...
Darakshan Rashid, Raza Imam, Ufaq Khan et al.· 0 citations
Radiological artificial intelligence has advanced rapidly, yet most systems remain narrowly task-specific, data-intensive, and fragile under domain shift. Foundation models promise more transferable and data-efficient solutions, but existing approaches are limited in scale, evaluated narrowly, and often assume that a s...
C. U. Harsy, Tassilo Wald, Karol Gotkowski et al.· 0 citations
A domain generalization framework that utilizes a foreground-only histogram matching protocol to resolve the domain shift issue arising from disparate clinical sources is proposed, significantly outperforming prominent domain generalization paradigms, including MixStyle and Discrete-Fourier-Transform-based frameworks.
Hong-Yi Pan, Gorkem Durak, H. Aktas et al.· 0 citations
Lung ultrasound (LUS) is a bedside tool for assessing pulmonary edema in patients at risk due to heart failure or impaired kidney function. However, automated LUS analysis remains challenging because of speckle noise, imaging artifacts, and operator-dependent acquisition variability. In this work, we present a deep lea...
Alya Almsouti, Lotfi Abdelkrim Mecharbat, Noha Aboukhater et al.· arXiv.org· 0 citations
Results show that forgetting depends not only on how much the model changes, but also on which parts of the model are allowed to change, which means that forgetting still increases as more blocks are trained and remains severe when the full backbone is updated.
Amal Saqib, Tausifa Jan Saleem, N. Saeed et al.· 0 citations
Medical vision-language models (MVLMs) promise broad zero-shot generalization, yet their reliability collapses when confronted with unseen modalities and domains, precisely where clinical robustness matters most. To address this gap, we revisit test-time modality generalization from the perspective of Mixture-of-Expert...
Raza Imam, Darakshan Rashid, Yutong Xie et al.· 0 citations
A systematic evaluation of CL for MedVQA across diverse clinical objectives, including classification, multi-label classification, detection, cell counting, and report generation suggests that existing CL methods struggle to maintain stability-plasticity balance when tasks with different objectives and supervision form...
Mai A. Shaaban, Tausifa Jan Saleem, Alaa Mohamed et al.· arXiv.org· 0 citations
A novel multimodal RAG framework tailored for MedVQA is proposed, which leverages multimodal data, including medical images, reports, and generated captions, to provide more accurate clinical answers, and introduces a training paradigm that uses captions as auxiliary supervision, enhancing cross-modal alignment via con...
Mai A. Shaaban, M. Zarei, Adnan Khan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.