MedVCoT: Bridging the Modality Gap in Medical VQA Through Latent Visual Reasoning
This work proposes MedVCoT, which incorporates latent visual reasoning into the medical visual question answering (VQA) domain, and utilizes the specialized expertise of MedSAM to train a large vision-language model so that it can autonomously generate consistent and continuous latent visual tokens within Visual Chain-of-Thought.