Skip to content

Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models

Jul 2026 · arXiv.org · Vol abs/2607.13860 · 0 citations · 33 references
Computer Science

TL;DR

A large-scale structured reasoning dataset constructed via a novel slice-wise data synthesis paradigm that unlocks deep volumetric understanding and highly interpretable clinical logic without requiring computationally expensive 3D-specific pre-training is introduced.

Abstract

While Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in 2D medical image understanding, their extension to 3D volumetric imaging remains hindered by prohibitive annotation costs and dataset opacity. Current data formats, predominantly consisting of rigid Visual Question Answering (VQA) pairs or unstructured final clinical reports, typically fail to capture explicit clinical reasoning. To address this limitation, we introduce a large-scale structured reasoning dataset constructed via a novel slice-wise data synthesis paradigm. Inspired by the genuine diagnostic workflow of radiologists, this paradigm models visual cognition by decomposing the complex 3D reading process, translating global clinical priors into fine-grained, per-slice observations that are subsequently synthesized into an interpretable Chain-of-Thought (CoT). Crucially, this synthesized reasoning framework enforces essential clinical principles: sequential spatial tracking, multi-slice spatial awareness for artifact mitigation, and differential exclusion. To validate this approach, we instruction-tune a standard 2D-pretrained MLLM baseline using the synthesized data to enhance its volumetric comprehension. Comprehensive evaluations across multiple 3D medical benchmarks demonstrate that our method yields significant performance improvements over the 2D baseline. Furthermore, the resulting model exhibits robust spatial reasoning capabilities and rivals resource-intensive native 3D architectures, effectively bridging the performance gap. Ultimately, this data-centric strategy unlocks deep volumetric understanding and highly interpretable clinical logic without requiring computationally expensive 3D-specific pre-training. The complete repository, including datasets and training workflows, is publicly available at https://github.com/2020420145009/hounsfield.

View source

Similar papers

Review Jul 2026

Multimodal AI in healthcare: Review of vision-language foundation models for real-world medical applications.

A definitive taxonomy of the medical VLM landscape is provided, tracing the evolution from early Contrastive Alignment and Generative MLLMs to the cutting-edge frontiers of Dense Pixel-Grounding, Sparse Mixture-of-Experts (MoE), and Reasoning-Incentivized (RL) architectures.

Taha Razzaq, Murtaza Taj, Asim Iqbal · 0 citations
Preprint Aug 2026

LocAnyMed: Vision-Language Grounding for Multimodal Medical Images

LocAnyMed-CoT-20K is derived, a rationale-augmented subset that connects anatomical context, visual observations, and spatial conclusions through structured reasoning and further improves cross-source generalization through fine-tuning.

Zi-hao Wang, Tong Liu, Zhiwei Wang et al. · 0 citations
Jul 2026

RadSight: Towards Perceptually Reliable Multimodal Radiology Image Understanding

RadSight is proposed, a perception-driven MLLM built upon a dual 2D/3D encoder architecture that preserves native imaging spatial structures that achieves consistent improvements on public 2D and 3D medical benchmarks, further demonstrating that robust low-level visual perception is a critical foundation for reliable clinical understanding.

Jianqi Liu, Weiwei Cao, Wanxing Chang et al. · 0 citations
Jul 2026

Radiological VQA with Multimodal LLMs: Performance and Insights

The exponential growth in medical imaging volumes necessitates scalable, reliable diagnostic support systems capable of augmenting clinical workflows. This article presents a systematic quantitative evaluation of state-of-the-art Multimodal Large Language Models (MLLMs) for radiology Visual Question Answering (VQA), a task requiring integrated visual perception and clinical reasoning. We benchmark five leading models — GPT5-Nano, Gemini 3 Flash, Qwen3-VL-8B, LLaVA Next, and Llama 3.2 Vision — on the VQA-RAD dataset under a rigorous zero-shot protocol with standardized prompts and comprehensive precision–recall–F1 evaluation. Our empirical analysis reveals that Gemini 3 Flash achieves superior balanced performance (F1 = 0.78, Accuracy = 0.78, Recall = 0.83), while Qwen3-VL-8B attains the highest precision (0.78) while also maintaining competitive recall. These outcomes demonstrate that general-purpose MLLMs can perform competitively with specialized medical models in tasks such as modality and organ recognition, but still struggle with abnormality detection and complex clinical reasoning. The findings reinforce that MLLMs currently serve best as assistive decisionsupport tools rather than autonomous diagnostic agents, and highlight the potential of retrieval-augmented and context-aware strategies for improving clinical reliability and interpretability.

Cristovão Pessoa Cândido, Matheus Alves de Oliveira Lima, C. de Souza Baptista et al. · 0 citations
Preprint Jul 2026

UniMedSeg: Unified In-Context Learning for Multi-Paradigm 2D/3D Medical Image Segmentation

Medical image segmentation foundation models are expected to generalize across diverse clinical scenarios, yet existing universal methods remain fragmented by prompt paradigms and spatial dimensions. Visual in-context learning, interactive segmentation, and language-guided segmentation are typically handled by paradigm-specific models, while 2D and 3D images are also modeled separately. Such isolation prevents heterogeneous annotations and data from being jointly absorbed by a single scalable model and limits cross-paradigm knowledge transfer. To address this bottleneck, we propose UniMedSeg, a Transformer-centric universal segmentation framework that maps visual examples, geometric interactions, language instructions, and 2D/3D images into a shared sequence space, enabling heterogeneous medical supervision to be jointly learned through a unified in-context interface without prompt- or dimension-specific branches. To overcome the long-sequence memory bottleneck caused by visual contexts, we introduce Decoupled Split Attention, which reduces attention complexity to linear while preserving hardware-friendly computation and focused context-target interaction. Extensively trained and evaluated on a large corpus curated from 27 public datasets, UniMedSeg achieves state-of-the-art performance across visual in-context, interactive, and language-guided segmentation without task-specific fine-tuning, demonstrating strong generalization on diverse held-out tasks. The code and model weights are publicly available at https://github.com/Lii1228/UniMedSeg

Yunzhou Li, Jiesi Hu, Yanwu Yang et al. · 0 citations
Review Aug 2026

Volumetric Radiology AI in the Era of Multimodal Large Language Models

This Review examines volumetric foundation models, language alignment and compression strategies, and agentic systems that extend MLLMs through planning, tools, memory, and workflow interaction, and introduces a Claim-Design-Validation framework to assess whether technical, workflow, and clinical claims are matched by appropriate design and validation.

Zanting Ye, Shengyuan Liu, Xin Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.