Skip to content

Category

computer vision

3,022 papers

#machine learning Preprint Oct 2026

Readout Blindness: VLM Scores Miss the Spatial Direction Their Frozen Encoders Retain

CLIP-like vision-language models remain a cornerstone of multimodal systems, yet their scores stay near chance on directed spatial relations, such as whether one object is left of another. We call this failure readout blindness and analyze, theoretically and empirically, why deployed scores miss the direction: when sco...

Guang-Yuan Li, Tian-Ming Du, Yan Jiang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

How well do routinely collected demographic and clinical variables aid point-of-care lung ultrasound TB classification

We consider the fusion of lung ultrasound images with routinely-collected clinical and demographic data for the purpose of automated tuberculosis (TB) screening using deep-learning. Such deep-learning based screening tools for TB could meaningfully support the health care system in Africa, where the burden of disease i...

Joshua M. Jansen van V\"uren, Christiaan M. Geldenhuys, Devendra S. Parihar et al. · 0 citations
#machine learning Preprint Oct 2026

Fitting Vision Adapters at Frontier Scales

Training a small projector between a frozen vision encoder and language model is an established approach to multimodal learning. As the parameter count of language models scales dramatically, we revisit which vision capabilities this approach can add while keeping their pretrained weights fixed. Here we train a 50M par...

Jaehoon Lee, Harry B. Partridge, M. Jayasekara et al. · 0 citations
#machine learning Preprint Open access Oct 2026

The Poisoned Conversation: Privacy-Leaking Watermarks in Unified Multimodal Models

Multimodal models are increasingly shifting toward unified architectures that understand and generate text, images, and other modalities within a shared conversational context. This design enables fluid interaction across modalities, but it also changes the privacy threat model: Information revealed in one part of a co...

Tobias Braun, Jonas Henry Grebe, Emil Sivic et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Unmentioned Checklist Findings Change How Reinforcement Learning Appears to Improve Chest Radiograph Report Checking

Automated checks of radiology reports may rely on AI-generated checklists that leave findings unmentioned. We used reinforcement learning to train a vision-language model to fill in a 12-finding checklist from a chest radiograph without seeing the sentence under test; a separate checking model judged the sentence from...

Ali Vosoughi, Akhil Kasturi, Chenliang Xu et al. · 0 citations
#machine learning Preprint Open access Oct 2026

TIRMamba: A Thermal-Prior-Modulated State-Space Network for Sub-Million-Parameter Infrared Image Super-Resolution

Infrared image super-resolution is currently led by Mamba-based networks with 26 to 37 million parameters, which are difficult to deploy on the airborne and handheld platforms where thermal imaging is most needed. This paper presents TIRMamba, a network with 896K to 910K parameters for single-channel thermal imagery. A...

Chun-An Lin, Tsung-Jung Liu, Yen-Chieh Ouyang · 0 citations
#machine learning Preprint Open access Oct 2026

One Tile, Multiple Instances: Rethinking MIL for Sparse Diagnostic Evidence

In weakly supervised Whole Slide Image (WSI) classification, feature extractors typically compress each image tile into a single global embedding. Consequently, slide-level aggregators are restricted to this coarse tile scale, concealing fine-grained sub-tile evidence from the attention mechanism. We introduce DI-MIL,...

Runsheng Liu, Cheng Jin, Hao Jiang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

MOXIE: Discovering Alternative Explanations for Biomedical Image Classifiers

Segment-based explanation methods such as LIME return a single explanation for each prediction, computed from one fixed image segmentation. This hides two important facts: a prediction can be supported by many different sets of image segments, and the segmentation itself shapes which explanations can be found. We intro...

Abiha Tahsin Chowdhury, Rahul Dubey · 0 citations
#machine learning Preprint Open access Oct 2026

WASP: Weakly Aligned Spatiotemporal Pairs for Fetal Brain MRI-Ultrasound Learning

Magnetic Resonance Imaging (MRI) is widely regarded as the optimal sensor for fetal brain analysis due to its superior soft-tissue contrast and anatomical detail. However, its high cost and operational burden make it invasive and difficult to obtain at scale. Ultrasound (US), in contrast, is cheap, safe, and routinely...

Francesco Correnti, Gabriele Magrini, Marco Mistretta et al. · 0 citations
#machine learning Preprint Oct 2026

Homogeneous Semantic Alignment and Hierarchical Expert Routing for Radiology Report Generation

Radiology report generation (RRG) aims to convert medical images into diagnostic texts to assist in clinical decision-making and alleviate the workload of physicians. Although existing methods have made extensive progress in cross-modal interaction and the incorporation of external priors, the distribution shift of und...

Er-Jian Zhang, Jia-Yuan Ma, Lie-Jun Wang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Where Does the Semantic Gain Come From? A Reproduction and Extension of Semantic Knowledge-driven Contrastive Learning for Long-Tailed Recognition

Semantic Knowledge-driven Contrastive Learning (SKCL) uses a language model to decide which classes are related, and pulls each image towards the prototypes of its semantic neighbours. On CIFAR-100-LT (beta = 100) it reports 54.02% top-1 accuracy, 2.01 points above Balanced Contrastive Learning (BCL), the method it bui...

Sushrut Ghimire · 0 citations
#machine learning Preprint Open access Oct 2026

Scaling 3D Visual Grounding in Abdominal CT

Visual grounding models can enhance radiology workflows by linking report findings to image regions. This is particularly valuable for 3D CT, where findings often occupy a tiny fraction of the volume. Training 3D grounding models requires large sets of paired phrases and regions, and building such datasets is expensive...

Sam Church, Danyal Maqbool, Joshua D. Warner et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.