Skip to content

Category

computer vision

3,022 papers

#machine learning Preprint Sep 2026

SynCo: Learning Cross-Modal Synergy by Contrasting Interaction Residuals

Multimodal contrastive learning is a dominant paradigm for learning transferable representations from unlabeled data, but standard objectives primarily capture information that is redundant between modalities. Partial Information Decomposition (PID) shows that task-relevant information in multimodal data decomposes int...

Yavuz Yarici, Ghassan AlRegib · 0 citations
#machine learning Preprint Open access Sep 2026

VCRE-Fib: View-Conditioned Regional Evidence for Fine-Grained Ultrasound Grading of Schistosoma japonicum-Associated Liver Fibrosis

Accurate assessment of Schistosoma japonicum-associated liver fibrosis is essential for disease management and long-term follow-up in endemic regions. Ultrasound provides non-invasive imaging, but complex local echogenic patterns and anatomical structures make fine-grained grading challenging. Existing deep learning me...

Ziyang Xu, Shuli An, Hao Zhou et al. · 0 citations
#machine learning Preprint Sep 2026

USAI-Quant: A Quantitative Reasoning Benchmark for Vision-Language Models in Built Environments

The USAI-Quant benchmark is developed, the first benchmark designed to quantitatively evaluate VLM's reasoning capabilities on built environment metrics via remote sensing imagery, and reveals that current state-of-the-art models consistently fall short on numeric reasoning tasks.

Dong-Dong Wang, Q. Song, Yu-Zhou Chen et al. · 0 citations
#machine learning Preprint Sep 2026

CT-OPD: Counterfactual Trace On-Policy Distillation for Diffusion Vision-Language Models

Counterfactual Trace On-Policy Distillation (CT-OPD), which combines completed teacher responses with trajectory masks from the current student, consistently enhances multimodal understanding and reasoning capabilities, and improves both visual understanding and image generation.

Long Qian, Bing-Ke Zhu, Jia-Qi Wei et al. · 0 citations
#machine learning Preprint Sep 2026

OmniMoE-VL: A Sparse Vision-Language Model with Coupled Visual-Depth Routing

OmniMoE-VL, a sparse VLM with a coupled visual-depth routed projector, enables question-dependent visual access while preserving the native visual-token sequence, and complements token-level expert routing in the vision and language stacks.

Long Qian, Bing-Ke Zhu, Jia-Qi Wei et al. · 0 citations
#machine learning Preprint Open access Sep 2026

An Empirical Study on What Matters for Viewpoint-Generalizable Policies in Visual Imitation Learning

Visual imitation learning is a promising approach to training robot manipulation policies capable of completing a wide variety of tasks. However, policies today remain brittle to viewpoint perturbations, making deployment in diverse environments a challenge. We present a controlled empirical study of which design choic...

Mino Nakura, Sriram Krishna, Yufei Wang et al. · 0 citations
#machine learning Preprint Sep 2026

JEPA Learns What the Mask Leaves Unrecoverable

Joint-embedding predictive architectures are unusually sensitive to how the input is masked: block masks work, scattered masks do not, and the explanations are empirical. We give a measurement account. A mask is a linear measurement, and in a compactly supported wavelet basis every atom whose support lies inside the hi...

Peng Xie, Amr Alanwar · 0 citations
#machine learning Preprint Sep 2026

Federated Subspace Guided Vision-Language-Action Policy Distillation for Non-IID Multi-Robot Manipulation

Federated learning offers a natural way for multiple robots to jointly improve manipulation policies without requiring centralized access to training demonstrations. However, non-IID task and environment distributions can induce representation drift and mutually incompatible robot-policy updates, making naive parameter...

Biprodip Pal, Kaushik Roy, Yan-Ming Zhu et al. · 0 citations
#machine learning Preprint Sep 2026

Kernel-Based Steering of CLIP with Vision-Language Model Preferences

Large vision-language models (VLMs) can judge visual similarity, but their judgments are not directly available as compact image embeddings for efficient comparison. We study how to transfer these preferences into CLIP while retaining its image--text capabilities. We introduce ASK, a kernel-based steering method that l...

Sajjad Ghiasvand, Haniyeh Ehsani Oskouie, Sina Mansouri et al. · 0 citations
#machine learning Preprint Sep 2026

Contamination, Prior, or Evidence? Decomposing and Training Evidence Use in Whole-Slide Vision-Language Models

Pathology vision-language models (VLMs) are conventionally evaluated by accuracy, but accuracy alone does not measure evidence use: it may conflate dataset contamination, prior knowledge, and image evidence. In a motivating study of lymph-node metastasis prediction, we found that most public pathology VLMs showed minim...

Wen-Hao Zhang, Zhong-Liang Zhou, Shi-Yuan Zhang et al. · 0 citations
#machine learning Preprint Sep 2026

Panoptic Scene Program Diffusion Transformer

Modern text-to-image models produce high-fidelity images but still struggle with compositional prompts that require instance identity, attribute ownership, counting, spatial ordering, and role-sensitive relations. We introduce Panoptic Scene Program Diffusion Transformer (PSP-DiT), a diffusion-transformer architecture...

C. Maduabuchi · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.