Skip to content

Category

computer vision

3,022 papers

#machine learning Preprint Sep 2026

PoE-Fuse: Precision-Weighted Expert Fusion for Bi-Temporal Change Understanding

Bi-temporal change understanding, which localizes and characterizes what changed between two satellite images, is central to disaster response and environmental monitoring, spanning change detection, building localization, and damage assessment. Strong vision-language models address these tasks, but adapting them typic...

Haruki Watase, Shunya Nagashima, T. Nishimura · 0 citations
#machine learning Preprint Sep 2026

Learning Social Navigation from Internet Videos in the Policy State Space

Training robust social-navigation policies requires simulators with diverse scene layouts, terrain, and human motion, but constructing such environments and specifying pedestrian behavior is costly. We propose an efficient pipeline that converts ordinary monocular walking videos directly into closed-loop social-navigat...

Jia-Ming Wang, Duc Thang Nguyen, Ji-Zhuo Chen et al. · 0 citations
#machine learning Preprint Open access Sep 2026

Multi-task learning for the automatic grading of enlarged perivascular space burden using MRI

Enlarged perivascular spaces (PVS) visible in brain magnetic resonance imaging (MRI) are increasingly thought to be linked to poor brain health. PVS are elongated structures of less than 3 mm in diameter and can be numerous. To reflect the incidence of PVS, radiologists visually score their burden following a clinical...

Jesse Phitidis, William N. Whiteley, Joanna M. Wardlaw et al. · 0 citations
#machine learning Preprint Sep 2026

The Domain Is a Residue: Adapting Self-Supervised Features, Not Generators

Clearing fog, rain or snow from footage, or turning renders into photographs, must remove the source domain and keep the scene. Unpaired translators carry it through because their generator sees the source appearance (pixels, a near-invertible latent or a control map) and keeps it. A DINO feature map fixes what is in t...

Thomas Deixelberger, Markus Steinberger · 0 citations
#machine learning Preprint Sep 2026

Seeing Is Not Addressing: Auditing Linguistic Access to Frozen Visual Geometry

Visual distinctions are often finer than those reflected in linguistic conceptualization. Vision-language models exhibit a similar asymmetry: a distinction can remain discriminable in frozen image geometry while being weakly addressable through the native text interface. We study this gap by separating visual discrimin...

Woosang Jeon, Jiwon Yang, Chung Soo et al. · 0 citations
#machine learning Preprint Sep 2026

Sparse cubical complexes for efficient topology-preservation in image data

Persistent homology (PH) is a frequently used tool for extracting and preserving topological information from image data, particularly in image segmentation, where preservation of topological structures is important. However, despite its general applicability across dimensionality, domains, and target structures, the r...

Alexander H. Berger, Marco Fontana, Daniel Rueckert et al. · 0 citations
#machine learning Preprint Sep 2026

Improved Distributional Diffusion Models

Distributional Diffusion Models (DDMs) replace the standard mean-prediction denoiser with a \emph{distributional} denoiser trained via a scoring rule objective, learning a stochastic approximation to $p(x_1 \mid x_t)$ rather than its conditional mean. However, scaling DDMs to modern image-generation settings faces two...

Tommaso Martorella, Alexandre Galashov, F. Krause et al. · 1 citation
#machine learning Preprint Sep 2026

Multi-Depth Temporal Fusion for Feedforward, Locally Trained Spiking Neural Networks

We propose a new spiking neural network (SNN) design to process static images and event streams using time-to-first-spike (TTFS) latencies. Our key research question is which architectural choices best accommodate local and online learning in multi-layer convolutional SNNs. This question is addressed via an original fr...

Aidin Attar, Eleonora Cicciarella, Michele Rossi · 0 citations
#machine learning Preprint Sep 2026

GleanVID: Complementary Token Selection for Efficient Video Large Language Models

Video Large Language Models (VideoLLMs) have achieved strong video understanding capabilities but incur substantial inference overhead due to the large number of visual tokens. Existing VideoLLM token compression methods largely rely on selection-independent scoring, overlooking cross-frame complementarity and conseque...

Shuo Yang, Chang-Bai Li, Rui Tang et al. · 0 citations
#machine learning Preprint Open access Sep 2026

You Cannot Recover What Was Never Measured: Quantifying the Information Ceiling of Ultra-Low-Field MRI Super-Resolution

Generative super-resolution models can turn portable 64 mT MRI into images that look like 3T scans, and the field evaluates them with PSNR, SSIM, and pixelwise uncertainty, most often on pairs built by synthetically degrading high-field images. Prior work acknowledges that these models hallucinate and that the problem...

Prathamesh Pradeep Khole, Shreya Handa, Utkarsh Gupta et al. · 0 citations
#machine learning Preprint Sep 2026

Seeing What Should Be Heard: Diagnosing and Repairing Cross-Modal Shortcuts in Omni-Modal LLMs

Omni-modal large language models (LLMs) are expected to answer a question using the modality it explicitly refers to. However, existing training paradigms rarely verify whether models actually follow this modality, because multimodal inputs from the same sample often provide redundant evidence for the same answer. In t...

Yue-Ran Ma, Ronghao Lin · 0 citations
#machine learning Preprint Open access Sep 2026

Losing the name before the box: measuring and repairing what narrow fine-tuning costs a detector outside its deployment vocabulary

A detector pretrained on a broad corpus is fine-tuned on a narrow domain, its in-domain accuracy improves, and it ships. We ask what happens meanwhile to its coverage of objects the vocabulary never names, which in obstacle detection and inspection carry the risk. No in-domain test set holds an example of one. We give...

Trung Minh Bui, Jongsul Moon, YoungOuk Kim et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.