Skip to content

Category

computer vision

3,022 papers

#machine learning Preprint Sep 2026

Dexterous Tactile World Model

World models for manipulation are typically trained from video, yet the events that determine how manipulation unfolds, such as making and releasing contact, are difficult to observe visually and are often easier to sense through touch. We present the Dexterous Tactile World Model (DTWM), a video world model for future...

Zi-Yao Zeng, Xiatao Sun, Hao Wang et al. · 0 citations
#machine learning Preprint Sep 2026

3D Point Tracking with State Space Models

This paper proposes a 3D point tracker accurate in those absolute terms and operating within a single commodity GPU, pose-free, monocular budget, exceeding strong feed-forward trackers and a companion analysis explains why several published trackers lose most of their accuracy under this budget.

Masahiro Ogawa, Qi An, Atsushi Yamashita · 0 citations
#machine learning Preprint Open access Sep 2026

Residual-Stream Burden Shapes Representation Learning in Diffusion Transformers

In diffusion-based generation, a neural network can be trained to predict the clean data, the noise, or the velocity from a noisy input. These prediction targets are interconvertible and describe the same generative process, yet plain Diffusion Transformers operating on large pixel patches succeed with clean prediction...

Tongtong Liang, Siqi Kou, Ziqiao Xi et al. · 0 citations
#machine learning Preprint Sep 2026

ViBR-WM: Visual Bayesian Regression for World Modeling

Across four forecasting tasks spanning object motion, vegetation greenness and solar power, ViBR-WM achieves lower mean overall physical-target error than Temporal Straightening, ConvLSTM, PredRNN and SimVP on every task.

Ji-Fan Li, Ning Ning · 0 citations
#machine learning Preprint Sep 2026

CLIMB-flow: Coupled Linear Inverse posterior sampling via Multiscale-Based flow

Diffusion models are now widely used in Bayesian inverse problems in imaging as priors, where latent diffusion models are often used for larger scale problems to keep the computational complexity and model-size manageable. Unfortunately, the auto-encoder based compression results in loss of spatial detail. In addition,...

Ze-Qiu Yu, Rui Yuan, Mathews Jacob · 0 citations
#machine learning Preprint Sep 2026

One Attack to Fool Them All: Highly Transferable Black-Box Adversarial Attacks on Frontier MLLMs

Adversarial attacks have long posed a fundamental threat to machine learning systems. As multimodal large language models (MLLMs) rapidly evolve and become widely deployed, assessing their vulnerability to such attacks is essential for their safe use. In this work, we investigate whether a single adversarial image can...

Sen Nie, Jie Zhang, Zhong Ling Wang et al. · 0 citations
#machine learning Preprint Sep 2026

Constrained Edit Fields for Training-Free Flow Editing

Constrained Edit Fields (CEF) achieves state-of-the-art Structure Distance, background LPIPS, and background MSE with both Stable Diffusion 3.5 Medium and FLUX, while retaining competitive instruction alignment.

Jing-Xuan Kang, Yin-Song Wang, Che Liu et al. · 0 citations
#machine learning Preprint Sep 2026

Prompt-Anchored Residual Adaptation for Biomedical Vision-Language Models

PARA is proposed, which retains the frozen prompt prediction as a support-invariant semantic anchor and incorporates a visual prediction learned from the support set through an anchor-relative residual and achieves state-of-the-art performance in both few-shot classification and base-to-novel generalization.

Jing-Xuan Kang, Qian-Ying Yue, Che Liu et al. · 0 citations
#machine learning Preprint Open access Sep 2026

A Visual Classification Dataset and Model Evaluation for Historical Manuscript Illustrations

Historical manuscript illustrations preserve rich visual evidence of past cultures. They depict people, animals, plants, diagrams, music notations, and decorative forms. Although large digitization projects have made many manuscripts available online, the material itself remains difficult to explore at scale. Extractio...

Yoav Evron, Michal Bar-Asher Siegal, Michael Fire · 0 citations
#machine learning Preprint Open access Sep 2026

Parameter-Efficient 3D Segmentation of Liver and Liver tumors: Depthwise factorization Scales Better Than Dense Convolution with Spatial Dimensionality

Three-dimensional dense convolutional networks are the strongest performers on volumetric medical image segmentation, but their parameter counts scale poorly: moving a dense k x k convolution to k x k x k multiplies its weights by k. We observe that depthwise separable factorization does not share this penalty. Because...

Adham M. Alkhadrawi, Mohammed A. B. Mahmoud · 0 citations
#machine learning Preprint Sep 2026

Distributed Hydrological Modeling in the Feature Space

This work introduces feature-space routing: a topology-aware state-space operator embedded directly in the forecasting dynamics, which preserves the physical connectivity of the river system while allowing the propagated state itself to be learned end-to-end.

Mohamad Hakam Shams Eddin, M. L. Taccari, Yi-Kui Zhang et al. · 0 citations
#machine learning Preprint Open access Sep 2026

Is H&E Image-to-Spatial Transcriptomics Simpler Than It Looks?

Predicting spatial gene expression from routine H&E histology offers a scalable route toward spatial molecular profiling. Recent work has pursued increasingly sophisticated architectures to capture spatial context and richer expression structure. At the same time, simple estimators have shown strong performance in seve...

Duc T. Nguyen, Thanh Ha Do, Phuong M. Cao et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.