Skip to content

Category

computer vision

3,022 papers

#machine learning Preprint Open access Oct 2026

Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions

Using behavioural science, health interventions focus on behaviour change by providing a framework to help patients acquire and maintain healthy habits that improve medical outcomes. In-person interventions are costly and difficult to scale, especially in resource-limited regions. Digital health interventions offer a c...

Manuela Gonz\'alez-Gonz\'alez, Soufiane Belharbi, Muhammad Osama Zeeshan et al. · 0 citations
#machine learning Preprint Open access Oct 2026

A Sobel-Gradient MLP Baseline for Handwritten Character Recognition

This study examines how much handwritten-character information is retained by a deliberately simple first-order edge representation. Instead of learning spatial filters, each input image is transformed by the fixed Sobel-Feldman operator into signed horizontal and vertical derivative maps, which are independently norma...

Azam Nouri · 0 citations
#machine learning Preprint Open access Oct 2026

Flow Map Denoisers: Traversing the Distortion-Perception Plane for Inverse Problems

Image restoration faces a fundamental tradeoff: methods that minimize error produce blurry reconstructions, while those that maximize perceptual quality yield sharp but less faithful images. Existing approaches either commit to a single operating point on this distortion perception (DP) frontier or require paired-data...

Nicolas Zilberstein, Morteza Mardani, Santiago Segarra · 0 citations
#machine learning Preprint Open access Oct 2026

How Far Does a Shared Linear Map Go? Probing Feature-Space Manipulability for Image Editing

Understanding how image-space transformations manifest in a model's internal representations is a longstanding goal in representation analysis. Prior work has shown that geometric transformations can often be captured by learned linear operators between feature maps, but it remains unclear whether this extends to photo...

Elias Krey, Nils Neukirch, Nils Strodthoff · 0 citations
#machine learning Preprint Open access Oct 2026

Can AI Understand the Language of Origami?

Building AI systems that can plan, act, and create in the physical world requires more than pattern recognition. Such systems must reason about the generative mechanisms and constraints governing physical processes, using structured representations that connect observations, actions, and their effects. Yet, many existi...

Naaisha Agarwal, Yihan Wu, Xin Guan et al. · 0 citations
#machine learning Preprint Oct 2026

XGenAct: Geometry-Enhanced World Action Models through Cross-Task Generation

World action models (WAMs) have advanced robot control by predicting how observations and actions evolve over time. Despite this progress, RGB and action based future prediction does not explicitly address the spatial understanding needed for robot manipulation. Existing efforts often add a limited set of spatial predi...

Ting-Ting Du, Zi-Yao Wang, Guoheng Sun et al. · 0 citations
#machine learning Preprint Oct 2026

Iterating Consistency Models: Stability, Error Bounds and Noise Schedules

Consistency models (CMs) have become a leading approach for generating high-quality samples in few steps. However, adding steps can improve or degrade sample quality in ways that are highly sensitive to the schedule and that existing theory does not fully explain. To provide accuracy guarantees and guide CM sampler des...

Alessio Spagnoletti, A. Haji-Ali, Andrés Almansa et al. · 0 citations
#machine learning Preprint Oct 2026

From Patching to Pruning Visual Computation in Vision Language Models

Vision language models (VLMs) incur substantial inference cost because every visual token is processed by the attention and MLP projections of every decoder layer, even when token-specific visual computation is unnecessary at many depths. We introduce Patch-to-Prune (P2P), inspired by Mechanistic Interpretability, a tr...

Rahul Chowdhury, Timothy Rupprecht, Xuan Shen et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Wrong Organ, Right Physics: Transferring Echocardiography Pretraining to Lung Ultrasound for Tuberculosis Screening

Lung ultrasound (LUS) is attractive for tuberculosis (TB) screening at primary-care level, but labelled cohorts are small. Echocardiography carries no such constraint, while sharing the same underlying ultrasound imaging physics, signal processing and B-mode appearance as LUS. We ask whether an encoder pretrained on th...

Christiaan M. Geldenhuys, Joshua M. Jansen van V\"uren, V\'eronique Suttels et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Does Physics Live in the Activations? Localizing Physical Quantities in Video Diffusion Models

Video generation models produce strikingly realistic sequences and are increasingly proposed as world models, yet recent benchmarks reveal pronounced deficits in their physical reasoning. This raises the question of whether these models internalize physical principles or merely reproduce familiar motion patterns. We ad...

Jonas Kneifl, Jakub Skalski, Bart{\l}omiej Twardowski et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Confidence-Gated Cloud-Edge Cascade Triage via Variational Risk Minimization for Medical Imaging

Emergency chest X-ray (CXR) triage has a structural modality gap: reports arrive after triage decisions, yet multimodal foundation models require image-text inputs. We present Variational Risk Minimization (VRM), a distillation framework that treats LVLM-generated report variants as Monte Carlo samples of latent clinic...

Xinye Yang, Zhusi Zhong, Scott Collins et al. · 0 citations
#machine learning Preprint Oct 2026

Adaptive Second-Order Solvers for Fast Stochastic Diffusion Sampling

Diffusion models rely on numerical solvers requiring time-discretization, which has a large influence on the tradeoff between sampling cost and quality. However, the computational difficulty of the reverse process varies along the sampling trajectory and across data distributions, making the choice of discretization im...

E. Kemperman, Luca Ambrogioni · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.