Skip to content

Category

computer vision

3,022 papers

#machine learning Preprint Open access Sep 2026

CAMEO: A Class-Activation-Mapped Equitable Overlay Framework for Fair and Robust Deep Learning-based Skin Condition Diagnosis

Deep learning classifiers for dermoscopic skin lesions often reach high in-distribution accuracy while quietly relying on spurious background cues such as skin tone, device vignetting, and embedded rulers, rather than on lesion morphology. This undermines robustness and fairness across skin tones. This work asks whethe...

Youssef Attia, Debasmita Mukherjee · 0 citations
#machine learning Preprint Open access Sep 2026

StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks

Vision-language-action (VLA) models perform well on shorter-horizon manipulation tasks but still struggle with long-horizon tasks that require multiple dependent manipulations from a single command. Online reinforcement learning (RL) can improve these policies through environment interaction, yet many existing methods...

Ziyi Yin, Sangmin Woo, Kang Zhou et al. · 0 citations
#machine learning Preprint Sep 2026

One-Step Next-Latent Prediction Is Not a World Model

Next-latent prediction fits a map from the current embedding to the next one. LeNEPA carries this objective to time series, replacing the stop-gradient of next-embedding prediction with the isotropy penalty of LeJEPA. A world model is a transition kernel that can be rolled out. The one-step regression identifies a cond...

Shi-Tong Wang, Zhongang Cai, Yu Hong · 0 citations
#machine learning Preprint Open access Sep 2026

Exploring Learning Models for Topological Relationship Recognition from Image Data

Figuring out how objects relate to each other, like whether they touch, overlap, stay completely separate or one sits inside another, matters a lot in fields like GIS, biomedical imaging, and robotics. Even though machine learning has come a long way, people haven't really focused on spotting these topological relation...

Saptak Das, Monidipa Das · 0 citations
#machine learning Preprint Open access Sep 2026

Hardware-Aware Functional Kolmogorov-Arnold Networks for Efficient Medical Image Enhancement and Segmentation

Functional Kolmogorov-Arnold Networks (FunKAN) achieve state-of-the-art accuracy on MRI Gibbs artifact removal and anatomical segmentation, but their 11.6 M parameters and 8.7 GFLOPs are too large for edge medical devices. We present FunKANLite, a two-stage, hardware-aware compression of FunKAN for point-of-care use. F...

Mohammad Sadegh Sirjani · 0 citations
#artificial intelligence Preprint Open access Sep 2026

PrototypeNAS: Rapid Design of Deep Neural Networks for Microcontroller Units

Enabling efficient deep neural network (DNN) inference on edge devices with different hardware constraints is a challenging task that typically requires DNN architectures to be specialized for each device separately. To avoid the huge manual effort, one can use neural architecture search (NAS). However, many existing N...

Mark Deutel, Simon Geis, Axel Plinge · 0 citations
#machine learning Preprint Open access Sep 2026

Procedural Core: A Compact Recurrent Initialization for Vision Transformers

Transformers are typically trained from random initialization, requiring all their capabilities to emerge from large-scale optimization. Recent work showed that a small amount of abstract procedurally generated data can help acquire generic inductive structure at low cost. However, this adds a pretraining stage that mu...

Zachary Shinnick, Christian Intern\`o, Hemanth Saratchandran et al. · 0 citations
#machine learning Preprint Sep 2026

Scaling Full Conformal Image Classifiers

Conformal prediction provides set-valued predictions with distribution-free coverage guarantees, making it attractive for high-stakes image classification. However, split conformal prediction is data-inefficient, while full conformal prediction (FCP), despite its stronger statistical efficiency, is computationally proh...

Julio Silva-Rodríguez, E. Konukoglu · 0 citations
#machine learning Preprint Open access Sep 2026

PE-OPSD: Internalizing Prompt Enhancement into Flow-matching Models via On-Policy Self-Distillation

Text-to-image users often provide concise and underspecified prompts, whereas generative models benefit from detailed textual conditions for reliable instruction following. Existing systems bridge this gap with Prompt Enhancers (PEs) that rewrite raw prompts at inference time, introducing additional latency and leaving...

Mingfeng Lin, Chengfei Cai, Lin Xu et al. · 0 citations
#machine learning Preprint Open access Sep 2026

On the spectral properties of generative denoiser Jacobians

Generative denoising models, such as diffusion and flow-matching, learn to sample from complex distributions by training a deep neural network denoiser to recover clean data from noise-corrupted samples. While such models are typically compared on the quality of their synthesized samples, these metrics provide limited...

Alexandros Graikos, Nebojsa Jojic, Dimitris Samaras · 0 citations
#machine learning Preprint Sep 2026

The Decision Value of Perception Compute

This work introduces DEEP (Decision Evaluation for Escalated Perception), a benchmark that scores pre-escalation allocators against this oracle under selection, latency and energy budgets, charging each allocator for its own computation.

Cong Pham Hoang, Ho Viet Duc Luong · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.