Skip to content

Category

computer vision

3,022 papers

#machine learning Preprint Open access Oct 2026

DART: A Degradation-Aware Recurrent Transformer for Archival Film Restoration

Archival film restoration is a challenging problem because historical footage contains compound degradations such as scratches, dust, blur, noise, flicker, and photometric aging, while clean reference videos are unavailable. Existing video restoration methods largely treat these degradations implicitly, reconstructing...

Miko{\l}aj Jastrz\k{e}bski, Wojciech Koz{\l}owski, Kamil Adamczewski · 0 citations
#machine learning Preprint Open access Oct 2026

Mask-supervised Object-centric Representation Learning with LeJEPA

Self-supervised image encoders deliver strong features for downstream tasks but need many images for training. A natural remedy to counter this is to make each image count for more. A scene contains many objects, and given masks from human annotators or an off-the-shelf segmentation model, pre-training can focus on ali...

Jakob Geusen, Ender Konukoglu · 0 citations
#machine learning Preprint Open access Oct 2026

Fine-Grained Caching for Diffusion Transformers with Few Calibration Conditions

Diffusion transformers require repeated denoiser evaluations, making image and video generation computationally expensive. We propose a training-free framework that uses a few calibration conditions to construct a fixed, fine-grained module-reuse schedule without schedule search. The design is motivated by an empirical...

Zihao Wu, Bohan Zeng, Yuanxing Zhang · 0 citations
#machine learning Preprint Open access Oct 2026

MIND: Microstructure INverse Design with Generative Hybrid Neural Representation

The inverse design of microstructures plays a pivotal role in optimizing metamaterials with specific, targeted physical properties. While traditional forward design methods are constrained by their inability to explore the vast combinatorial design space, inverse design offers a compelling alternative by directly gener...

Tianyang Xue, Longdu Liu, Lin Lu et al. · 0 citations
#machine learning Preprint Open access Oct 2026

The Label Complexity of Useful Class-Conditional Prediction Sets under Distribution Shift

Prediction sets can make deployed classifiers safer by returning several plausible labels when a single prediction is uncertain. Their value depends on classwise reliability: average coverage can meet its target while rare or difficult classes fail repeatedly. This concern is sharper after distribution shift, when cali...

Weijia Han, Lisha Qu, Tianxin Zhou et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Flash-WAM: Modality-Aware Distillation for World Action Models

World-action models (WAMs) jointly generate future video and robot actions through iterative diffusion, achieving strong performance on manipulation benchmarks but requiring tens of denoising steps, a cost that precludes real-time control. Step distillation has emerged as the natural remedy, but off-the-shelf methods b...

Arman Akbari, Ci Zhang, Arash Akbari et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Sample-wise Targeted Adversarial Attacks on Test-time Adaptation

Test-time adaptation (TTA) mitigates distribution shifts by adapting models to unlabeled test inputs, but also exposes them to adversarial manipulation. Existing class-wise targeted attacks remain suboptimal for stealthy exploitation in this setting: since TTA operates on batches, forcing a subset of samples toward a t...

Phuc Duc Nguyen, Quang Duc Nguyen · 0 citations
#machine learning Preprint Open access Oct 2026

Resisting Adversarial Attacks in Deep Neural Networks using Diverse Decision Boundaries

The security of deep learning (DL) systems is an extremely important field of study as they are being deployed in several applications due to their ever-improving performance to solve challenging tasks. Despite overwhelming promises, the deep learning systems are vulnerable to crafted adversarial examples, which may be...

Manaar Alam, Shubhajit Datta, Debdeep Mukhopadhyay et al. · 0 citations
#machine learning Preprint Oct 2026

RealtimeWAM: One-Step Asynchronous World Action Models

World Action Models (WAMs) incorporate visual representations from video generation backbones to guide action prediction. Recent efficient WAMs adopt Mixture-of-Transformers (MoT) architectures and compute video representations once for reuse by the action expert. However, intra-expert iteration (\ie, multi-step action...

Cheng-Tao Lv, Jin-Yang Du, Shu-Yi Feng et al. · 0 citations
#machine learning Preprint Oct 2026

NeuroCBIR: A Fast and Accurate Image Retrieval System for Whole-Brain and Region-Specific MRI

Content-based image retrieval (CBIR) in neuroimaging enables the identification of structurally similar brain scans, supporting diagnosis, prognosis, and treatment planning; however, existing methods are often limited to small datasets, single brain regions, or coarse class labels, thereby restricting their clinical ut...

Félix Nieto-del-Amor, Jing-Ru Fu, J.-Sebastian Muehlboeck et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Environmental sensor readings in two crop disease image datasets identify the session in which each image was taken

Integrating environmental sensor data with leaf imagery is widely reported to boost crop disease classification accuracy. In this work, we reveal that these reported gains are often artifacts of dataset construction: because a single sensor reading is shared across many images collected in a single session (one farm on...

Sungwoo Kang · 0 citations
#machine learning Preprint Open access Oct 2026

Dual Variational Autoencoders for Efficient Sim-to-Real Transfer in Low-Cost Robotic Navigation

Vision-based autonomous navigation for low-cost robots remains a fundamental challenge, primarily due to the significant gap between simulated training environments and real-world operational conditions. Direct policy transfer from simulation is often ineffective, while training exclusively on real data is impractical....

\'Alvaro D\'iez (Department of Computer Science, Artificial Intelligence, University of Alicante) et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.