Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Open access Oct 2026

On-Board Anomaly Detection for Efficient Marine Environmental Monitoring

Marine ecosystems are impacted by various threats such as oil spills, algal blooms, and sediment floods, which disrupt habitats, wildlife, and human activities. Advances in satellite imagery and Artificial Intelligence (AI) have enhanced our capabilities for early detection and mitigation of such hazards. In this paper...

Thomas Goudemant, Clotilde Szywala, Benjamin Francesconi et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

LoGo: Local-Global Rewards for Consistent Long-Horizon Video Generation

Camera-controlled video models are rapidly advancing toward long generation horizons and complex camera control. A key failure mode is 3D inconsistency: as the camera moves, objects lose permanence and scene structures shift. Existing post-training techniques, which assign a single scalar reward to the entire generatio...

Ziqi Ma, Shreya Sharma, Mohamed El Banani et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Rethinking What to Cache in Few-Step Diffusion Transformers: Solver-Aware Target Selection

Diffusion Transformers (DiTs) can generate high-quality images and videos, but generating each sample requires multiple costly DiT forward passes. Two common ways to accelerate DiT sampling are step distillation, which reduces the number of sampling steps, and caching, which skips some DiT evaluations by reusing a tens...

Shuo Yang, Lihao Fang, Yi Zhang et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Weave Forcing: Compositional Memory Routing for Interactive Long Video Generation

Recent advances in autoregressive video generation have improved temporal consistency over extended durations, yet interactive storytelling requires more than continuous scene extension: a new shot may combine characters and backgrounds from different historical shots. Whole prompt retrieval can overlook the distinct r...

Zi-Yi Wang, Jun-Chi Yao, Heqian Qiu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Preserving Anatomical Continuity: Three-Stage Pipeline for Colon Segmentation in 3D Abdominal CT Scans

Accurate colon segmentation from CT images is essential for colorectal disease analysis, yet deep learning based methods often produce disconnected predictions due to complex anatomy. This study introduces a three-stage, topology-preserving segmentation pipeline to address this issue. The first stage performs initial d...

Deshan Kalupahana, Sonit Singh, Praveen Ravindran et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally

A targeted adversarial perturbation can drive a vision-language model's (VLM's) teacher-forced training loss for a fixed target caption to near zero, yet the same model, allowed to generate freely, produces the original, correct description with no trace of the target. We call this dissociation the train/inference gap,...

Arun Josephraj Arokiaraj, Ze-Kun Wu, A. Koshiyama · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ForestQuery: Boundary-Aware and Spatially Anchored Query Learning for Unified Forest Point Cloud Segmentation

Forest point cloud segmentation is fundamental for fine-grained 3D forest scene understanding, yet remains challenging due to irregular tree structures, severe occlusions, density variations, and ambiguous instance boundaries. Recent query-based forest segmentation methods have shown promise for unified semantic and in...

Zhihao Zhan, Le Tao, Yifei Tian et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Consecutive Posterior Fusion for Diffusive Recovery of Unobservable Image Structures

Solving severely ill-posed imaging inverse problems requires recovering image structures that are unobservable or weakly constrained by the measurements. Diffusion models provide expressive learned priors for inferring such missing information, while posterior sampling incorporates measurement consistency along the rev...

Elena Morotti, Davide Evangelista, Elena Loli Piccolomini · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Uncertainty as a Proxy for Semantic Correctness in Diffusion-Based Medical Image Synthesis

Diffusion models can synthesise contrast-enhanced CT (CECT) from non-contrast CT (NCCT), avoiding contrast administration and its environmental and patient-access costs. However, visually realistic images are not necessarily anatomically correct, and the pixel-intensity and feature-space similarity metrics used to asse...

Yuxuan Ou, Konstantinos Kamnitsas, OxAAA Study et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Evolving Hybrid Quantum-Classical Architectures for Image Classification

Hybrid quantum classical neural networks integrate parameterized quantum circuits (PQCs) with established deep learning architectures, but their performance depends strongly on the choice of quantum circuit architecture, a choice that remains largely manual. Most existing approaches rely on hand-designed or fixed circu...

Devroop Kar, Daniel Krutz, Travis Desell · 0 citations
#artificial intelligence Preprint Oct 2026

Contextual Flow Matching: Adaptive Step Selection in Flow Models for Efficient Visual Generation

Flow Matching enables high-quality visual generation via continuous-time dynamics, but inference remains costly due to multiple sequential function evaluations. Existing acceleration methods reduce the number of function evaluations but often introduce additional training overhead, degrade quality, or fail to account f...

D. J. Bajpai, Arun Verma, M. Hanawal · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Foresight: planning future perception in streaming VLMs without retraining

Existing streaming vision-language models (VLMs) continuously perceive and reason over visual streams, but their computational pathways remain fixed throughout inference. Consequently, they cannot adapt computation to evolving scene dynamics, where different future events demand different levels and forms of perception...

Ashok Prasad Neupane, Dipan Bartaula, Ankit Belbase et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.