Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Sep 2026

From Scores to Samples: Elastic Forcing for Autoregressive Video Generation

This framework minimizes maximum mean discrepancy in frozen self-supervised video representation spaces in frozen self-supervised video representation spaces, using a hybrid Nystr\"om--Monte Carlo estimator to balance approximation bias and sampling variance.

Chi Zhang, Yue-Yi Liu, Hao-Yan Shi et al. · 1 citation
#artificial intelligence Preprint Open access Sep 2026

Spectral Super-Resolution using Spatial-Spectral Residual Operator Networks

Spectral super-resolution of multispectral satellite images can enable high temporal- and spatial-resolution hyperspectral satellite imagery at a modest cost, significantly increasing the applicability of hyperspectral remote sensing. This task is inherently ill-posed, making it well-suited for deep learning-based meth...

Seokhyun Chin · 0 citations
#artificial intelligence Preprint Sep 2026

eval-unlearn: Benchmarking unlearning in Text-to-Image Diffusion Models

Eval-unlearn is an open-source Python library providing a unified, reproducible benchmarking framework for concept unlearning in T2I Diffusion models, integrating twelve published unlearning techniques spanning fine-tuning, closed-form model editing, and inference-time intervention.

Mansi, Nikhil Raghavan, Zi-Xia Huang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Spatial Grafting: Grounding 3D Features for Flow-Matching Robot Policies

Spatial Grafting is proposed, a versatile, lightweight spatial module that binds frozen reconstruction features to metric, robot-relative geometry and injects them into the flow-matching action expert through cross-attention without modifying the host's perceptual pathway, so the host retains the full benefit of its pr...

Ding-Sheng Liu, Yang-Zheng Wu, Mahboubeh Asadi et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Generative AI-Based Data Augmentation for Oral Lesion Classification: The PhotoMOCI Dataset and Benchmark

Early detection of oral cancer via photographic imaging presents a promising avenue for large-scale oral cavity screening. However, the development of robust deep learning models is frequently hampered by the scarcity of high-quality, annotated datasets. To address this limitation, a novel and well-curated resource, th...

Marco Parola, Mario G. C. A. Cimino, Sabrina Senatore · 0 citations
#artificial intelligence Preprint Sep 2026

ReCAT: Remember, Count, and Time: Structured Recurrent Memory for Robot Manipulation

Comparisons within ReCAT show that the observation encoder and every-block memory conditioning are needed for this performance, and that update rules developed for efficient sequence modeling behave differently as robot memory: additive updates have the highest observed success on counting and timing, and delta-rule up...

Pankhuri Vanjani, M. Hatab, Can Mizrakli et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

CarveMix-RC: Addressing Rare-Class Imbalance Through Lesion-Aware Synthetic Augmentation for Brain Metastasis Segmentation

Accurate segmentation of post-treatment brain metastases is essential for treatment planning, longitudinal disease monitoring, and quantitative assessment of therapeutic response. The BraTS-MET 2026 Task 1 challenge introduces a clinically relevant segmentation problem involving four anatomically distinct tumor subregi...

Md Shibly Sadique, Md Fayaz Bin Hossen, Michael L. Evans et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut

AI agents increasingly carry out long-horizon professional work, but their evaluations rarely require a finished creative deliverable. To this end, we introduce Timeline-Bench, a benchmark of 56 real video-editing tasks, each asking an agent to turn raw production material into a finished video. Tasks range from select...

Gunin Gupta, Nirmit Arora, Pavan Tankala · 0 citations
#artificial intelligence Preprint Sep 2026

OPIS: An Input-Grounded Benchmark for Multi-Object Memory in Video World Models

Video world models must preserve the visual state of the world over time, but existing evaluation protocols often rely on generated histories, video reference, or selected revisit viewpoints that can confound the assessment of a model's true memory capability. To address this, we introduce OPIS, an input-grounded bench...

Hao Wang, Tao Yu, Liu-Zhou Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Still There, No Longer Seen: Exposing Compression-Induced Risk in Large Vision-Language Models

Visual token compression reduces the inference cost of Large Vision-Language Models (LVLMs). However, aggregate robustness measures do not reveal whether a particular adversarial failure is induced by compression or inherited from the underlying model. We define a compression-specific failure (CSF) as an adversarial in...

Qian-Kun Li, Yuechen Zhang, Bo-Wen Chen et al. · 0 citations
#artificial intelligence Open access Sep 2026

SPIDER: Multi-Layer Semantic Token Pruning and Adaptive Sub-Layer Skipping in Multimodal Large Language Models.

Multimodal Large Language Models face significant efficiency challenges that stem from two distinct yet coupled sources: data redundancy and computational redundancy. While most methods focus on data redundancy by pruning visual tokens from the output of the visual encoder or computing redundancy in LLM decoders using...

Tian-Xiang Chen, Zhentao Tan, Zi Ye et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.