Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Sep 2026

EviRover: Reinforcing Agentic Perception Beyond a Glance

Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge on fine-grained visual details or requi...

Kai-Xuan Fan, Kai-Tuo Feng, Tian-Shuo Peng et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models

World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action prediction, as redu...

Qi Lyu, Jiahua Dong, Hao Shen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories

Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spanning months or years, reprocessing raw clips for every query becomes computationally prohibitive. Memory systems offer a scalable alternative by compacting videos into text...

Guang-Zhi Xiong, Xin-Yuan Zhang, Xiao Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

GateSPINE: Gated Cross-View Fusion for Lumbar Spine MRI Report Generation

Automated report generation can ease the burden radiolo gists face when interpreting multi-sequence MRI studies. Unlike CT, MRI examinations comprise multiple sequences and imaging planes, each con tributing complementary diagnostic information. Existing methods en code a study as a single volume and combine multiple a...

Van Nguyen Hoang, Cuong Vuong Tuan, Trang Mai Xuan et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LongEmo: Towards Emotion Understanding and Reasoning in Long Videos

While recent Multimodal Large Language Models (MLLMs) have shown promise in affective computing, their reasoning capabilities are largely confined to short video clips with limited interactions. However, real-world emotions are not merely isolated instantaneous reactions but dynamic and cumulative processes deeply shap...

Shuo Zhang, Yifan Zhou, Han-Yu Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Less Data, Better Timing: Student-Curriculum Coupling for VLM On-Policy Distillation in Temporal Video Grounding

On-policy distillation (OPD) provides dense supervision directly on student-generated trajectories, making it an effective post-training strategy for vision-language models in temporal video grounding (TVG). However, existing pipelines typically construct the training curriculum from a fixed teacher and the initial stu...

Jiacheng Qiu, Yunsoo Kim, Ruichen Xu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

WARP: A Unified Benchmark for Invisible Image Watermarking -- Robustness and Protection Against Attacks

Digital image watermarking is increasingly critical in media contexts, as emerging regulations and industry practices require marking AI-generated content and ensuring traceable sources to prevent manipulation or misuse. Recent advances in invisible watermarking methods highlight the need to update existing benchmarkin...

Khaled Abud, A. Yakushev, Aleksandr Akimenkov et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception

Hour-scale audio-visual question answering is constrained by a context dilemma: dense whole-recording encoding rapidly exhausts context limits, whereas uniform temporal compression severely dilutes fine-grained acoustic and visual evidence. We introduce LEAP, a framework where the model retrieves its own evidence witho...

Ju-Yi Lin, Zhi-Qiang Lao, Jia-Li Cui et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding

Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making long-video understanding expensive for the vision encoder and the language model. Existing methods often compress visual tokens after dense encoding, creating a mism...

Yulong Liu, Xiao-Tian Han, Jun-Yuan Shang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Let the Carrier Carry the Attack: Preserving the Subject in Adversarial Image Generation

Strong unrestricted adversarial attacks can distort the primary object of an image, hereafter referred to as the subject. To preserve subject integrity without compromising attack magnitude, we introduce the carrier: a secondary visual element that provides an auxiliary region to facilitate the attack under global clas...

Linfeng Jiang, Steven McDonagh, Yuhang Chen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

When Masking Helps or Hurts Robustness in Compressed CLIP: A Pre-Deployment Diagnostic

This paper demonstrate that whether masking-based token pruning helps or hurts worst-group robustness can be predicted before deployment, without labels or fine-tuning. A systematic study of semantic masking across 8 spurious-correlation benchmarks shows its effect on worst-group accuracy is highly unstable: it improve...

Muhammad Zawish, Steven Davy · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ShieldCLIP: Selective Safety Alignment for Harmful Content Mitigation in Multimodal Foundation Models

Multimodal encoders such as CLIP underlie many downstream systems, but their web-scale training data embed harmful associations that safety alignment must suppress without unnecessarily changing benign representations. Because ethical and practical constraints prevent collecting real unsafe content at scale, existing d...

Tobia Poppi, Silvia Cappelletti, Samuele Poppi et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.