Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Open access Oct 2026

A Framework for Egocentric and Exocentric Procedural Understanding via Temporal Segmentation and Semantic Abstraction

Long-horizon ego/exo data contains rich procedural evidence, but are redundant, noisy, and costly to process or retain. We propose a compact framework that converts continuous multimodal workplace video into a structured Procedural State Memory, implemented as a Work Environment Model (WEM). Inspired by event segmentat...

Vivek Chavan, J\"org Kr\"uger · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Spatial Lifting for Dense Prediction

We present Spatial Lifting (SL), a novel methodology for dense prediction tasks. SL operates by lifting standard inputs, such as 2D images, into a higher-dimensional space and subsequently processing them using networks designed for that higher dimension, such as a 3D U-Net. Counterintuitively, this dimensionality lift...

Mingzhi Xu, Tao Zhou, Yong Li et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

STATERA: Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets

Vision models pretrained for frame-level appearance often struggle to infer hidden physical properties from motion. We study center-of-mass (CoM) localization for opaque, asymmetric rigid bodies from short monocular videos, where surface cues and point tracking are unreliable under self-occlusion. We propose STATERA, w...

Animesh Varma · 0 citations
#artificial intelligence Preprint Oct 2026

VISTA: A Visual Harness for Reasoning in an Interactive World

We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly...

Qiu-Shi Han, Ke-Ya Hu, Lin-Lu Qiu et al. · 5 citations · ⚡1
#artificial intelligence Preprint Open access Oct 2026

VideoEvolve: Evolving Agent Harnesses for Video Temporal Grounding

Video temporal grounding aims to localize events in videos from natural-language queries. For agents built around frozen video-language models, the harness determines how queries guide temporal predictions and how those predictions are refined. Manually refining these harnesses requires diagnosing grounding failures an...

Bingjun Luo, Yuhuan Fan, Jialin Guo et al. · 0 citations
#artificial intelligence Preprint Oct 2026

CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement

Visual grounding localizes an object described by language with a bounding box. Most multimodal grounding models compress target identification, spatial reasoning, and boundary estimation into one terminal prediction. Free-form rationales make reasoning linguistically explicit but do not necessarily expose measurable,...

Dong-Wei Sun, Yu-Jie Zhang, Bo-Wen Yao et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Towards Reliable Vision-Language Models for Autonomous Driving

Vision-Language models (VLMs) are increasingly being explored in autonomous driving for tasks such as scene understanding, driving reasoning, decision-making, and end-to-end driving. As their role becomes more prominent, ensuring their robustness and reliability is increasingly important. In real-world conditions, visu...

Manasa Mariam Mammen, Priyanka Mary Mammen, Zafer Kayatas et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

AiSearch: Interactive Multi-Modal Search with VLMs

Modern retrieval systems must both be automated and interactive, allowing users to search and refine results in real time. We present AiSearch, a flexible multimodal retrieval framework that leverages the zero shot capabilities of Vision Language Models (VLMs) for natural language search over images and videos. AiSearc...

Ali Koksal, Mei Chee Leong, Vicky Sintunata et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Frozen Scenes, Shifting Winners: Configuration Fragility in Text-to-3D Evaluation

Can a text-to-3D leaderboard change when every generated scene stays fixed? We audit this question for rendered-image evaluation, where camera settings and caption wording become part of the measurement protocol. Across 300 frozen scenes from six generators, we vary eight render and caption factors for 19 alignment eva...

Anson Y. Lam, Shuqing Li, Michael R. Lyu · 0 citations
#computer vision Preprint Open access Oct 2026

Direct Translation between Sign Languages

Sign language translation has made substantial progress between sign and spoken languages, while translation across sign languages remains less explored. Translating directly between sign languages could support communication across signing communities without requiring a shared written language. A cascade of sign-to-t...

Zetian Wu, Bowen Xie, Wuyang Meng et al. · 0 citations
#computer vision Preprint Sep 2026

MCD: Causal Distillation of Multimodal In-Context Learning in Large Vision-Language Models

Large vision-language models (LVLMs) exhibit strong multimodal in-context learning (ICL) capabilities, yet this ability degrades substantially as model size decreases. Knowledge distillation offers a natural way to bridge this gap, but existing methods primarily align output distributions or hidden representations dire...

Yanshu Li, Jia-Qian Li, Can-Ran Xiao et al. · 0 citations
#computer vision Preprint Open access Oct 2026

LoopVL: Recurrent Visual Intelligence

We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through shared modules. We train LoopVL from scratch through language pre-training, multimodal train...

Zhe Qian, Ziyang Gong, Zhongxing Xu et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.