Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Sep 2026

Exo2EgoHOI: Hand-Object-Interaction Aware Exocentric-to-Egocentric Video Generation

Egocentric videos of human manipulation provide valuable visual experience for embodied intelligence, yet collecting such data at scale is costly. Exocentric-to-egocentric video generation offers a scalable alternative by transforming abundant third-person manipulation videos into first-person observations. However, ex...

Hong-Jia Zhai, Xi-Yu Zhang, Hao-Ran Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

After a Decade: Bringing Shadow Removal into the Real World with Agentic Training Data

Shadow removal looks nearly solved on established benchmarks, yet remains brittle in the real world. Models have advanced; the paired training data they rely on have barely changed in nearly a decade. The reason is simple: obtaining a shadow-free target requires removing the occluder while keeping the scene, camera, an...

Shilin Hu, Jingyi Xu, Dimitris Samaras et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Retargeting Motions to Diverse Skeletons via Learnable Flattening

Cross-structural motion retargeting aims to transfer motion between different skeletal topologies. Despite recent progress, existing state-of-the-art models struggle with reliability in zero-shot settings, i.e. skeletons with different topologies which were unseen during training, and recent Transformer-based attempts...

Kia-Jung Yang, Fabian H. Sinz, Paweł A. Pierzchlewicz · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Caption-Mediated Perceived-Safety Estimation for Pedestrian Routing

This paper presents an explainable approach to pedestrian routing, in which perceived safety is estimated from street-level imagery through an explicit natural-language intermediate representation. A vision--language model caption is generated and stored before any scoring is undertaken, and the perceived-risk class is...

Simon Parkinson, Paloma Liu, Wei Zheng et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Does Gradient Conflict Predict the Understanding--Generation Trade-off? A Controlled Audit of Conflict-Metric Validity in Unified Multimodal Models

Unified multimodal models (UMMs) are increasingly designed around gradient conflict between understanding and generation objectives. The premise that reducing these metrics improves the downstream understanding-generation trade-off has never been tested directly. We audit it in a controlled testbed, GRIDUMM, which mirr...

Shuyang Jiang, Fucheng Deng, Yuchuan Luo et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Colorectal Cancer Segmentation with Adaptive Augmentation and Multi-Resolution Ensemble Models

Colorectal cancer (CRC) is the second most deadly and third most common cancer, and the leading cause of death among gastrointestinal cancers. Early diagnosis is crucial for the treatment of this cancer and increasing the survival rates. Although CRC is more common in developed regions, its occurrence is also increasin...

\"Umit Mert \c{C}a\u{g}lar, Alptekin Temizel · 0 citations
#artificial intelligence Preprint Open access Oct 2026

GaugeVLM: Structuring Spatial Supervision with Measured Geometric Interventions

Vision-language models (VLMs) can contradict themselves across views of the same spatial relation and fail to respond when that relation changes. Addressing these failures requires supervision that captures error magnitude and geometric dependencies across observations, both of which remain implicit in training on indi...

Hongbo Wang, Zihan Lin, Wenkui Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Raw Imagery Impacting Your AI: Should You Care?

Onboard AI is gaining interest for space applications such as vessel, wildfire, and cloud detection, where real-time processing can improve mission reactivity and reduce downlink needs. However, onboard models may operate on raw or minimally processed imagery rather than on restored ground products. This study evaluate...

A. Dorise, Marjorie Bellizzi, Stéphane May · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Life-Bench: A Benchmark and Knowledge Graph Framework for Multimodal Personalization Beyond Concept Recognition

As large language models increasingly power personal assistants, users expect them to reason over multimodal life histories, from recognizing people to understanding events to aggregating patterns, yet existing benchmarks primarily target concept-level recognition. We introduce Life-Bench, a fully synthetic, human-veri...

Xia Hu, Honglei Zhuang, Brian Potetz et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?

Delegating complete application development to coding agents requires preserving the intended design rather than simply producing plausible outputs through naive prompting. Game development provides a demanding testbed, as long-form Game Design Documents (GDDs) describe requirements that must work together across game...

Seonho Lee, Wonryeol Jeong, Alberto Cereser et al. · 0 citations
#artificial intelligence Preprint Sep 2026

UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement

Modern multimodal models bring generation and understanding into a single unified system, which enables them to provide and learn from their own feedback. Motivated by this unified capacity, we introduce UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback...

Fang Wu, Dan-Lei Xing, Yan-Jie Huang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ChartRevise: A Dataset and Evaluation Protocol for Exact Chart Editing via Code

Chart editing requires cross-modal edit grounding, realizing a requested visual change in the code that draws it, with necessary related updates and without altering unrelated content. Existing benchmarks emphasize either code executability or chart quality, but their metrics do not clearly distinguish request completi...

Jia-Xiang Tang, Yi Zhou, Chad DeLuca et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.