Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Open access Sep 2026

MeanFlowAdvantage: Stable Reward Fine-Tuning for Few-Step Average-Velocity Generators

MeanFlow enables efficient few-step generation by predicting interval-average velocities, but this representation creates a mismatch for reward fine-tuning: existing advantage-based objectives are typically defined on instantaneous velocities or equivalent $x_0$-space predictions, whereas inference directly uses the le...

Haocheng Tang, Tianchi Xie, Xingqiao Lin · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Emergent Specialization in Populations of Self-Supervised Collaborative Vision Experts Without a Shared Gate or Cross-Agent Gradients

Can a population of neural networks develop a useful division of labor without a shared gate or gradients between agents? We study a setting where each network has its own weights, trains independently on the same heterogeneous data, and can ask another agent for help through a forward pass. Unlike mixtures of experts,...

Aram Davtyan, Pablo Acuaviva, Sebastian Stapf et al. · 0 citations
#computer vision Preprint Jun 2026

Fusing Complementary Multi-view Features for Screen-Based Eye Tracking

Current multi-view gaze estimation remains limited by existing datasets, insufficient exploitation of complementary cross-view information, and evaluation focused primarily on average gaze error. We address these limitations through a more systematic study of multi-view gaze estimation. First, we introduce PrismGaze, a...

Chang Liu, Jia-Qi Liu, Cheng-Wen Zhang et al. · 1 citation
#computer vision May 2026

Low Latency Gaze Tracking via Latent Optical Sensing

The proposed system establishes a new operating point in the accuracy-latency-compute trade-off for latency- and resource-constrained gaze tracking, and highlights the potential of task-driven optical sensing for ultra-low-latency, computationally efficient human-computer interaction systems.

Yidan Zheng, Matheus Souza, Kaizhang Kang et al. · 0 citations
#computer vision Preprint Nov 2025

MILE: A Mechanically Isomorphic Hand Exoskeleton and Visuotactile Robotic Hand for Data Collection in Dexterous Manipulation

MILE, a teleoperation-based data-collection system comprising the wearable MILE exoskeleton and the mechanically corresponding MILE-Tac robotic hand, and trained paired ACT and DP policies with and without tactile input on MILE-collected demonstrations for downstream imitation learning.

Jinda Du, Jie-Ji Ren, Qiao-Jun Yu et al. · 6 citations
#computer vision Nov 2025

HAGI++: Head-Assisted Gaze Imputation and Generation

HAGI++ is presented, a multi-modal diffusion-based imputation method that, for the first time, leverages integrated head-orientation sensors to exploit the natural correlation between head and eye movements and enables more complete, accurate eye-gaze recordings in real-world contexts, enhancing gaze-based analysis and...

Chu-Han Jiao, Zhi-Ming Hu, Andreas Bulling · 1 citation
#computer vision Preprint Sep 2026

Reliability-Gated Fusion of Consumer Head and Foot IMUs for Lower-Body 3D Pose

The channel-gated model is the most accurate of the authors' learned fusion arms on clean data and its gates suppress the natively biased foot-orientation channels on clean real data without test-time supervision and flag dropout bursts at 0.92-0.999 AUROC.

Zhi-Lin Guo, Bo-Qiao Zhang, O. Urbán et al. · 0 citations
#computer vision Preprint Sep 2026

One Sensor, Whole Body - 3D Body Pose from a Single Consumer Earbud IMU

A multimodal capture pipeline is built that records four-view RGB-D video together with an AirPods head IMU and two Striv insole IMUs, synchronize the streams post-hoc, and generate pseudo-ground-truth with SAM 3D Body, yielding a 35-take single-subject benchmark spanning gait, turning, vertical, everyday, and clinical...

Zhi-Lin Guo, Bo-Qiao Zhang, O. Urbán et al. · 1 citation
#computer vision Preprint Open access Sep 2026

Toward a Culturally Adapted Chinese Language Agent: A Wizard-of-Oz Study of Nonverbal Behavior in Chinese-German Intercultural Interaction

Successful intercultural communication requires more than grammatical competence. It demands sensitivity to culturally embedded social norms whose violation triggers subtle but meaningful nonverbal responses. For German learners of Mandarin Chinese, acquiring this sensitivity is critical yet poorly supported by existin...

Siddhant Jain, Anna Lea Reinwarth, Dimitra Tsovaltzi et al. · 0 citations
#computer vision Preprint Open access Sep 2026

SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening

The rapid growth of online video platforms and AI-generated content has made reliable video guardrails a key challenge for safety and real-world deployment. While most videos can be screened through fast pattern recognition, a small subset requires deeper reasoning over temporally complex content and nuanced policy con...

Shahriar Kabir Nahin, Hadi Askari, Muhao Chen et al. · 0 citations
#computer vision Preprint Jan 2026

PEST: Parameter Efficient Steering of Blackbox VLMs via Agentic Few-shot Alignment for Hateful Meme Moderation

In this work, we examine hateful memes from three complementary angles - how to detect them, how to explain their content and how to intervene them before being posted - by applying a range of strategies built on top of generative AI models. To the best of our knowledge, explanation and intervention have typically been...

Naquee Rizwan, Subhankar Swain, Paramananda Bhaskar et al. · 4 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.