MeanFlow enables efficient few-step generation by predicting interval-average velocities, but this representation creates a mismatch for reward fine-tuning: existing advantage-based objectives are typically defined on instantaneous velocities or equivalent $x_0$-space predictions, whereas inference directly uses the le...
Can a population of neural networks develop a useful division of labor without a shared gate or gradients between agents? We study a setting where each network has its own weights, trains independently on the same heterogeneous data, and can ask another agent for help through a forward pass. Unlike mixtures of experts,...
Aram Davtyan, Pablo Acuaviva, Sebastian Stapf et al.· 0 citations
Current multi-view gaze estimation remains limited by existing datasets, insufficient exploitation of complementary cross-view information, and evaluation focused primarily on average gaze error. We address these limitations through a more systematic study of multi-view gaze estimation. First, we introduce PrismGaze, a...
Chang Liu, Jia-Qi Liu, Cheng-Wen Zhang et al.· 1 citation
The proposed system establishes a new operating point in the accuracy-latency-compute trade-off for latency- and resource-constrained gaze tracking, and highlights the potential of task-driven optical sensing for ultra-low-latency, computationally efficient human-computer interaction systems.
Yidan Zheng, Matheus Souza, Kaizhang Kang et al.· arXiv.org· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
MILE, a teleoperation-based data-collection system comprising the wearable MILE exoskeleton and the mechanically corresponding MILE-Tac robotic hand, and trained paired ACT and DP policies with and without tactile input on MILE-collected demonstrations for downstream imitation learning.
Jinda Du, Jie-Ji Ren, Qiao-Jun Yu et al.· 6 citations
HAGI++ is presented, a multi-modal diffusion-based imputation method that, for the first time, leverages integrated head-orientation sensors to exploit the natural correlation between head and eye movements and enables more complete, accurate eye-gaze recordings in real-world contexts, enhancing gaze-based analysis and...
Chu-Han Jiao, Zhi-Ming Hu, Andreas Bulling· IEEE Transactions on Visuali...· 1 citation
The channel-gated model is the most accurate of the authors' learned fusion arms on clean data and its gates suppress the natively biased foot-orientation channels on clean real data without test-time supervision and flag dropout bursts at 0.92-0.999 AUROC.
Zhi-Lin Guo, Bo-Qiao Zhang, O. Urbán et al.· 0 citations
A multimodal capture pipeline is built that records four-view RGB-D video together with an AirPods head IMU and two Striv insole IMUs, synchronize the streams post-hoc, and generate pseudo-ground-truth with SAM 3D Body, yielding a 35-take single-subject benchmark spanning gait, turning, vertical, everyday, and clinical...
Zhi-Lin Guo, Bo-Qiao Zhang, O. Urbán et al.· 1 citation
The design, integration, and field deployment of an AI-assisted collaborative inspection cell at the Silverline kitchen-appliance factory is presented, developed within the AI-PRISM project.
Asya Ünal, Amr Okasha, Ege Çirakman et al.· 0 citations
Successful intercultural communication requires more than grammatical competence. It demands sensitivity to culturally embedded social norms whose violation triggers subtle but meaningful nonverbal responses. For German learners of Mandarin Chinese, acquiring this sensitivity is critical yet poorly supported by existin...
Siddhant Jain, Anna Lea Reinwarth, Dimitra Tsovaltzi et al.· 0 citations
The rapid growth of online video platforms and AI-generated content has made reliable video guardrails a key challenge for safety and real-world deployment. While most videos can be screened through fast pattern recognition, a small subset requires deeper reasoning over temporally complex content and nuanced policy con...
Shahriar Kabir Nahin, Hadi Askari, Muhao Chen et al.· 0 citations
In this work, we examine hateful memes from three complementary angles - how to detect them, how to explain their content and how to intervene them before being posted - by applying a range of strategies built on top of generative AI models. To the best of our knowledge, explanation and intervention have typically been...
Naquee Rizwan, Subhankar Swain, Paramananda Bhaskar et al.· 4 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.