Visual emotion recognition plays a critical role in human–computer interaction and mental health applications. Although existing Vision–Language Models (VLMs) alleviate the limitations of conventional vision models in high-level semantic understanding, they still face three main challenges: limited emotional semantic understanding, insufficient visual emotional perception capability, and high computational costs when deploying both models simultaneously. To address these issues, a large language model-assisted distillation–fusion framework (VERLADF) is proposed, which introduces emotion instruction data generated by GPT to fine-tune a VLM, thereby enhancing its emotional semantic understanding capability. Furthermore, we transfer the visual emotion discrimination knowledge of a conventional vision model into the VLM using a distillation module while keeping the VLM frozen during training, which reduces the computational costs. Following that, we design a fusion and prediction module that adaptively fuses predictions from the instruction-tuned VLM and the distillation module for final emotion recognition. The experimental results on the Abstract, ArtPhoto, Emotion6, and FI datasets demonstrate that VERLADF achieves recognition accuracies of 36.71%, 52.38%, 74.73%, and 79.69%, respectively, significantly outperforming many methods in the literature and demonstrating the effectiveness of the proposed framework.
Scaling generalist vision-language-action (VLA) policies is severely bottlenecked by the inherent heterogeneity of embodied data, which spans diverse robot morphologies, camera configurations, and low-level action spaces. Existing paradigms typically address this mismatch through explicit action retargeting, human-to-robot video synthesis, or dataset-specific adaptation branches, fundamentally hindering the joint learning of a unified policy. We introduce UCAG-P, a camera-centric unified action formulation that structurally aligns heterogeneous embodied datasets into a shared geometric action space. Rather than treating robot-specific commands as the shared policy target, UCAG-P represents manipulation through camera-observable anchor motion in image and camera-frame coordinates, treating robot arms, humanoids, and human hands as different embodiments of a common action schema. A geometry-conditioned action translator combines predicted motion with target-embodiment kinematics to produce executable controls. The resulting decoupled architecture allows a shared VLA policy to learn transferable manipulation geometry while retaining embodiment-specific controllability. UCAG-P is trained on 4.03K hours of robot and simulation data and 2.34K hours of human demonstrations. A single checkpoint reaches 98.3% on LIBERO, 88.7% and 89.2% on RoboTwin Easy and Hard, 82.0% zero-shot on LIBERO-Plus, and 62.0% on RoboCasa GR-1, without benchmark-specific fine-tuning.
Xiaomi Embodied Intelligence Team, University of Macau Shaoqing Xu, Fang Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.