The Multimodal Interactive Motion Encoder (MIME) is introduced, which, to the authors' knowledge, represents the first dedicated multimodal encoder designed specifically for two person interactive motion.
Abstract
Text-motion representation learning has advanced rapidly, with growing interest in multi person interactions for animation, AR/VR, and embodied AI. These settings require representations that align language with both individual actor dynamics and the relationships between actors. We introduce the Multimodal Interactive Motion Encoder (MIME), which, to our knowledge, represents the first dedicated multimodal encoder designed specifically for two person interactive motion. MIME captures individual and shared structure using stream based co-attention with explicit interaction features and curriculum based contrastive training. On Inter-X text-motion retrieval, MIME consistently outperforms early and late fusion baselines across gallery sizes, achieving a 12.8% relative improvement in text-to-motion R@1 at a 2,000-sample gallery. We further evaluate MIME as a frozen auxiliary prior within TIMotion and InterMask on the unseen InterHuman dataset. MIME improves semantic alignment metrics while maintaining comparable FID in TIMotion. These results show that interaction aware multimodal encoding improves multi person motion retrieval and transfers across datasets to support downstream motion generation.
Recent advances in video diffusion models have spurred interest in human-object interaction (HOI) video generation, which demands fine-grained control over interaction logic beyond single-subject animation. However, existing HOI methods rely heavily on explicit motion control, limiting scalability and generalization across diverse objects and interactions. In this study, we propose AgentHOI, a text-driven HOI video generation following a thinking-before-generation framework that bridges the gap between high-level textual intent and physical execution through multi-agent reasoning over perception, interaction, and motion planning. Building upon the generated interaction plans, we further strengthen text-driven motion understanding. We introduce an implicit text-motion alignment strategy that distills text-to-motion priors into the video diffusion model, enabling robust HOI synthesis without explicit motion inputs at inference. Experiments show that AgentHOI significantly improves interaction naturalness, object appearance preservation, and adherence to complex textual instructions across challenging object-centric scenarios such as wearing and riding. The code is available at https://github.com/bone-11/agenthoi.
Ziyao Huang, Shunkai Li, Juan Cao et al.· arXiv.org· 1 citation
MaP provides a simple and effective solution for enhancing motion-centric video reasoning without model training or architectural modification, and consistently improves average motion-reasoning accuracy.
Xikai Sun, Kebin Liu, Haotian Wang et al.· 0 citations
COMET is a temporally grounded framework that systematically strengthens video MLLMs through explicit temporal representation, appearance-motion fusion, and direction-aware optimization and achieves consistent overall improvements with a pronounced motion-temporal bias.
Chenghua Zhu, Zhaolu Kang, Qifan Shi et al.· 0 citations
This work introduces ActionLMM, a memory-augmented vision-language model for long-video action summarization that aligns visual and motion modalities through joint representation learning and leverages a novel dual-memory mechanism to retain both local motion details and global temporal structure.
Rui-Rui Li, Dari Abdullah Alrwoaily, Turgut Sofuyev et al.· International Conference on...· 0 citations
PaMG is introduced, a framework that leverages the local features of body part motions, achieving better performance in both global and local text-driven human motion generation and editing and validating the effectiveness of the approach.
Xin Guo, Yifan Zhao, Jia Li· Science China Information Sc...· 0 citations
This work introduces CamChoreo, a benchmark of 4,229 real single-shot clips with expert-annotated temporal segments, and proposes CamDistill, which distills the same geometric knowledge into lightweight camera tokens during training and removes the 3D model at inference.
Dazhao Du, Shiyan Du, Jian Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.