Skip to content
Preprint

MIME: Multimodal Interactive Motion Encoder

Jul 2026 · 0 citations · 54 references
Computer Science

TL;DR

The Multimodal Interactive Motion Encoder (MIME) is introduced, which, to the authors' knowledge, represents the first dedicated multimodal encoder designed specifically for two person interactive motion.

Abstract

Text-motion representation learning has advanced rapidly, with growing interest in multi person interactions for animation, AR/VR, and embodied AI. These settings require representations that align language with both individual actor dynamics and the relationships between actors. We introduce the Multimodal Interactive Motion Encoder (MIME), which, to our knowledge, represents the first dedicated multimodal encoder designed specifically for two person interactive motion. MIME captures individual and shared structure using stream based co-attention with explicit interaction features and curriculum based contrastive training. On Inter-X text-motion retrieval, MIME consistently outperforms early and late fusion baselines across gallery sizes, achieving a 12.8% relative improvement in text-to-motion R@1 at a 2,000-sample gallery. We further evaluate MIME as a frozen auxiliary prior within TIMotion and InterMask on the unseen InterHuman dataset. MIME improves semantic alignment metrics while maintaining comparable FID in TIMotion. These results show that interaction aware multimodal encoding improves multi person motion retrieval and transfers across datasets to support downstream motion generation.

View source

Similar papers

Jul 2026

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment

Recent advances in video diffusion models have spurred interest in human-object interaction (HOI) video generation, which demands fine-grained control over interaction logic beyond single-subject animation. However, existing HOI methods rely heavily on explicit motion control, limiting scalability and generalization across diverse objects and interactions. In this study, we propose AgentHOI, a text-driven HOI video generation following a thinking-before-generation framework that bridges the gap between high-level textual intent and physical execution through multi-agent reasoning over perception, interaction, and motion planning. Building upon the generated interaction plans, we further strengthen text-driven motion understanding. We introduce an implicit text-motion alignment strategy that distills text-to-motion priors into the video diffusion model, enabling robust HOI synthesis without explicit motion inputs at inference. Experiments show that AgentHOI significantly improves interaction naturalness, object appearance preservation, and adherence to complex textual instructions across challenging object-centric scenarios such as wearing and riding. The code is available at https://github.com/bone-11/agenthoi.

Ziyao Huang, Shunkai Li, Juan Cao et al. · 1 citation
Preprint Aug 2026

COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models

COMET is a temporally grounded framework that systematically strengthens video MLLMs through explicit temporal representation, appearance-motion fusion, and direction-aware optimization and achieves consistent overall improvements with a pronounced motion-temporal bias.

Chenghua Zhu, Zhaolu Kang, Qifan Shi et al. · 0 citations
Conference Aug 2026

ActionLMM: captioning long-video actions with memory-augmented VLMs

This work introduces ActionLMM, a memory-augmented vision-language model for long-video action summarization that aligns visual and motion modalities through joint representation learning and leverages a novel dual-memory mechanism to retain both local motion details and global temporal structure.

Rui-Rui Li, Dari Abdullah Alrwoaily, Turgut Sofuyev et al. · 0 citations
Jul 2026

PaMG: adaptive part-based motion generation and editing from text

PaMG is introduced, a framework that leverages the local features of body part motions, achieving better performance in both global and local text-driven human motion generation and editing and validating the effectiveness of the approach.

Xin Guo, Yifan Zhao, Jia Li · 0 citations
Preprint Aug 2026

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation

This work introduces CamChoreo, a benchmark of 4,229 real single-shot clips with expert-annotated temporal segments, and proposes CamDistill, which distills the same geometric knowledge into lightweight camera tokens during training and removes the 3D model at inference.

Dazhao Du, Shiyan Du, Jian Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.