Jul 2026· ACM Transactions on Multimedia Computing, Communications, and Applications (TOMCCAP)· 0 citations· 65 references
TL;DR
This work proposes POP-TMR which pretrains body part representations for fine-grained text-motion retrieval and introduces HumanML3D+, an enhanced benchmark that provides accurate positive annotations for text queries and includes text descriptions at varying levels of detail, enabling more systematic performance assessment.
Abstract
Text-motion retrieval has gained increasing research attention, yet several critical challenges remain such as data scarcity, limited fine-grained matching capabilities, and inadequate evaluation protocols. To address these issues, we propose POP-TMR which pretrains body part representations for fine-grained text-motion retrieval. Our approach leverages large-scale human motion datasets to pretrain a spatio-temporal transformer-based motion encoder, enabling more generalizable motion features. In addition to matching global motion and text representations, we propose a local branch to capture detailed body part features for enhancing spatial-aware cross-modal alignment. To improve evaluation, we introduce HumanML3D+, an enhanced benchmark that provides accurate positive annotations for text queries and includes text descriptions at varying levels of detail, enabling more systematic performance assessment. Extensive experiments on KIT-ML, HumanML3D and HumanML3D+ benchmarks demonstrate that POP-TMR outperforms state-of-the-art methods. Furthermore, we showcase its effectiveness in additional downstream applications, including text-to-motion generation evaluation, human interaction recognition and zero-shot moment retrieval. Data, code and pretrained model are publicly available at https://lin-kayla.github.io/POP_TMR/.
This work proposes a lightweight granularity-aware model anchored at a frozen standard-caption-aligned retrieval model that improves mixed-granularity retrieval without compromising standard-caption performance, and believes that its MRBench provides a comprehensive testbed for advancing motion-language alignment evaluation.
Fulong Liu, Liang Xu, Chengqun Yang et al.· 0 citations
PaMG is introduced, a framework that leverages the local features of body part motions, achieving better performance in both global and local text-driven human motion generation and editing and validating the effectiveness of the approach.
Xin Guo, Yifan Zhao, Jia Li· Science China Information Sc...· 0 citations
This work proposes FineMoLA, a weakly supervised framework that learns fine-grained frame--phrase correspondence directly from clip-level annotations, and efficiently infers pseudo frame-level alignments without human labeling.
Tongyan Wang, Zhengyuan Li, Muhan Lin et al.· 0 citations
This work introduces ActionLMM, a memory-augmented vision-language model for long-video action summarization that aligns visual and motion modalities through joint representation learning and leverages a novel dual-memory mechanism to retain both local motion details and global temporal structure.
Rui-Rui Li, Dari Abdullah Alrwoaily, Turgut Sofuyev et al.· International Conference on...· 0 citations
This work proposes Video-GAR, a novel framework that reframes the conventional retrieval task from superficial matching to generative understanding, positing that the capability for query reconstruction evidences deep semantic comprehension.
Mingjin Kuai, Qianyin Xiao, Juncheng Li et al.· Annual International ACM SIG...· 0 citations
This work proposes Self-SiMS, a self-similarity-based Moment Proposal and Scoring that exploits intrinsic relationships within videos, enabling robust span generation and scoring and introduces a query-aware MLLM-based reasoning stage to further sharpen alignment between text and video.
Jihyun Lee, Cheol-Ho Cho, Woojin Jun et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.