Pretraining Body Part Representations for Text-Motion Retrieval
This work proposes POP-TMR which pretrains body part representations for fine-grained text-motion retrieval and introduces HumanML3D+, an enhanced benchmark that provides accurate positive annotations for text queries and includes text descriptions at varying levels of detail, enabling more systematic performance assessment.