Skip to content

PaMG: adaptive part-based motion generation and editing from text

Jul 2026 · Science China Information Sciences · Vol 69 · 0 citations · 52 references

TL;DR

PaMG is introduced, a framework that leverages the local features of body part motions, achieving better performance in both global and local text-driven human motion generation and editing and validating the effectiveness of the approach.

View source

Similar papers

2025

PyraMotion: Attentional Pyramid-Structured Motion Integration for Co-Speech 3D Gesture Synthesis

Generating full-body human gestures encompassing face, body, hands, and global movements from audio is crucial yet challenging for virtual avatar creation. Existing systems tokenize gestures with fixed frame-count for each token, predicting tokens of single scale from the input audio. However, expressive human gestures consist of varied patterns with different frame lengths, and different body parts exhibit motion patterns of varying durations. Existing systems fail to capture motion patterns across body parts and temporal scales due to the fixed frame-count setting of their gesture tokens. Inspired by the success of the feature pyramid technique in the multi-scale visual information extraction, we propose a novel framework named PyraMotion and an adaptive multi-scale feature capturing model called Attentive Pyramidal VQ-VAE (APVQ-VAE). Objective and subjective experiments demonstrate that the PyraMotion outperforms state-of-the-art methods in terms of generating natural and expressive full-body human gestures. Extensive ablation experiments highlight that the self-adaptiveness integration through attention maps contributes to performance.

Zhizhuo Yin, Y. Tsui, Pan Hui · 6 citations · ⚡1
Open access Jul 2026

Pretraining Body Part Representations for Text-Motion Retrieval

This work proposes POP-TMR which pretrains body part representations for fine-grained text-motion retrieval and introduces HumanML3D+, an enhanced benchmark that provides accurate positive annotations for text queries and includes text descriptions at varying levels of detail, enabling more systematic performance assessment.

Kejun Lin, Shizhe Chen, Anwen Hu et al. · 0 citations
Preprint Jul 2026

MoSAIC: Aligned Intervention Supervision for Part-Local Motion Style Transfer

It is demonstrated that MoSAIC improves the response--preservation trade-off required for selective and controllable part-local motion editing, and is presented as a latent diffusion framework for part-local reference-conditioned motion style transfer.

N. Amini, Kevin Desai · 0 citations
Book Open access Jul 2026

Motion4Motion: Motion Transfer Across Subjects at Inference

This work explores the motion transfer from one video to another, which is crucial in animation for diverse characters. Previously, video motion transfer has been largely explored between human and human-like characters, enabling a lot of applications in digital creation. However, these approaches encounter a main limitation. Specifically, related technical pipelines heavily rely on a predefined human skeleton structure and accordingly require skeleton-conditional model training. On the one hand, these methods are difficult to generalize to diverse characters, such as animals from different species, while preserving their unique motion styles. On the other hand, labeled data in diverse skeletons is limited, which additionally restricts the large-scale training for the task. In this paper, we jump out of the skeleton-based motion transfer framework and propose a training-free motion transfer framework, named Motion4Motion. Motion4Motion models the motion flow of the character in a video instead of skeletons, which makes motion transfer across species easier. Extensive experimental results and novel applications show our methods outperform baselines impressively.

Ling-Hao Chen, Zixin Yin, Duomin Wang et al. · 0 citations
Aug 2026

Text-to-Motion Generation With Discrete Representations and Large Language Models.

This work investigates a simple yet effective conditional generative framework for text-to-motion generation and proposes T2M-GIT+, which employs a non-autoregressive method to generate discrete motion representations in parallel, and is therefore more efficient than T2M-GPT+ while achieving comparable results.

Jianrong Zhang, Yang Zhang, Xiaodong Cun et al. · 0 citations
Preprint Aug 2026

Spatial Temporal Synergy: Balancing Change and Invariance in Text Driven 3D Human Motion Editing

This work proposes Change and Invariance Motion Editing (CIME), a unified framework that comprehensively decouples change and invariance into spatial pose and temporal rhythm dimensions and introduces the Riemannian Non-uniform Integral Manifold Mapping module.

Shaohui Lin, Zhenwu Shi, Jingyu Gong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.