Jul 2026· ACM Transactions on Graphics· Vol 45, pp. 1 - 13· 0 citations· 67 references
Computer Science
TL;DR
STyMo is presented, a few-shot approach that learns motion style from only seconds of paired data and trains in one to two minutes, to decompose style into two components: a static component capturing time-invariant posture, and a temporal component capturing frame-wise dynamics.
Abstract
Supporting a wide variety of motion styles is critical for creating diverse virtual characters, but current methods either require large stylized datasets or pre-trained models that cannot generalize beyond their training distribution. We present STyMo, a few-shot approach that learns motion style from only seconds of paired data and trains in one to two minutes. Our key insight is to decompose style into two components: a static component capturing time-invariant posture, and a temporal component capturing frame-wise dynamics. This decomposition yields an interpretable system where posture intensity, temporal exaggeration, and per-body-region style can be adjusted at runtime. Furthermore, the reduction in required training data and computation time structurally permits an iterative authoring workflow. To ensure robustness on arbitrary inputs, we further introduce a stylizability gate that automatically prevents artifacts on out-of-distribution motions. We demonstrate results across diverse motion styles, from subtle emotional variations to exaggerated character archetypes, and release our processed paired dataset to facilitate future research. The source code used in this paper can be found at: https://github.com/facebookresearch/STyMo
This work proposes Directable Motion Paraphrasing (DMP), a novel motion retargeting framework based on the concept of motion paraphrasing, analogous to text paraphrasing, where the core semantics of a motion are preserved while allowing expressive, user-directed variations.
Sunmin Lee, Davis Rempe, Yifeng Jiang et al.· ACM Transactions on Graphics· 0 citations
Transferring animations between characters with diverse skeletal structures is challenging. Traditional retargeting pipelines rely on fixed correspondences, canonical skeletons, or human-centric datasets, which can lead to artifacts when applied across heterogeneous morphologies. We introduce a framework for cross-morphology motion transfer with semantic style alignment that uses morphology-agnostic control signals (e.g., velocity, angular velocity, relative height) to align behaviors across species. Our method supports all-to-all retargeting: motions from any source can be mapped to any trained target while preserving target-specific style. For each target morphology, we train a Vector Quantized VAE and an autoregressive sequence model to construct a compact, morphology-specific codebook that captures stylistic priors. This modular design scales to new morphologies without retraining existing models and allows optional user control (e.g., phase, velocity scaling) for fine-grained alignment. Experiments across bipeds and quadrupeds demonstrate accurate, plausible, and style-faithful motion transfer, establishing a scalable approach to retargeting across arbitrary skeletal topologies.
Alexios Mylordos, J. L. Pontón, Nuria Pelechano et al.· IEEE Transactions on Visuali...· 0 citations
It is demonstrated that MoSAIC improves the response--preservation trade-off required for selective and controllable part-local motion editing, and is presented as a latent diffusion framework for part-local reference-conditioned motion style transfer.
Video-text models adapted from image-text architectures (e.g., CLIP) frequently exhibit temporal blindness, the inability to perceive fundamental cues like order, direction, and motion dynamics. Standard datasets mask this limitation by enabling models to exploit static spatial shortcuts. To systematically evaluate this, we introduce XTE-Bench, a diagnostic probe revealing that even large-scale video-language models struggle with basic temporal reasoning, indicating that parameter scaling alone is insufficient to resolve this flaw. To address this, we propose Cross-Modal Temporal Edits (XTE), a self-supervised framework that injects precise temporal supervision. By performing synchronized video-text transformations, XTE generates hard temporal negatives without manual annotation. We instantiate this with ViTAL-X, a lightweight model that equips frozen image-text backbones with temporal awareness while preserving their foundational spatial knowledge. Across six temporal benchmarks, ViTAL-X achieves state-of-the-art performance. Utilizing only 0.4B parameters and 1M training clips, ViTAL-X outperforms 7B-parameter models and surpasses baselines trained on 600x more data. These results demonstrate that targeted, high-quality temporal alignment provides a highly efficient alternative to pure scaling.
V. SethuramanT., Savya Khosla, O. Susladkar et al.· 0 citations
Motion warping is a core technique in character animation that enables the adaptation of existing motion data to novel spatio-temporal constraints. Conventional motion warping methods often rely on heuristic modifications that can violate physical consistency or introduce visual artifacts. More recent learning-based editing approaches improve realism, but many of them encode motion into tightly entangled latent space, which makes them struggle to balance editing flexibility and content preservation. To address this, we propose a novel deep motion warping framework that explicitly disentangles the motion structure from global and stylistic attributes for intuitive motion editing. Our key insight is to leverage learned phase features as a continuous and robust representation of the underlying structure, and explicitly disentangle motion into root velocity, phase, and learned latent variables using a phase-conditioned diffusion autoencoder. This design supports a wide range of editing operations, including root motion warping, motion exaggeration, time warping, and style transfer by directly manipulating decoupled components, without requiring paired training data. Extensive experiments demonstrate that our approach enables high-level, flexible motion editing while strictly preserving the structural consistency and physical plausibility of the source motion
Bowen Zheng, Linjun Wu, Xinwei Jiang et al.· International Conference on...· 0 citations
This work builds a bidirectional control flow between the style and the content, also adjusting the style towards the content, in which case the style-content collision is alleviated and the dynamics of the style is better preserved in the integration.
Zhe Li, Yisheng He, Lei Zhong et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.