Dec 2024· arXiv.org· Vol abs/2412.09901· 0 citations· 49 references
Computer Science
TL;DR
This work builds a bidirectional control flow between the style and the content, also adjusting the style towards the content, in which case the style-content collision is alleviated and the dynamics of the style is better preserved in the integration.
Abstract
Generating motion sequences conforming to a target style while adhering to the given content prompts requires accommodating both the content and style. In existing methods, the information usually only flows from style to content, which may cause conflict between the style and content, harming the integration. Differently, in this work we build a bidirectional control flow between the style and the content, also adjusting the style towards the content, in which case the style-content collision is alleviated and the dynamics of the style is better preserved in the integration. Moreover, we extend the stylized motion generation from one modality, i.e. the style motion, to multiple modalities including texts and images through contrastive learning, leading to flexible style control on the motion generation. Extensive experiments demonstrate that our method significantly outperforms previous methods across different datasets, while also enabling multimodal signals control. The code of our method will be made publicly available.
Generating motion sequences conforming to a target style while adhering to the given content prompts requires accommodating both the content and style. In existing methods, the information usually only flows from style to content, which may cause conflict between the style and content, harming the integration. Differently, in this work we build a bidirectional control flow between the style and the content, also adjusting the style towards the content, in which case the style-content collision is alleviated and the dynamics of the style is better preserved in the integration. Moreover, we extend the stylized motion generation from one modality, i.e. the style motion, to multiple modalities including texts and images through contrastive learning, leading to flexible style control on the motion generation. To further boost the performance, we advance the motion diffusion to motion-aligned temporal latent diffusion by developing a novel motion VAE. Extensive experiments demonstrate that our method significantly outperforms previous methods across different datasets, while also enabling multimodal signals control. The code of our method will be made publicly available.
Zhe Li, Yisheng He, Lei Zhong et al.· IEEE Transactions on Image P...· 18 citations· ⚡2
This paper presents an innovative method that leverages user-specified action paths to guide the 4D scene generation that dynamically synchronizes motions in the action path domain with their corresponding contents in the time domain.
Guo-Wei Yang, Qun-Ce Xu, Zhao Wei et al.· Science China Information Sc...· 0 citations
Multi-modality motion stylization presents a solution to the challenge of generating flexible, stylized motion based on multimodal style inputs. Historically, motion stylization has grappled with the difficulty of balancing content and style, often prioritizing one at the expense of the other. This paper addresses the complex challenge of multi-modality content-style duality, achieving a sophisticated integration that both preserves and enhances the core narrative through nuanced stylistic modifications. We propose a Multi-modality Latent Diffusion Model (MM-LDM), a novel framework that leverages diffusion models under multi-modality conditions, including motion style (text-based or motion-based), motion content, and motion trajectory components. A central innovation in our approach is the introduction of a Multi-condition Denoiser, which carefully balances the preservation of primary content with the dynamic integration of style and trajectory as secondary conditions. This multi-modality guidance mechanism, implemented during the denoising process, ensures that new styles are seamlessly integrated with the original content. It gives rise to more authentic and cohesive motion stylization outcomes, establishing a new benchmark in computer animation. To further refine the control over the text-based motion style, we introduce an LLM parser that converts broad motion descriptions into detailed, part-specific representations. By decomposing the human body into movement-related parts, our method significantly enhances the precision and effectiveness of text-based motion stylization, enabling fine-grained control over individual body parts. Our model's effectiveness and generalization capabilities have been rigorously validated through extensive experiments, including text-based motion stylization and generating stylized motion with video sources, which all demonstrate the potential of our MM-LDM to advance the state-of-the-art motion stylization.
Wenfeng Song, Xingliang Jin, Shuai Li et al.· IEEE Transactions on Pattern...· 0 citations
Recent advances in generative models have achieved remarkable performance in text- and image-conditioned editing. However, preserving the content of a given image while referencing style patterns from another remains challenging, often leading to uncontrollable stylization results. In this paper, we approach image stylization from the perspective of continuous control, aiming to enable modern Diffusion Transformer (DiT)-based multi-reference editing models to (1) faithfully preserve the semantic structure of the content image, (2) render strong stylization effects, and (3) smoothly transition between the two. To this end, we propose a simple yet effective two-stage training strategy along with a style-strength-aware spline formulation. Specifically, in the first stage, the model is trained to produce strongly stylized outputs while preserving the content semantics as much as possible. In the second stage, with the base model frozen, we learn a set of anchor projectors that map various stylization strengths into the model parameter space. During inference, by performing style-strength-aware spline interpolation in a low-rank space, our method enables continuous control over stylization strength, even though the model is trained with only a few discrete strength levels. Extensive experiments demonstrate that our method supports precise and continuous manipulation of stylization strength while generating high-fidelity results with modern DiT models. Project page: https://reychiaro.github.io/StyleController.
This work introduces stylized phase manifolds—a compact, interpretable latent representation that disentangles motion content, the temporal structure, and style and develops a diffusion‐based motion generator that enables fine‐grained control over semantic, temporal, and stylistic aspects of motion.
Jingyuan Li, Peizhuo Li, A. Aristidou et al.· Computer graphics forum (Pri...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.