SyncDreamer : Controllable and Expressive Avatar Generation Beyond the Talking Head
SyncDreamer is presented, a unified diffusion Transformer framework that generates identity-preserving and emotionally expressive talking avatars from only a single image, speech audio, and text prompt, and an RL-based Cross-Modal Prompt Enhancer grounding textual cues in visual context for fine-grained motion con-trol.