Integrating Temporal Supervision and Self-Attention for Audio-Driven Head Synthesis
Abstract
NeRF-based talking head methods can render individual frames with impressive fidelity, yet the assembled videos often flicker. The reason is structural: each frame is optimized independently, so nothing in the training objective ties frame t to frame t − 1. We present TemporalTalk, which closes this gap with three training-time mechanisms: a face-masked temporal consistency loss with stop-gradient and delayed activation, a multi-head self-attention module that replaces the fixed convolutional attention over the audio context window, and cosine learning-rate annealing with linear warmup. None of the three touches the inference path, so runtime cost is unchanged. On the May benchmark, TemporalTalk reduces temporal flicker measured by tLPIPS (the mean Learned Perceptual Image Patch Similarity between consecutive generated frames) by 24% relative to SyncTalk, while also improving PSNR by 0.14 dB, LPIPS (Learned Perceptual Image Patch Similarity) by 3.0%, and landmark distance by 2.3%. An ablation study shows that all three mechanisms contribute, with the temporal loss accounting for the largest share of the gain.