Integrating Temporal Supervision and Self-Attention for Audio-Driven Head Synthesis
NeRF-based talking head methods can render individual frames with impressive fidelity, yet the assembled videos often flicker. The reason is structural: each frame is optimized independently, so nothing in the training objective ties frame t to frame t − 1. We present TemporalTalk, which closes this gap with three trai...