Component-Aware Spatio-Temporal Adaptation of Frozen Foundation Models for Video Deepfake Detection
This work proposes a parameter-efficient, video-based deepfake detection framework that leverages a frozen foundation model encoder coupled with a lightweight spatio-temporal decoder, complemented by a Bidirectional Spatio-Temporal decoder that models local temporal transitions and bidirectional temporal dependencies across sampled frames, enabling robust temporal reasoning within each video clip.