Post-Training VLMs for Video Mistake Detection
This work proposes the first video-language-model post-training technique for mistake detection, which uses a tailored reward function to encourage the model to identify discrepancies between an instruction and the corresponding video, and generalizes especially well to unseen procedures.
Federico Spurio, Olga Zatsarynna, Lars Doorenbos et al.
· 0 citations