Jul 2026
OSEF: One-Step Evidence Fusion for Cross-Video Scene Procedure Planning
This work introduces Cross-Video Scene Procedure Planning (CVSPP): given an answer-redacted start-goal query and K candidate videos, a model must retrieve the supporting video, localize the relevant window, and predict the action sequence.
Zhentong Ye, Lei Zhang, Sijia Zhou et al.
· arXiv.org · 0 citations