Incentivizing Vision Language Models to Search for Long Video Question Answering
VSeek is introduced, an agentic framework that transforms long-video question answering (LVQA) from a passive, single-pass perception task into a multi-turn retrieval process and proposes a novel neuro-symbolic approach that bridges open-ended natural language with discrete visual verification.