Temporal Coherence in Video-Language Models for Long-Form Narrative Understanding
Video-language systems perform well on short clips and poorly on material that unfolds over minutes or hours, and the gap has proved resistant to increases in model size. This article argues that the difficulty is structural rather than one of capacity. Methods developed for clips of a few seconds inherit two assumptions, namely that a small set of sampled frames represents the whole and that the language attached to a segment describes only that segment, and both assumptions fail once a question depends on events separated in time. We propose a framework that separates temporal coherence into four levels, covering perceptual continuity, event segmentation, entity persistence and causal narrative structure, and we place published architectures within it. Analysis of the cost of full attention over long token sequences shows why hierarchical, memory based and state space designs have replaced dense attention for extended input. We then examine evaluation, where diagnostic studies have shown that a large share of questions on standard benchmarks can be answered from a single frame, which means reported accuracy overstates temporal ability. The article closes with six problems that stand between current systems and reliable narrative understanding.