Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing
This work presents a novel framework that addresses two critical limitations of existing methods: inadequate modeling of hierarchical temporal structure and inability to handle complex many-to-many correspondences between modalities by introducing a multi-scale temporal convolutional encoder that captures motion patterns across different temporal granularities.