CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating
Concentrate and Concentrate (CaC) is a coarse-to-fine anomaly reward model based on Vision-Language Models that first conducts a global temporal scan to anchor anomalous time windows, then performs fine-grained spatial grounding within the localized interval, and finally derives robust judgments via structured spatiote...