GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid Reasoning
GHR-VLM, a visual grounded hybrid reasoning framework for zero-shot transit-bus video analytics, is proposed, motivated by the observation that explicit visual grounding can improve VLM reasoning by converting long surveillance streams into compact, passenger-centered spatiotemporal evidence.
Kaicong Huang, Weiheng Oh, Jack M. Reilly et al.
· 0 citations