Skip to content
Conference

Efficient Frame Retrieval for Traffic Law Question Answering over Dashcam Videos

Aug 2026 · International Conference on Multimedia Analysis and Pattern Recognition · pp. 730-735 · 0 citations · 24 references

Abstract

Traffic Law Question Answering over dashcam videos requires both accurate visual evidence selection and reliable multimodal reasoning. This task is especially challenging in Vietnamese traffic scenes, where road environments are dense, diverse, and highly dynamic. In this paper, we present a resource-efficient two-stage framework for traffic-law question answering over dashcam videos, developed for the "RoadBuddy: Understanding the Road through Dashcam AI" track of ZaloAI Challenge 2025. First, a CLIP ViT–based frame retriever reduces visual redundancy by selecting four query-relevant frames from each video. Second, the selected frames are processed by a Qwen3-VL 8B model adapted with LoRA for task-specific multiple-choice answer prediction. Despite its lightweight design, our method achieves competitive results, with scores of 0.70617 on the public leaderboard and 0.702 on the private leaderboard, ranking top 3 among 241 participating teams. These results show that efficient frame retrieval, when combined with parameter-efficient task adaptation, offers a practical solution for traffic-law question answering over real-world dashcam videos.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.