UAV-borne imaging has advanced from megapixel to gigapixel sensors, shifting aerial perception from recognizing individual targets to understanding entire dynamic scenes. We characterize this demand as Wide-area Spatio-temporal Scene Understanding (WSTU), which requires wide-area coverage, per-target resolution, and te...
Yu-Hang Zhu, Mei-Yi Zhu, Yun-Kai Dang et al.· 0 citations
Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially in high-resolution...
Xinye Zhao, Yun-Kai Dang, Yun-Chen Wu et al.· 0 citations
This work proposes B rain-inspired Supervised Supervised Reflection (BUS), a label-free training framework to enhance reflective reasoning capability in challenging image analysis and validate that backward prediction capability is critical for VLM reasoning.
Jiacheng Yang, Tongying Xiao, Yun-Kai Dang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.