Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially in high-resolution...
Xinye Zhao, Yun-Kai Dang, Yun-Chen Wu et al.· 0 citations
Riemannian Geometry-Sensitive Quantization (RGSQ), which formulates quantization as a reconstruction problem under a unified Fisher-Riemannian metric, enabling standard unimodal PTQ methods to evaluate multimodal quantization error under their original assumptions.
Zhi-Ping Wu, Dong-Dong Ren, Yang Zhou et al.· 0 citations
This work proposes B rain-inspired Supervised Supervised Reflection (BUS), a label-free training framework to enhance reflective reasoning capability in challenging image analysis and validate that backward prediction capability is critical for VLM reasoning.
Jiacheng Yang, Tongying Xiao, Yun-Kai Dang et al.· 0 citations
This paper first proves that TFS is a sufficient condition for weight disentanglement, and finds that TFS also gives rise to an observable geometric consequence: weight vector orthogonality, which positions TFS as the common cause for both the desired functional outcome and a measurable geometric property.
Shan Liu, Yuehan Yin, Lei Wang et al.· arXiv.org· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.