This work studies an interesting problem: how to achieve fine visual perception under lower cost without larger images, and builds this framework on the advanced SigLIP 2 model, which consistently delivers stronger results than the baseline model, especially on OCR-related tasks.
Direct Advantage Amplification (DAA), which amplifies the advantages of hard-to-sample correct responses on hard prompts, as obtained by Dynamic Sampling, is proposed, which ensures that, when Dynamic Sampling is used, these hard-to-sample responses can be effectively capitalized on, implying higher training efficiency.
Si-Yuan Gan, Yu-Hang Li, Xiran Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.