Large vision-language models (LVLMs) exhibit strong multimodal in-context learning (ICL) capabilities, yet this ability degrades substantially as model size decreases. Knowledge distillation offers a natural way to bridge this gap, but existing methods primarily align output distributions or hidden representations dire...
Yanshu Li, Jia-Qian Li, Can-Ran Xiao et al.· 0 citations
In agentic AI systems, frozen foundation models are increasingly deployed as closed-weight API endpoints, making downstream adaptation possible only through the inputs and inference procedures surrounding the model. As a result, for each input query, two coupled decisions largely determine both answer quality and token...
Xi Xiao, Yun-Bei Zhang, Chen Liu et al.· 0 citations
Inspired by the Information Bottleneck principle, Prompted Information Bottlenecks (PIB) is introduced, a framework that regularizes layer-wise compression-sufficiency trade-offs and promotes a more coherent cross-layer information path.
Yuqi Li, Xi Xiao, Yun-Bei Zhang et al.· arXiv.org· 10 citations
This work injects two complementary semantic priors into Visual prompt tuning, a cascaded scheme that integrates both priors throughout ViT adaptation, and proposes a cascaded scheme that integrates both priors throughout ViT adaptation.
Xi Xiao, Xing-Jian Li, Cheng Han et al.· Trans. Mach. Learn. Res.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.