Visual-token compression for vision--language models is posed almost entirely as a selection problem: decide which tokens to keep and discard the rest. The criteria that work best rank tokens by the attention the language model pays them, which makes the ranking a function of the question being asked. That is invisible...
Hong-Bo Zhang, Zi-Hao Yang, Liu-Yang Song et al.· 0 citations
Global vision--language similarities compress an image and a caption into one vector, preserving semantics but not which word corresponds to which region or how those regions are arranged; a model can recognize every word and object yet prefer a compositionally incorrect caption. We argue that a frozen encoder retains...
Liu-Yang Song, Yi Zhang, Zhong-Yi Deng et al.· 0 citations
FutureBridge is presented, which ranks joint LLM-SLM token candidates according to how well they support the SLM's subsequent reasoning, and indicates that token selection benefits from modeling whether the receiving SLM can use each candidate to continue reasoning, rather than relying on the LLM's local preference alo...
Quanquan Li, Hongbo Zhang, Yihe Chi et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.