Deploying large vision-language models (VLMs) onboard satellites enables onboard data processing and reduces raw data downlink. However, onboard inference faces two resource challenges. Limited onboard memory and energy require model compression and distributed deployment. Dynamic resource availability requires fast de...
Tong Quan, Yuan-Long Wan, Hua-Sen He et al.· 0 citations
Video large language models (VideoLLMs) achieve strong video understanding performance, but their inference remains expensive due to the large number of redundant spatio-temporal visual tokens in long videos. Existing token pruning methods alleviate this cost by reducing redundant tokens, yet most of them rely on segme...
Meng-Jie Zhang, Qi-Hui Zhu, Tao Zhang et al.· 0 citations
LLM-based multi-agent systems (MAS) increasingly use latent collaboration to avoid the information loss and repeated encoding-decoding overhead of natural-language communication. However, directly forwarding all sender latents makes the receiver-side context scale with both the number of agents and the reasoning length...
Shi-Nan Zhang, Tao Zhang, Qi-Hui Zhu et al.· 0 citations
This work finds that the deep-layer features of a lightweight speculative model exhibit strong consistency with the target model in the selection of critical tokens for recomputation, and proposes SpecCache, which employs deep-layer hidden-state norms from a speculative model as a proxy to guide the critical token sele...
Zijian Wen, Tao Zhang, Shuangwu Chen et al.· Annual Meeting of the Associ...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.