A large audio language model (LALM) turns a minute of speech into 750-1,500 tokens and prefills every one. Image-token pruning often cuts after the language model's first layers, where image tokens draw little attention. Audio tokens draw much more attention there, and their ranking is still far from final, so audio ne...
Kyoungjun Park, Yun-Zhe Li, Li-Li Qiu· 0 citations
Speculative decoding accelerates autoregressive generation by using a smaller drafter to propose tokens for batched verification by a larger target. However, conventional speculative decoding couples drafting to the target's evolving verified prefix, serializing drafting and verification. We ask whether this dependency...
Yun-Zhe Li, Kyoungjun Park, Hong-Zi Zhu et al.· 0 citations
OmniCache is proposed, a unified hierarchical caching framework that performs multidimensional feature reuse through Token Cache, Frame Cache, Block Cache, and Layered Cache that reuses spatial features in temporal layers and temporal features in spatial layers, while Layered Cache captures cross-step redundancy at the...
Zhaoyuan He, Muhammad Muaz, Lili Qiu· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.