Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI...
Fang-Zhi Zhong, Xue-Rui Qiu, Yu-Qi Pan et al.· 0 citations
SOLO is the first local learning method to show such memory and throughput gains in billion-parameter language-model pretraining, and becomes a practical alternative to backpropagation for large-scale pretraining.
Bo-Jian Yin, Shu-Rong Wang, Yu-Qi Pan et al.· 0 citations
PISA is proposed, a block-sparse attention mechanism that employs a pyramid Top-$K selection strategy, and develops hardware-aware Triton kernels for both training and inference, fusing hierarchical routing and LogSumExp scoring without materializing the query-key score matrix.
Bo-Hao Tang, Zhen Qin, Yu-Qi Pan et al.· 0 citations
Modular TTT is proposed, a framework that represents the inner learner as a directed acyclic graph and exposes the fast-weight network, loss function, learning rate, weight decay, and normalization as explicit design dimensions and finds that small learning-rate initialization, weight decay, and a single-layer nonlinea...
Bo-Hao Tang, Zhen Qin, Yu-Qi Pan et al.· 1 citation
LoGo, a token-level dynamic local-global attention mechanism that uses attention span as a direct proxy for attention budget allocation, is proposed and results suggest that learned token-level span allocation is an effective and scalable way to improve the long-context performance-compute trade-off.
Yu-Qi Pan, Zheng Li, Bo-Hao Tang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.