Sparse attention reduces core-attention computation, but its indexers still incur repeated selection work and per-layer key-cache storage. Reusing selected indices across layers reduces this overhead but constrains multiple layers to the same token set. We introduce LatentIndex, which extends the latent-sharing princip...
Zhao-Hui Wang, Zhi-Xin Pan, Fan-Xu Meng et al.· 0 citations
Polymers are fundamental to modern materials science because their backbone chemistry, monomer composition, and chain architecture can be systematically tuned to achieve a virtually unlimited range of mechanical, thermal, electronic, and optical properties. To accelerate the discovery and design of polymeric material...
Amberbir Alemayoh, Zhi-Xin Pan, Ning Wang· Polymer Science & Techno...· 0 citations
A workload characterization study of 13 419 CNN configurations on two GPU platforms under GPU telemetry reveals that energy, latency, and memory exhibit fundamentally distinct scaling behaviors: energy and latency diverge by 3x under high computational demand, and cross-GPU transferability differs by target.
This work proposes HISA (Hierarchical Indexed Sparse Attention), a plug-and-play replacement for the indexer that rewrites the search path from a flat token scan into a two-stage hierarchical procedure: a block-level coarse filtering stage that scores pooled block representations to discard irrelevant regions, followed...
Yufei Xu, Fan-Xu Meng, Fan Jiang et al.· arXiv.org· 10 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.