DynaCore is presented, a unified architecture for efficient LLM serving via system-architecture co-design that substantially reduces service-level latency over quantization and reconfigurable accelerators, and proposes disaggregated quantization, applying dual-side quantization to prefill and weight-only quantization t...
Cong Guo, Chi-Yue Wei, Bo-Wen Duan et al.· 0 citations
Transformer attention mechanisms pose significant scalability challenges due to quadratic complexity in sequence length, and existing accelerators remain bottlenecked by dense arithmetic and data movement. This paper proposes CAMformer, a hardware accelerator that reinterprets attention as an associative memory operati...
Tergel Molom-Ochir, Benjamin F. Morris, Mark Horton et al.· IEEE Transactions on Circuit...· 0 citations
This study addresses challenges with Vortex, an architecture compatible with systolic-array-based accelerators with minimal hardware overhead, bridging the gap between extreme compression and efficient inference, and proposes codebook-wise contextual sparsity to align with VQ execution.
Haoxuan Shan, Cong Guo, Bo-Wen Duan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.