DynaCore is presented, a unified architecture for efficient LLM serving via system-architecture co-design that substantially reduces service-level latency over quantization and reconfigurable accelerators, and proposes disaggregated quantization, applying dual-side quantization to prefill and weight-only quantization t...
Cong Guo, Chi-Yue Wei, Bo-Wen Duan et al.· 0 citations
The rapid advancement of deep learning has presented significant energy efficiency challenges to the conventional von Neumann architecture. In-memory computing (IMC) architectures based on emerging non-volatile memory (eNVM) are widely regarded as a promising solution for accelerating neural network training due to the...
Peng Dang, You-Na Huang, Yintao He et al.· 1 citation
This study addresses challenges with Vortex, an architecture compatible with systolic-array-based accelerators with minimal hardware overhead, bridging the gap between extreme compression and efficient inference, and proposes codebook-wise contextual sparsity to align with VQ execution.
Haoxuan Shan, Cong Guo, Bo-Wen Duan et al.· 0 citations
A voltage-aware ReRAM-based accelerator (VARA) is proposed, along with its accompanying design methodology, that reduces the average total system energy consumption and improves the average system energy efficiency and outperforming existing state-of-the-art accelerators for sparse-activation optimization.
Monte Carlo tree search (MCTS) enables artificial intelligence (AI) decision-making, but requires 55-300 W on conventional processors, limiting edge deployment. In-memory computing (IMC) is energy-efficient on regular workloads but has been considered incompatible with irregular multi-phase algorithms. We introduce pha...
Tergel Molom-Ochir, Benjamin F. Morris, Yintao He et al.· arXiv.org· 0 citations
The growing demand for long-context LLM inference has exposed a critical bandwidth–capacity trade-off in memory systems, rendering single-tier PIM architectures ineffective. HBM-PIMs offer high bandwidth but limited capacity, while DIMM-PIMs provide scalability at the cost of lower bandwidth; neither satisfies the thro...
Shixin Zhao, Lian Liu, Xiangwen An et al.· IEEE transactions on compute...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.