Skip to content

Author

Yintao He

We have 6 of 28 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Preprint Oct 2026

A Shape-Adaptive Architecture with Disaggregated Quantization for Efficient LLM Serving

DynaCore is presented, a unified architecture for efficient LLM serving via system-architecture co-design that substantially reduces service-level latency over quantization and reconfigurable accelerators, and proposes disaggregated quantization, applying dual-side quantization to prefill and weight-only quantization t...

Cong Guo, Chi-Yue Wei, Bo-Wen Duan et al. · 0 citations
Preprint Sep 2026

An Emerging NVM-Based On-Chip Training Architecture with Non-Ideality Mitigation Through Bipolar Weight Distributions

The rapid advancement of deep learning has presented significant energy efficiency challenges to the conventional von Neumann architecture. In-memory computing (IMC) architectures based on emerging non-volatile memory (eNVM) are widely regarded as a promising solution for accelerating neural network training due to the...

Peng Dang, You-Na Huang, Yintao He et al. · 1 citation
Preprint Sep 2026

Vortex: Bridging Extreme Compression and Efficient LLM Inference

This study addresses challenges with Vortex, an architecture compatible with systolic-array-based accelerators with minimal hardware overhead, bridging the gap between extreme compression and efficient inference, and proposes codebook-wise contextual sparsity to align with VQ execution.

Haoxuan Shan, Cong Guo, Bo-Wen Duan et al. · 0 citations
Preprint Aug 2026

VARA: A Voltage-Aware ReRAM-Based Accelerator for Energy-Efficient Computing

A voltage-aware ReRAM-based accelerator (VARA) is proposed, along with its accompanying design methodology, that reduces the average total system energy consumption and improves the average system energy efficiency and outperforming existing state-of-the-art accelerators for sparse-activation optimization.

Peng Dang, Yin-Tao He, Huawei Li · 0 citations
Jul 2026

Multi-primitive in-memory computing for Monte Carlo tree search

Monte Carlo tree search (MCTS) enables artificial intelligence (AI) decision-making, but requires 55-300 W on conventional processors, limiting edge deployment. In-memory computing (IMC) is energy-efficient on regular workloads but has been considered incompatible with irregular multi-phase algorithms. We introduce pha...

Tergel Molom-Ochir, Benjamin F. Morris, Yintao He et al. · 0 citations
Sep 2026

HydraPIM: A Heterogeneous PIM Architecture for Efficient Attention in Long-Context LLMs

The growing demand for long-context LLM inference has exposed a critical bandwidth–capacity trade-off in memory systems, rendering single-tier PIM architectures ineffective. HBM-PIMs offer high bandwidth but limited capacity, while DIMM-PIMs provide scalability at the cost of lower bandwidth; neither satisfies the thro...

Shixin Zhao, Lian Liu, Xiangwen An et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.