Skip to content

Author

Haoxuan Shan

We have 2 of 12 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Preprint Oct 2026

A Shape-Adaptive Architecture with Disaggregated Quantization for Efficient LLM Serving

DynaCore is presented, a unified architecture for efficient LLM serving via system-architecture co-design that substantially reduces service-level latency over quantization and reconfigurable accelerators, and proposes disaggregated quantization, applying dual-side quantization to prefill and weight-only quantization t...

Cong Guo, Chi-Yue Wei, Bo-Wen Duan et al. · 0 citations
Preprint Sep 2026

Vortex: Bridging Extreme Compression and Efficient LLM Inference

This study addresses challenges with Vortex, an architecture compatible with systolic-array-based accelerators with minimal hardware overhead, bridging the gap between extreme compression and efficient inference, and proposes codebook-wise contextual sparsity to align with VQ execution.

Haoxuan Shan, Cong Guo, Bo-Wen Duan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.