Skip to content

Author

Feng Zhou

We have 1 of 25 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Jul 2026

Efficient inference and computational optimization of large language models for intelligent signal processing

The paper outlines an efficient scheme of inference in LLM by synergistically using post-training quantization, key-value cache compression, speculative decoding, and Flash Attention to fill the gap between the state-of-the-art LLM capabilities and the latency constraints of the intelligent signal processing systems.

Feng Zhou · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.