Efficient inference and computational optimization of large language models for intelligent signal processing
The paper outlines an efficient scheme of inference in LLM by synergistically using post-training quantization, key-value cache compression, speculative decoding, and Flash Attention to fill the gap between the state-of-the-art LLM capabilities and the latency constraints of the intelligent signal processing systems.