Breaking Memory Wall for Fast Edge LLM Inference Using Contextual Sparsity
Deploying Large Language Models (LLMs) on memory-constrained edge servers to serve requests from mobile devices is challenging due to their substantial resource demands. The Key-Value (KV) cache and Feed-Forward Network (FFN) parameters consume the majority of available memory. However, existing methods typically rely...