Event-Driven Language Models with Sparse Neural Activity for Neuromorphic Hardware
This work introduces a method that induces sparse neural activity in heavily quantized linear-attention models with minimal performance loss, and positions sparse, quantized linear-attention models as a natural fit for deploying LLMs on event-driven multi-core platforms.
Simon Richter, Ruhai Lin, Jason Yik et al.
· 0 citations