Tailoring LLM Weight Compression For PIM Architectures
This work investigates lightweight BF16 weight compression schemes tailored for PIM-based LLM inference by focusing on exponent-oriented compression methods that exploit the locality and redundancy present in BF16 exponent fields.
Sabiha Tajdari, Akhil Shekar, Kevin Skadron et al.
· 0 citations