Conference
Open access
2026
Grouped Adaptive Weight Sharing (GAWS): An Inference-Efficient Adaptation Method for Large Language Models
GAWS is proposed, a novel adapter design based on structured Kronecker product decomposition that is positioned as a Pareto-efficient solution for deploying adapted LLMs in latency-sensitive settings, balancing the low latency of compressed adapters with the accuracy of LoRA.
Eman Alsuradi, Junhyung Lee, Kyeng-Hun Lee et al.
· Annual Meeting of the Associ... · 0 citations