Open access
Jul 2026
Evaluating large language model compression: a comparative analysis on state-of-the-art models across diverse hardware platforms
Quantization, pruning, and parameter-efficient fine-tuning methods enabled models to match or outperform models with up to 18 times the parameters on SQuAD v2 while reducing storage as well as optimizer overhead and often improved task performance even when full fine-tuning failed.
Dominik Hildebrand, Benjamin Kiefer, Andreas Zell
· Artificial Intelligence Revi... · 0 citations