GNR-Q: Inference-Time State-Conditioned Reconstruction for Memory-Efficient Quantized Language Models
Low-bit quantization substantially reduces the memory footprint of large language models(LLMs), but the associated loss of numerical precision can distort internal representations anddegrade downstream prediction quality. We investigate whether part of this lost representationquality can be reconstructed dynamically at...