Compress and Forget: bitsandbytes Quantization Amplifies Proactive Interference in LLMs
The results suggest that bitsandbytes 4-bit quantization can impose an additional cost on applications relying on long, updatable, semantically dense contexts, even when aggregate benchmark accuracy appears largely unaffected.
Shayan Shahrabi-Farahani, D. Rahmati
· 0 citations