How Many Bits Does a Fact Need? Memory Load, Competing Alternatives and Weight Quantization in a Small Transformer
This work investigates how post-training weight quantization affects the knowledge memorized by a language model. To this end, we train a transformer with 1.1 million parameters on synthetic biographies, a setting in which the amount of information is known exactly, and we vary both the information load and the number...