A Layer Importance Metric for Quantization Accounting for the Speed-Quality Trade-off in Autoregressive Models
This work proposes a composite metric that combines two orthogonal criteria: information retention and throughput gains and finds that it allocates more resources to the most expressive layers compared to evolutionary search, specialized accelerators, or Shapley-value-based approaches that require expensive approximate inference.
A. Safronov
· 0 citations