Adaptive Log-Space (AL) quantization for non-negative states is introduced and its results make state topology and update semantics first-class design constraints for optimizer quantization.
Abstract
Optimizer-state quantization is commonly designed for Adam's dense, parameter-aligned first- and second-moment arrays. This abstraction breaks for memory-efficient optimizers, whose states may be factored, confidence-modulated, or maintained in a projected space, so similar reconstruction error can produce different update error. We formulate optimizer-state quantization as a joint problem over representation, topology, and update semantics. We then introduce Adaptive Log-Space (AL) quantization for non-negative states. AL fits each block's observed nonzero logarithmic interval and reserves a separate code for exact zero, enforcing $q = 0 \Leftrightarrow x = 0$; signed momentum and state precision remain independently selectable. Controlled probes show that adaptive ranges reduce update error and temporal drift, exact-zero reservation preserves dormant states, and state topology constrains useful block granularity. End-to-end language-model training evaluates the resulting policy across dense, factored, confidence, and projected optimizer states. On TinyLlama-1.1B, AL8 with uniform 8-bit momentum reaches 72.90 perplexity versus 73.54 for bitsandbytes 8-bit AdamW, with comparable optimizer-state storage and higher throughput. CAME matches reference-level final perplexity across three seeds when its non-negative states use AL16, while a semantic grouping-and-protection policy closes most of quantized Adafactor's 100K-step late-loss gap. These results make state topology and update semantics first-class design constraints for optimizer quantization.
DAMP uses both quantization-error energy and decay-based persistence to identify high-risk channels during offline calibration and stores these channels at higher precision and the remainder in INT8, the first to study post-training quantization of recurrent states in GDN and KDA based language models.
Tao Zhang, Jian-Chao Tan, Ping-Wei Sun et al.· 0 citations
SCHUROPT is introduced, which analytically eliminates the suffix's optimal continuous response, yielding an exact groupwise quadratic with Schur-complement curvature, and achieves the highest mean zero-shot accuracy among the evaluated backpropagation free PTQ baselines.
Gunjun Lee, Sehwan Son, Younjoo Lee et al.· 0 citations
QUASAR is introduced, a QAT method that continuously performs lightweight, loss-aware reconstruction in the training loop to lower the loss floor and improve the resulting low-bit model, establishing QUASAR's objective as a principled optimization target.
Vincent Counathe, Ben Athiwaratkun, C. De Sa et al.· 1 citation
This work forms AdamW as a finite-horizon input--state--output (ISO) system whose state contains the model parameters and first- and second-moment estimates, and derives an exact multistep error decomposition and establishes first-order finite-horizon accuracy under local smoothness and controlled activation switching.
RDQ (Residual Distribution Quantization), a PTQ framework whose central contribution is Cascaded Error Compensation, a sequential calibration procedure that captures the actual drifted activations each layer receives and fits per-channel AWQ-style scales against those drifted inputs, with scales folded into preceding RMSNorm weights for exact mathematical equivalence at zero inference overhead.
Support Vector Generation is introduced, a kernel-based framework that converts a frozen language model into an interpretable, training-free classifier for zero-and few-shot learning and suggests that SVG offers a viable path toward efficient, interpretable NLP systems under compute constraints.
Shohei Ohsawa· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.