This novel progressive quantization framework combines multi-bit assistive teacher models with self-knowledge distillation to stabilize BNN training and integrates matching structured pruning with an asymmetric Binary Weight Network scaling factor, thereby reducing quantization errors while maintaining hardware efficiency.
Evaluations show that QACT provides a high-accuracy solution when retraining is affordable, whereas CSVQ offers efficient compression under strict computational budgets.
Zhao-Qing Li, Haoning Xu, Zengrui Jin et al.· IEEE Transactions on Audio,...· 0 citations
A comprehensive PTQ framework is presented that addresses the problem of compressing LLM weights through three core innovations: a calibration process guided by Kullback-Leibler divergence minimization to preserve the original weight distribution, a learnable codebook optimization mechanism employing noise substitution...
Bao Tan Duy Huynh, Takashi Tsunakawa, Masafumi Nishida· International Conference on...· 0 citations
Findings confirm that combining complementary compression strategies yields substantially better performance-efficiency trade-offs than any single technique applied in isolation.
Upma Sharma Archana· International Journal of Res...· 0 citations
Quantization-Aware Pre-Training (QAPT) can increase the inference efficiency of DNNs, but a problematic behaviour known as rounding boundary weight oscillation can introduce detrimental noise into the training process and significantly reduce convergence speed. While existing methods can reduce this detrimental noise,...
Ning-Feng Yang, T. Aamodt· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.