A novel approach to lossless convolutional neural network compression via progressive knowledge distillation-incorporated low-rank compression.
Model compression is widely used to deploy large neural networks on resource-constrained edge devices. Among existing techniques, low-rank composition is theoretically grounded in approximation theory and provides a strong basis for preserving model performance after compression. However, in practice, even state-of-the...