DRL-Enabled Polymorphic Acceleration Framework for Flexible and Energy-Efficient Hybrid-Float Deep Learning Inference in Mobile Computing
Hybrid-float quantization has emerged as a promising solution for efficient deep neural network inference on mobile platforms, but its practical deployment is still limited by three challenges: compute–transmission imbalance, rapidly growing mapping complexity, and the tradeoff between intermediate-data movement and ha...