Practical Modeling for Split DNN Inference on Near-Edge Accelerators
Splitting complex model inference between multiple computing devices can overcome latency and energy constraints at the edge. Newer edge accelerator devices with higher computing capacity and energy efficiency, enable more fine-grained offload throughout layers of the network, leading to the potential for multiple spli...