C2fDeploy: Function-Preserving Graph Rewriting to Eliminate Runtime Split Overhead on FPGA Deep Learning Processing Units
Efficient deployment of neural-network detectors on field-programmable gate array (FPGA) accelerators depends not only on model complexity but also on compiler-visible graph structure. On deep learning processing unit (DPU) platforms, unsupported operators can fragment execution between accelerator and host execution d...