Skip to content
Open access

C2fDeploy: Function-Preserving Graph Rewriting to Eliminate Runtime Split Overhead on FPGA Deep Learning Processing Units

Sep 2026 · Electronics · 0 citations · 20 references

Abstract

Efficient deployment of neural-network detectors on field-programmable gate array (FPGA) accelerators depends not only on model complexity but also on compiler-visible graph structure. On deep learning processing unit (DPU) platforms, unsupported operators can fragment execution between accelerator and host execution domains. We present C2fDeploy, a training-free rewrite for the YOLOv8 C2f block that moves channel splitting from the activation graph to an offline partition of the trained projection and batch-normalization parameters. The transformed block preserves the 32-bit floating-point (FP32) function without retraining or additional parameters. On the GRAZPEDWRI-DX fracture-detection task, the original and rewritten graphs produced identical FP32 test metrics and comparable 8-bit integer (INT8) accuracy. Direct tensor-level FP32 comparison at the outputs of all eight rewritten C2f blocks yielded an aggregate mean absolute error of 9.73×10−8 and a relative L2 error of 1.97×10−7, providing numerical verification beyond detection-level metrics. Compilation for the Kria KV260 consolidated nine DPU subgraphs into one and removed the C2f-related host-side slicing operations. In a same-checkpoint whole-XModel benchmark, C2fDeploy improved whole-XModel graph execution throughput by 95.2× and reduced energy per execution on the 5 V system-on-module (SOM) rail by 98.2%. Direct runtime profiling further showed that DPU compute-unit busy-time utilization increased from 0.36% to 90.75%, while system-wide CPU utilization decreased by 89.5%. Aggregate APM-observed external-memory bandwidth increased from 40.20 to 3250.56 MB/s as accelerator execution became more continuous, whereas normalized APM-observed traffic decreased by 14.8% per graph execution. These results show that compiler-aware graph rewriting can remove deployment bottlenecks without changing the trained detector.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.