Skip to content
Preprint

CARB: A Characterization-Guided Framework for CNN Inference Cost Prediction and Deployment Screening

Aug 2026 · 0 citations · 27 references
Computer Science

TL;DR

A workload characterization study of 13 419 CNN configurations on two GPU platforms under GPU telemetry reveals that energy, latency, and memory exhibit fundamentally distinct scaling behaviors: energy and latency diverge by 3x under high computational demand, and cross-GPU transferability differs by target.

Abstract

Accurate pre-deployment estimation of CNN inference cost--energy, latency, and peak memory--is increasingly critical as models are deployed on resource-constrained GPU platforms. Existing approaches rely on FLOPs, latency measurements, or single-device profiling as energy proxies, overlooking the non-linear interactions between architectural design and hardware load. We present a workload characterization study of 13 419 CNN configurations on two GPU platforms (RTX 5090 and RTX 3080) under GPU telemetry, revealing that energy, latency, and memory exhibit fundamentally distinct scaling behaviors: energy and latency diverge by 3x under high computational demand, and cross-GPU transferability differs by target--energy and latency require platform-specific models while memory transfers well across the two tested platforms. Building on these characterization findings, we develop CARB, a cascade-blended ensemble that jointly predicts all three targets with R2 ~0.99, and a two-stage deployment screening workflow that eliminates over 90% of candidates in seconds, reducing large design spaces to a Pareto-prioritized shortlist validated against real hardware.

View source

Similar papers

Book Open access Aug 2026

A Differentiable Simulator for Optimizing Time-Domain Analog CNN Accelerators

This work introduces a differentiable, GPU-accelerated CNN inference simulator that backpropagates through the analog accumulation trajectory, and allows gradient-based exploration of high-dimensional, heterogeneous hardware-parameter settings.

Mark Horton, Changwoo Park, Tergel Molom-Ochir et al. · 0 citations
Open access Aug 2026

Power Prediction Modelling for CPU–GPU Heterogeneous Platforms Based on Heterogeneous Feature Sequence Modelling

Central processing unit-graphics processing unit (CPU-GPU) heterogeneous platforms are widely deployed for AI-related, data-processing, and compute-intensive workloads. Their practical deployment is constrained by fluctuating node power, cross-device coordination overhead, and uneven energy efficiency. This paper devel...

J.-W. Hu · 0 citations
Aug 2026

Hardware-Aware Neural Network Deployment on Multi-Core in-Memory Computing Systems: A Compiler Perspective

Conventional compute-in-memory (CIM) deployment flows usually assume a fixed crossbar geometry, although DNN layers often have different channel counts and matrix shapes. The resulting shape mismatch leaves part of the array capacity unused and can increase the number of split-and-transfer operations during inference....

Kaiwen Deng, Sifan Sun, Hanjie Liu et al. · 0 citations
Preprint Aug 2026

Achieving Near-Zero-Overhead Multi-Model Hierarchical Classification in Real-Time Detection Pipelines

This work presents a five-step methodology for zero GPU fallback DLA INT8 deployment of classification backbones, comprising architecture adaptation, manual dynamic range workaround to rescue TensorRT's implicit quantization, and generalizes to any detection-classification edge pipeline.

V. Raju · 0 citations
Book Open access Aug 2026

Application-Driven System Technology Co-Optimization for 2.5D Edge AI Platforms

Edge AI systems face stringent run-time constraints and tight energy budgets, demanding comprehensive optimization of computing platforms. Nonetheless, existing approaches often focus only on the co-design of hardware and software, but rarely link algorithmic opportunities with system- and technology-level optimization...

Anna Burdina, David Mallasén, A. Levisse et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.