Skip to content
Open access

Impact of GPU Architecture and VRAM on Image Generation: A Study of Energy Efficiency in Heterogeneous Edge Nodes

Jul 2026 · Anais do XVIII Simpósio Brasileiro de Computação Ubíqua e Pervasiva (SBCUP 2026) · 1 citation · 22 references

TL;DR

These findings demonstrate that contemporary generative workloads can be effectively supported by decentralized edge infrastructures, providing practical insights for the design of energy-efficient and heterogeneous local AI systems.

Abstract

The rapid evolution of generative artificial intelligence has substantially increased the computational demands of image synthesis models, traditionally restricting their execution to centralized cloud infrastructures. In response to concerns related to data privacy, energy consumption, cost, and dependency on hyperscale providers, this work investigates the feasibility of executing state-of-the-art generative image models at the network. We present a quantitative performance evaluation of two representative model families, Stable Diffusion XL (SDXL) and Z-Image Turbo, executed on heterogeneous hardware, including high-end consumer GPUs, mobile-class devices, and legacy workstation GPUs from NVIDIA and AMD. The analysis focuses on latency, power consumption, resource utilization, and the impact of software stack optimizations, such as attention mechanisms and backend frameworks, under realistic hardware constraints. Results show that software-level optimizations are the primary factors determining inference viability, often outweighing raw computational throughput. While modern GPUs benefit from optimized attention mechanisms and improved energy efficiency, legacy and lower-power devices remain viable when combined with optimized runtimes and model compression techniques. These findings demonstrate that contemporary generative workloads can be effectively supported by decentralized edge infrastructures, providing practical insights for the design of energy-efficient and heterogeneous local AI systems.

Read PDF

Similar papers

Conference Open access Sep 2026

Analyzing the Impact of Architectural Design Decisions on Performance Across Generations of NVIDIA GPUs

The high-performance computing industry is moving beyond an era in which each generation of GPU provides uniform performance gains across all applications. The growing importance of AI is driving GPU architecture towards greater specialization, with more silicon devoted to Tensor Cores and reduced-precision arithmetic....

Matthew Tindale, I. Karlin, Tobias Salamon et al. · 0 citations
Aug 2026

Modeling throughput and power consumption for real-time concurrent vision application deployments on edge

A mathematical model is proposed to predict throughput and energy consumption for concurrently executing CV workloads on edge GPU accelerators and can be integrated into functional simulation frameworks for edge–cloud deployment studies.

Abhinaba Chakraborty, D. Colle, M. Pickavet et al. · 0 citations
Open access Jul 2026

Comparative Performance Analysis of Workload on Enterprise GPUs with Consumer Platforms Accelerated by CUDA Graphs

This work investigates the feasibility of reproducing benchmarks originally run on datacenter GPUs such as the NVIDIA A100 and RTX 8000 using consumer-grade graphics cards, focusing on the NVIDIA GeForce RTX 3050 and GTX 1060 with CUDA Graphs support. Seven NAS Parallel Benchmarks (BT, LU, SP, EP, IS, MG, and CG) are e...

Leandro L. Retzlaff, Calebe C. Pereira, Helena P. Veltri et al. · 0 citations
Preprint Aug 2026

Architecting the Next Generation of Asynchronous, Distributed GPUs for the AI Era

The rapid evolution of machine learning workloads has fundamentally transformed GPU hardware, driving architectures toward Multi-Chip Module (MCM) topologies, asynchronous execution primitives, and persistent, multi-phase kernel behaviors. Despite these shifts, cycle-level simulation infrastructure has lagged behind, l...

Junrui Pan, Wei-Li An, Cesar Avalos Baddouh et al. · 2 citations · ⚡2
Book Open access Jul 2026

Experience with NVIDIA GPUDirect Storage (GDS) in Academic HPC Environments: Challenges, Pitfalls, and Practical Limitations

The results show that GDS performance depends heavily on file sizes, access patterns, and storage backends, and the path from benchmark to production is longer and more expensive than marketing materials suggest.

Prabhjyot Saluja, Khemraj Shukla, Sam Fulcomer et al. · 0 citations
Preprint Aug 2026

DiffPower: GPU-Accelerated Differentiable Switching Power Analysis and Optimization

DiffPower translates design netlists into a PDK-agnostic bytecode representation, enabling analytical gradient computation via reverse-mode automatic differentiation, achieving up to a speedup over single-threaded CPU propagation on the largest evaluated design, with the GPU advantage growing with design scale.

Isaac Jacobson, Zhengjie Zhao, Rashmi Mehrotra et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.