Jul 2026· Anais do XVIII Simpósio Brasileiro de Computação Ubíqua e Pervasiva (SBCUP 2026)· 1 citation· 22 references
TL;DR
These findings demonstrate that contemporary generative workloads can be effectively supported by decentralized edge infrastructures, providing practical insights for the design of energy-efficient and heterogeneous local AI systems.
Abstract
The rapid evolution of generative artificial intelligence has substantially increased the computational demands of image synthesis models, traditionally restricting their execution to centralized cloud infrastructures. In response to concerns related to data privacy, energy consumption, cost, and dependency on hyperscale providers, this work investigates the feasibility of executing state-of-the-art generative image models at the network. We present a quantitative performance evaluation of two representative model families, Stable Diffusion XL (SDXL) and Z-Image Turbo, executed on heterogeneous hardware, including high-end consumer GPUs, mobile-class devices, and legacy workstation GPUs from NVIDIA and AMD. The analysis focuses on latency, power consumption, resource utilization, and the impact of software stack optimizations, such as attention mechanisms and backend frameworks, under realistic hardware constraints. Results show that software-level optimizations are the primary factors determining inference viability, often outweighing raw computational throughput. While modern GPUs benefit from optimized attention mechanisms and improved energy efficiency, legacy and lower-power devices remain viable when combined with optimized runtimes and model compression techniques. These findings demonstrate that contemporary generative workloads can be effectively supported by decentralized edge infrastructures, providing practical insights for the design of energy-efficient and heterogeneous local AI systems.
The high-performance computing industry is moving beyond an era in which each generation of GPU provides uniform performance gains across all applications. The growing importance of AI is driving GPU architecture towards greater specialization, with more silicon devoted to Tensor Cores and reduced-precision arithmetic....
Matthew Tindale, I. Karlin, Tobias Salamon et al.· Inquiry@Queen's Undergraduat...· 0 citations
A mathematical model is proposed to predict throughput and energy consumption for concurrently executing CV workloads on edge GPU accelerators and can be integrated into functional simulation frameworks for edge–cloud deployment studies.
Abhinaba Chakraborty, D. Colle, M. Pickavet et al.· Journal of Real-Time Image P...· 0 citations
This work investigates the feasibility of reproducing benchmarks originally run on datacenter GPUs such as the NVIDIA A100 and RTX 8000 using consumer-grade graphics cards, focusing on the NVIDIA GeForce RTX 3050 and GTX 1060 with CUDA Graphs support. Seven NAS Parallel Benchmarks (BT, LU, SP, EP, IS, MG, and CG) are e...
Leandro L. Retzlaff, Calebe C. Pereira, Helena P. Veltri et al.· Anais do LIII Seminário Inte...· 0 citations
The results show that GDS performance depends heavily on file sizes, access patterns, and storage backends, and the path from benchmark to production is longer and more expensive than marketing materials suggest.
Prabhjyot Saluja, Khemraj Shukla, Sam Fulcomer et al.· Practice and Experience in A...· 0 citations
DiffPower translates design netlists into a PDK-agnostic bytecode representation, enabling analytical gradient computation via reverse-mode automatic differentiation, achieving up to a speedup over single-threaded CPU propagation on the largest evaluated design, with the GPU advantage growing with design scale.
Isaac Jacobson, Zhengjie Zhao, Rashmi Mehrotra et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.