Skip to content

GS-Threads: An Energy-Efficient 3-D Gaussian Splatting Hardware Accelerator With Sparse-Aware Rasterization for Radiance Field Rendering

Oct 2026 · IEEE Transactions on Circuits and Systems Part 1: Regular Papers · Vol 73, pp. 6775-6786 · 0 citations · 50 references

Abstract

Real-time 3D Gaussian Splatting (3DGS) exhibits significant inefficiencies when mapped onto conventional tile-based rasterization, primarily due to the mismatch between anisotropic or finely projected Gaussian footprints and coarse rasterization tiles. This work presents GS-Threads, an energy-efficient 3DGS accelerator that exploits per-pixel sparsity via a Fine-grained Gaussian Rasterization Unit (FGRU). An <inline-formula> <tex-math notation="LaTeX">$8\times 8$ </tex-math></inline-formula> Early Test Unit (ETU) performs <inline-formula> <tex-math notation="LaTeX">$\sigma $ </tex-math></inline-formula> tests to generate hit bitmaps, while a shared Volume Rendering Unit (VRU) executes exponentiation and compositing only for valid pixels. Relative to a monolithic baseline, FGRU reduces LUT utilization by 56.6%, power by 28.6%, and rasterization latency by 32%-47% through sub-tile skipping. The design is implemented on a Xilinx ZCU104 FPGA with on-chip sorting and storage orchestration. At a clock frequency of 200 MHz and an <inline-formula> <tex-math notation="LaTeX">$800\times 800$ </tex-math></inline-formula> resolution, GS-Threads achieves 116 FPS with 2.69 W FPGA Programmable Logic (PL) power (23.1 mJ/frame on-chip). Compared with prior FPGA-based 3DGS accelerators, GS-Threads improves throughput by 1.7–<inline-formula> <tex-math notation="LaTeX">$2.4\times $ </tex-math></inline-formula> and reduces energy per frame by 2.3–<inline-formula> <tex-math notation="LaTeX">$3.1\times $ </tex-math></inline-formula>. Post-synthesis ASIC evaluation indicates that GS-Threads achieves <inline-formula> <tex-math notation="LaTeX">$1.5\times $ </tex-math></inline-formula> higher area-normalized throughput, and <inline-formula> <tex-math notation="LaTeX">$1.3\times $ </tex-math></inline-formula> higher energy-normalized area efficiency than the recent 3DGS ASIC, delivering competitive performance–energy efficiency relative to existing hardware neural rendering accelerators.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.