GS-Threads: An Energy-Efficient 3-D Gaussian Splatting Hardware Accelerator With Sparse-Aware Rasterization for Radiance Field Rendering
Abstract
Real-time 3D Gaussian Splatting (3DGS) exhibits significant inefficiencies when mapped onto conventional tile-based rasterization, primarily due to the mismatch between anisotropic or finely projected Gaussian footprints and coarse rasterization tiles. This work presents GS-Threads, an energy-efficient 3DGS accelerator that exploits per-pixel sparsity via a Fine-grained Gaussian Rasterization Unit (FGRU). An <inline-formula> <tex-math notation="LaTeX">$8\times 8$ </tex-math></inline-formula> Early Test Unit (ETU) performs <inline-formula> <tex-math notation="LaTeX">$\sigma $ </tex-math></inline-formula> tests to generate hit bitmaps, while a shared Volume Rendering Unit (VRU) executes exponentiation and compositing only for valid pixels. Relative to a monolithic baseline, FGRU reduces LUT utilization by 56.6%, power by 28.6%, and rasterization latency by 32%-47% through sub-tile skipping. The design is implemented on a Xilinx ZCU104 FPGA with on-chip sorting and storage orchestration. At a clock frequency of 200 MHz and an <inline-formula> <tex-math notation="LaTeX">$800\times 800$ </tex-math></inline-formula> resolution, GS-Threads achieves 116 FPS with 2.69 W FPGA Programmable Logic (PL) power (23.1 mJ/frame on-chip). Compared with prior FPGA-based 3DGS accelerators, GS-Threads improves throughput by 1.7–<inline-formula> <tex-math notation="LaTeX">$2.4\times $ </tex-math></inline-formula> and reduces energy per frame by 2.3–<inline-formula> <tex-math notation="LaTeX">$3.1\times $ </tex-math></inline-formula>. Post-synthesis ASIC evaluation indicates that GS-Threads achieves <inline-formula> <tex-math notation="LaTeX">$1.5\times $ </tex-math></inline-formula> higher area-normalized throughput, and <inline-formula> <tex-math notation="LaTeX">$1.3\times $ </tex-math></inline-formula> higher energy-normalized area efficiency than the recent 3DGS ASIC, delivering competitive performance–energy efficiency relative to existing hardware neural rendering accelerators.