DSTAR: Accelerating Diffusion Transformers via Spatial and Temporal Redundancy Reduction
DSTAR, a software-hardware co-design framework that accelerates DiT inference by reducing spatial and temporal redundancy and incorporates a sparse attention reuse mechanism to minimize redundant computation in attention layers, and design a specialized hardware accelerator which achieves high efficiency in both latency and energy consumption.