SPAE: Spectrally Guided Autoencoder for Pretrained Visual Latents
SPAE employs a compact bottleneck to distill stable semantic information while suppressing high-frequency components, thereby improving the alignment between DiT-generated latents and encoder latents, and achieves a favorable balance among visual understanding, generation quality, and reconstruction fidelity.