TRALF: training-free regional prompt adaptation and layer fusion for transformer-based text-to-image diffusion models
While recent Transformer-based diffusion models have significantly advanced text-to-image (T2I) synthesis, they inherently lack explicit spatial inductive bias. Consequently, generating complex scenes with multiple objects and fine-grained attributes often leads to severe "attribute leakage" and "spatial misalignment....