Enhanced Text-to-Image Editing with Multi-Step Control and Explainability
Abstract
: Text-driven image editing has advanced significantly in generating and modifying visual content. Existing approaches often face challenges in maintaining visual coherence across sequential edits and providing informative rationales for alterations. This approach develops an improved text-to-image editing system that enables users to apply sequential edits while preserving previous alterations, with an added option to undo edits when necessary. Through the combination of robust fine-tuning techniques and leading-edge visual understanding models, the framework enhances edit consistency, image quality, and user control. Qualitative and illustrative quantitative results demonstrate the effectiveness of the InstructPix2Pix-MB-FT model in performing instruction-driven image editing tasks, achieving high realism and fidelity in object modification, scene enhancement, and human feature changes. The developed method has the potential to be used in creative design, content generation, and visual storytelling.