Abstract. Current diffusion models struggle to achieve fine-grained remote sensing imagery (RSI) generation. This limitation fundamentally stems from their reliance on "flattened" text prompts, which overlook the inherent hierarchical structure of RSI. This paper proposes a fine-grained RSI generation method driven by expert knowledge and hierarchical captions. We first deconstruct RSI into a hierarchical "element-relation-scene" caption and employ an automatic caption optimization mechanism, grounded in spatial relation knowledge, to ensure high fidelity. Critically, we introduce a novel hierarchical caption encoding mechanism that efficiently injects decoupled hierarchical caption segments into the U-Net's cross-attention layers. This design enables the model to exert hierarchical and decoupled attentional control over the global scene, spatial layout, and geographical element details during the denoising process. Experiments demonstrate that, when combined with efficient fine-tuning algorithms such as LoRA, our method significantly outperforms traditional single-level captions across all six evaluation metrics, exemplified by the FID metric decreasing from 228.43 to 205.59 and the GSHPS metric increasing from 0.86 to 0.92. This research provides a new paradigm for controllable remote sensing scene generation, establishing an effective link between hierarchical semantic understanding and the progressive generation process of diffusion models.
Jiaxin Ren, Wanzeng Liu, Feng Zhang et al.· ISPRS Annals of the Photogra...· 0 citations
High-precision, pixel-level annotations are indispensable for remote sensing semantic segmentation and related tasks, yet producing such labels manually is prohibitively expensive. Although recent generative models can synthesize realistic remote sensing data, existing approaches typically either rely heavily on preexisting ground-truth masks as conditioning inputs or lack precise control over the spatial layout of the generated content. To address this gap, we propose Graph2Scene, a novel framework for the joint generation of remote sensing images and pixel-level labels driven by scene graphs. This framework establishes a flexible control mechanism that utilizes scene graphs derived from existing semantic labels during training to learn semantic priors, while enabling users to explicitly define object quantities, categories, and topological relationships for customized generation during inference. Graph2Scene adopts a two-stage cascade: Graph2Mask encodes the scene graph into textual prompts and employs an image-level low-rank adaptation (LoRA) to finetune a FLUX model for label generation; Mask2Scene uses a class-level LoRA strategy to learn fine-grained visual features and generates remote sensing images conditioned on the label. Experiments on our developed non-agricultural conversion process (NACP) dataset and the public LoveDA dataset show that Graph2Scene effectively produces structurally coherent and visually realistic remote sensing images together with accurate pixel-level annotations. The code will be made available at https://github.com/GeoRSAI/Graph2Scene
Shaoxuan Zhao, Xiaoguang Zhou, Dongyang Hou et al.· IEEE Geoscience and Remote S...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.