UniEarth: A Unified Foundation Model for Remote Sensing Image Generation and Understanding
Abstract
Remote sensing (RS) underpins geospatial intelligence by enabling large-scale Earth observation (EO). Existing RS approaches are typically developed along two separate lines: multimodal understanding (MMU) and image generation. Such separation limits parameter sharing and the potential synergy between understanding and generation. A unified foundation model can in principle support both tasks, but it still yields inferior performance in the RS domain due to omitting specific issues, including RS data scarcity, granularity mismatch between RS imagery and text, and optimization instability under long-tailed distributions. To address the above issues together, we propose a unified RS foundation model, called UniEarth, for joint RS-oriented understanding and generation. Built on the linear-complexity Mamba-2 architecture, UniEarth is data-efficient even with limited RS supervision. To stabilize joint optimization under coarse RS annotations, we design a cross-modal semantic anchoring mechanism on the intermediate hidden states of the shared backbone, balancing the generation and understanding pathways. Furthermore, UniEarth designs difficulty-aware bandit optimization (DA-UCB) to stabilize training by adaptively prioritizing low-resource categories in long-tailed distributions. Experiments on widely used RS benchmarks show that UniEarth consistently outperforms state-of-the-art methods on both understanding and generation, with a 44.6% relative reduction in Fréchet inception distance (FID) for image generation and a 20.1% relative improvement in CIDEr for captioning, while also achieving leading accuracy on visual question answering and scene classification.