Balancing Visual Fidelity and Detection Reliability: Scalable Remote Sensing Image Compression
The surging volume of high-resolution remote sensing (RS) images and the limited transmission capacity of the satellite-to-ground link impose a pressing challenge on image compression: how to maintain higher reconstruction fidelity at lower bit rates without compromising the reliability of downstream vision tasks (e.g., object detection). To address this challenge, we propose an end-to-end scalable remote sensing image compression (SRSIC) framework. Considering that downstream tasks prioritize semantic structure while visual interpretation requires textural details, we adopt a scalable framework to decouple these features. Specifically, the compressed bitstream is divided into a base layer and an enhancement layer. The base layer is dedicated to compact semantic features optimized for object detection via a feature transfer network, bypassing the need for complete decoding. The enhancement layer supplements residual details for high-fidelity image reconstruction. Furthermore, considering the complex scale variations characteristic of RS images, we design a multiscale asymmetric codec to extract multiscale features and employ an adaptive context entropy model to minimize redundancy. Experimental results on the DIOR dataset demonstrate that SRSIC achieves significant bitrate savings, while maintaining better image reconstruction quality and higher object detection accuracy.