Diffusion Enhancement for Cloud Removal in Ultra-Resolution Remote Sensing Imagery

CV SR and Denoising ET
2024年01月25日
本文提出了一种基于深度学习的云去除技术,旨在解决云层对光学遥感图像质量和有效性的影响。然而,现有的深度学习技术在准确重构图像的原始视觉真实性和详细语义内容方面遇到了困难。为了解决这个挑战,本文提出了数据和方法两方面的改进。在数据方面,建立了一个超高分辨率基准,名为CUHK云去除(CUHK-CR),空间分辨率为0.5米。这个基准包含了丰富的细节纹理和多样的云覆盖,为设计和评估云去除模型提供了坚实的基础。从方法的角度来看,提出了一种新的扩散增强框架(Diffusion Enhancement,DE)来执行渐进的纹理细节恢复,从而减轻了训练难度并提高了推理精度。此外,开发了一个权重分配网络(WA)来动态调整特征融合的权重,从而进一步提高性能,特别是在超高分辨率图像生成的情况下。此外,采用粗到细的训练策略,有效地加快了训练收敛速度,同时减少了处理超高分辨率图像所需的计算复杂度。在新建立的CUHK-CR和现有的数据集(RICE)上进行了大量实验,结果表明,所提出的DE框架在感知质量和信号保真度方面优于现有的基于深度学习的方法。
The presence of cloud layers severely compromises the quality and effectiveness of optical remote sensing (RS) images. However, existing deep-learning (DL)-based Cloud Removal (CR) techniques encounter difficulties in accurately reconstructing the original visual authenticity and detailed semantic content of the images. To tackle this challenge, this work proposes to encompass enhancements at the data and methodology fronts. On the data side, an ultra-resolution benchmark named CUHK Cloud Removal (CUHK-CR) of 0.5m spatial resolution is established. This benchmark incorporates rich detailed textures and diverse cloud coverage, serving as a robust foundation for designing and assessing CR models. From the methodology perspective, a novel diffusion-based framework for CR called Diffusion Enhancement (DE) is proposed to perform progressive texture detail recovery, which mitigates the training difficulty with improved inference accuracy. Additionally, a Weight Allocation (WA) network is developed to dynamically adjust the weights for feature fusion, thereby further improving performance, particularly in the context of ultra-resolution image generation. Furthermore, a coarse-to-fine training strategy is applied to effectively expedite training convergence while reducing the computational complexity required to handle ultra-resolution images. Extensive experiments on the newly established CUHK-CR and existing datasets such as RICE confirm that the proposed DE framework outperforms existing DL-based methods in terms of both perceptual quality and signal fidelity.
许愿