DCAT: Dual Cross-Attention-Based Transformer for Change Detection

Yuan Zhou; Chunlei Huo; Jiahang Zhu; Leigang Huo; Chunhong Pan

doi:10.3390/rs15092395

Remote Sensing (May 2023)

DCAT: Dual Cross-Attention-Based Transformer for Change Detection

Yuan Zhou,
Chunlei Huo,
Jiahang Zhu,
Leigang Huo,
Chunhong Pan

Affiliations

Yuan Zhou: School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 101408, China
Chunlei Huo: School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 101408, China
Jiahang Zhu: School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 101408, China
Leigang Huo: School of Computer and Information Engineering, Nanning Normal University, Nanning 530001, China
Chunhong Pan: National Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China

DOI: https://doi.org/10.3390/rs15092395
Journal volume & issue: Vol. 15, no. 9
p. 2395

Abstract

Read online

Several transformer-based methods for change detection (CD) in remote sensing images have been proposed, with Siamese-based methods showing promising results due to their two-stream feature extraction structure. However, these methods ignore the potential of the cross-attention mechanism to improve change feature discrimination and thus, may limit the final performance. Additionally, using either high-frequency-like fast change or low-frequency-like slow change alone may not effectively represent complex bi-temporal features. Given these limitations, we have developed a new approach that utilizes the dual cross-attention-transformer (DCAT) method. This method mimics the visual change observation procedure of human beings and interacts with and merges bi-temporal features. Unlike traditional Siamese-based CD frameworks, the proposed method extracts multi-scale features and models patch-wise change relationships by connecting a series of hierarchically structured dual cross-attention blocks (DCAB). DCAB is based on a hybrid dual branch mixer that combines convolution and transformer to extract and fuse local and global features. It calculates two types of cross-attention features to effectively learn comprehensive cues with both low- and high-frequency information input from paired CD images. This helps enhance discrimination between the changed and unchanged regions during feature extraction. The feature pyramid fusion network is more lightweight than the encoder and produces powerful multi-scale change representations by aggregating features from different layers. Experiments on four CD datasets demonstrate the advantages of DCAT architecture over other state-of-the-art methods.

Published in Remote Sensing

ISSN: 2072-4292 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Science
Website: http://www.mdpi.com/journal/remotesensing/

About the journal

Abstract

Keywords