2D3D-DescNet: Jointly Learning 2D and 3D Local Feature Descriptors for Cross-Dimensional Matching

Shuting Chen; Yanfei Su; Baiqi Lai; Luwei Cai; Chengxi Hong; Li Li; Xiuliang Qiu; Hong Jia; Weiquan Liu

doi:10.3390/rs16132493

Remote Sensing (Jul 2024)

2D3D-DescNet: Jointly Learning 2D and 3D Local Feature Descriptors for Cross-Dimensional Matching

Shuting Chen,
Yanfei Su,
Baiqi Lai,
Luwei Cai,
Chengxi Hong,
Li Li,
Xiuliang Qiu,
Hong Jia,
Weiquan Liu

Affiliations

Shuting Chen: Chengyi College, Jimei University, Xiamen 361021, China
Yanfei Su: School of Computer and Information Engineering, Xiamen University of Technology, Xiamen 361024, China
Baiqi Lai: Fujian Key Laboratory of Sensing and Computing for Smart Cities, School of Informatics, Xiamen University, Xiamen 361005, China
Luwei Cai: Queen’s Business School, Queen’s University Belfast, Belfast BT7 1NN, UK
Chengxi Hong: Chengyi College, Jimei University, Xiamen 361021, China
Li Li: Chengyi College, Jimei University, Xiamen 361021, China
Xiuliang Qiu: Chengyi College, Jimei University, Xiamen 361021, China
Hong Jia: Fujian Key Laboratory of Sensing and Computing for Smart Cities, School of Informatics, Xiamen University, Xiamen 361005, China
Weiquan Liu: Fujian Key Laboratory of Sensing and Computing for Smart Cities, School of Informatics, Xiamen University, Xiamen 361005, China

DOI: https://doi.org/10.3390/rs16132493
Journal volume & issue: Vol. 16, no. 13
p. 2493

Abstract

Read online

The cross-dimensional matching of 2D images and 3D point clouds is an effective method by which to establish the spatial relationship between 2D and 3D space, which has potential applications in remote sensing and artificial intelligence (AI). In this paper, we propose a novel multi-task network, 2D3D-DescNet, to learn 2D and 3D local feature descriptors jointly and perform cross-dimensional matching of 2D image patches and 3D point cloud volumes. The 2D3D-DescNet contains two branches with which to learn 2D and 3D feature descriptors, respectively, and utilizes a shared decoder to generate the feature maps of 2D image patches and 3D point cloud volumes. Specifically, the generative adversarial network (GAN) strategy is embedded to distinguish the source of the generated feature maps, thereby facilitating the use of the learned 2D and 3D local feature descriptors for cross-dimensional retrieval. Meanwhile, a metric network is embedded to compute the similarity between the learned 2D and 3D local feature descriptors. Finally, we construct a 2D-3D consistent loss function to optimize the 2D3D-DescNet. In this paper, the cross-dimensional matching of 2D images and 3D point clouds is explored with the small object of the 3Dmatch dataset. Experimental results demonstrate that the 2D and 3D local feature descriptors jointly learned by 2D3D-DescNet are similar. In addition, in terms of 2D and 3D cross-dimensional retrieval and matching between 2D image patches and 3D point cloud volumes, the proposed 2D3D-DescNet significantly outperforms the current state-of-the-art approaches based on jointly learning 2D and 3D feature descriptors; the cross-dimensional retrieval at TOP1 on the 3DMatch dataset is improved by over 12%.

Published in Remote Sensing

ISSN: 2072-4292 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Science
Website: http://www.mdpi.com/journal/remotesensing/

About the journal

Abstract

Keywords