Language guided 3D object detection in point clouds for MEP scenes

Junjie Li; Shengli Du; Jianfeng Liu; Weibiao Chen; Manfu Tang; Lei Zheng; Lianfa Wang; Chunle Ji; Xiao Yu; Wanli Yu

doi:10.1049/cvi2.12261

IET Computer Vision (Jun 2024)

Language guided 3D object detection in point clouds for MEP scenes

Junjie Li,
Shengli Du,
Jianfeng Liu,
Weibiao Chen,
Manfu Tang,
Lei Zheng,
Lianfa Wang,
Chunle Ji,
Xiao Yu,
Wanli Yu

Affiliations

Junjie Li: China Coal Shaanxi Yulin Energy & Chemical Co., Ltd. of China National Coal Group Co. Yulin Shanxi China
Shengli Du: China Coal Shaanxi Yulin Energy & Chemical Co., Ltd. of China National Coal Group Co. Yulin Shanxi China
Jianfeng Liu: China Coal Shaanxi Yulin Energy & Chemical Co., Ltd. of China National Coal Group Co. Yulin Shanxi China
Weibiao Chen: China Coal Shaanxi Yulin Energy & Chemical Co., Ltd. of China National Coal Group Co. Yulin Shanxi China
Manfu Tang: China Coal Shaanxi Yulin Energy & Chemical Co., Ltd. of China National Coal Group Co. Yulin Shanxi China
Lei Zheng: China Coal Shaanxi Yulin Energy & Chemical Co., Ltd. of China National Coal Group Co. Yulin Shanxi China
Lianfa Wang: China Coal Electric Co., Ltd of China National Coal Group Co. Beijing China
Chunle Ji: China Coal Electric Co., Ltd of China National Coal Group Co. Beijing China
Xiao Yu: IOT Perception Mine Research Center China University of Mining and Technology Xuzhou Jiangsu China
Wanli Yu: Institute of Electrodynamics and Microelectronics University of Bremen Bremen Germany

DOI: https://doi.org/10.1049/cvi2.12261
Journal volume & issue: Vol. 18, no. 4
pp. 526 – 539

Abstract

Read online

Abstract In recent years, contrastive language‐image pre‐training (CLIP) has gained popularity for processing 2D data. However, the application of cross‐modal transferable learning to 3D data remains a relatively unexplored area. In addition, high‐quality, labelled point cloud data for Mechanical, Electrical, and Plumbing (MEP) scenarios are in short supply. To address this issue, the authors introduce a novel object detection system that employs 3D point clouds and 2D camera images, as well as text descriptions as input, using image‐text matching knowledge to guide dense detection models for 3D point clouds in MEP environments. Specifically, the authors put forth the proposition of a language‐guided point cloud modelling (PCM) module, which leverages the shared image weights inherent in the CLIP backbone. This is done with the aim of generating pertinent category information for the target, thereby augmenting the efficacy of 3D point cloud target detection. After sufficient experiments, the proposed point cloud detection system with the PCM module is proven to have a comparable performance with current state‐of‐the‐art networks. The approach has 5.64% and 2.9% improvement in KITTI and SUN‐RGBD, respectively. In addition, the same good detection results are obtained in their proposed MEP scene dataset.

Published in IET Computer Vision

ISSN: 1751-9632 (Print); 1751-9640 (Online)
Publisher: Wiley
Country of publisher: United Kingdom
LCC subjects: Medicine: Medicine (General): Computer applications to medicine. Medical informatics; Science: Mathematics: Instruments and machines: Electronic computers. Computer science: Computer software
Website: https://ietresearch.onlinelibrary.wiley.com/journal/17519640

About the journal

Abstract

Keywords