PointBLIP: Zero-Training Point Cloud Classification Network Based on BLIP-2 Model

Yunzhe Xiao; Yong Dou; Shaowu Yang

doi:10.3390/rs16132453

Remote Sensing (Jul 2024)

PointBLIP: Zero-Training Point Cloud Classification Network Based on BLIP-2 Model

Yunzhe Xiao,
Yong Dou,
Shaowu Yang

Affiliations

Yunzhe Xiao: College of Computer Science and Technology, National University of Defense Technology, Changsha 410073, China
Yong Dou: College of Computer Science and Technology, National University of Defense Technology, Changsha 410073, China
Shaowu Yang: College of Computer Science and Technology, National University of Defense Technology, Changsha 410073, China

DOI: https://doi.org/10.3390/rs16132453
Journal volume & issue: Vol. 16, no. 13
p. 2453

Abstract

Read online

Leveraging the open-world understanding capacity of large-scale visual-language pre-trained models has become a hot spot in point cloud classification. Recent approaches rely on transferable visual-language pre-trained models, classifying point clouds by projecting them into 2D images and evaluating consistency with textual prompts. These methods benefit from the robust open-world understanding capabilities of visual-language pre-trained models and require no additional training. However, they face several challenges summarized as prompt ambiguity, image domain gap, view weight confusion, and feature deviation. In response to these challenges, we propose PointBLIP, a zero-training point cloud classification network based on the recently introduced BLIP-2 visual-language model. PointBLIP is adept at processing similarities between multi-images and multi-prompts. We separately introduce a novel method for point cloud zero-shot and few-shot classification, which involves comparing multiple features to achieve effective classification. Simultaneously, we enhance the input data quality for both the image and text sides of PointBLIP. In point cloud zero-shot classification tasks, we outperform state-of-the-art methods on three benchmark datasets. For few-shot classification tasks, to the best of our knowledge, we present the first zero-training few-shot point cloud method, surpassing previous works under the same conditions and showcasing comparable performance to full-training methods.

Published in Remote Sensing

ISSN: 2072-4292 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Science
Website: http://www.mdpi.com/journal/remotesensing/

About the journal

Abstract

Keywords