CLIP-based Point Cloud Classification via Point Cloud to Image   Translation

Shuvozit Ghose; Manyi Li; Yiming Qian; Yang Wang

arXiv:2408.03545·cs.CV·August 8, 2024

CLIP-based Point Cloud Classification via Point Cloud to Image Translation

Shuvozit Ghose, Manyi Li, Yiming Qian, Yang Wang

PDF

Open Access

TL;DR

This paper introduces PPCITNet, a novel method that translates point clouds into colored images with visual cues, enhancing CLIP-based classification accuracy for 3D point cloud data.

Contribution

It proposes a point cloud to image translation network and a viewpoint adapter, addressing limitations of existing CLIP-based methods for improved classification.

Findings

01

Outperforms state-of-the-art CLIP-based models on multiple datasets

02

Achieves higher accuracy in point cloud classification tasks

03

Demonstrates the effectiveness of visual cues in point cloud understanding

Abstract

Point cloud understanding is an inherently challenging problem because of the sparse and unordered structure of the point cloud in the 3D space. Recently, Contrastive Vision-Language Pre-training (CLIP) based point cloud classification model i.e. PointCLIP has added a new direction in the point cloud classification research domain. In this method, at first multi-view depth maps are extracted from the point cloud and passed through the CLIP visual encoder. To transfer the 3D knowledge to the network, a small network called an adapter is fine-tuned on top of the CLIP visual encoder. PointCLIP has two limitations. Firstly, the point cloud depth maps lack image information which is essential for tasks like classification and recognition. Secondly, the adapter only relies on the global representation of the multi-view features. Motivated by this observation, we propose a Pretrained Point…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

Topics3D Surveying and Cultural Heritage · Remote Sensing and LiDAR Applications · 3D Shape Modeling and Analysis

MethodsContrastive Language-Image Pre-training · Adapter