SAMPro3D

SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation

¹SSE, CUHKSZ ²FNii, CUHKSZ ³Microsoft Research Asia
^†Corresponding Author

3DV 2025

Abstract

We introduce SAMPro3D for zero-shot instance segmentation of 3D scenes. Given the 3D point cloud and multiple posed RGB-D frames of 3D scenes, our approach segments 3D instances by applying the pretrained Segment Anything Model (SAM) to 2D frames. Our key idea involves locating SAM prompts in 3D to align their projected pixel prompts across frames, ensuring the view consistency of SAM-predicted masks. Moreover, we suggest selecting prompts from the initial set guided by the information of SAM-predicted masks across all views, which enhances the overall performance. We further propose to consolidate different prompts if they are segmenting different surface parts of the same 3D instance, bringing a more comprehensive segmentation. Notably, our method does not require any additional training. Extensive experiments on diverse benchmarks show that our method achieves comparable or better performance compared to previous zero-shot or fully supervised approaches, and in many cases surpasses human annotations. Furthermore, since our fine-grained predictions often lack annotations in available datasets, we present ScanNet200-Fine50 test data which provides fine-grained annotations on 50 scenes from ScanNet200 dataset.

Qualitative Comparison

The qualitative comparison of our method, SAM3D ( zero-shot ), Mask3D ( fully-supervised ) and ScanNet200's annotations, across various scenes in the ScanNet200 validation set, from holistic to focused view.

BibTeX

@inproceedings{xu2025sampro3d, title={SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation}, author={Mutian Xu and Xingyilang Yin and Lingteng Qiu and Yang Liu and Xin Tong and Xiaoguang Han}, year={2025}, booktitle = {International Conference on 3D Vision (3DV)} }

SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation

3DV 2025

We introduce SAMPro3D for zero-shot 3D instance segmentation.

Abstract

Key Idea

The comparison of our key idea and others. Our method (b) locates SAM prompts in 3D, which aligns pixel prompts across frames, bringing the frame consistency of prompts and their masks, and can handle newly emerged instances.

Method Overview

Animated Qualitative Comparison

Qualitative Comparison

The qualitative comparison of our method, SAM3D ( zero-shot ), Mask3D ( fully-supervised ) and ScanNet200's annotations, across various scenes in the ScanNet200 validation set, from holistic to focused view.

The qualitative comparison of our method, SAM3D ( zero-shot ), Mask3D ( fully-supervised ) and ScanNet200's annotations, across various scenes in the ScanNet200 validation set, from holistic to focused view.

BibTeX