Pith. sign in

REVIEW 2 cited by

PointSeg: A Training-Free Paradigm for 3D Scene Segmentation via Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.06403 v5 pith:JTWUS7KP submitted 2024-03-11 cs.CV

classification cs.CV
keywords foundationmodelspointsegpromptsacrossdatasetsscenetraining-free
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Recent success of vision foundation models have shown promising performance for the 2D perception tasks. However, it is difficult to train a 3D foundation network directly due to the limited dataset and it remains under explored whether existing foundation models can be lifted to 3D space seamlessly. In this paper, we present PointSeg, a novel training-free paradigm that leverages off-the-shelf vision foundation models to address 3D scene perception tasks. PointSeg can segment anything in 3D scene by acquiring accurate 3D prompts to align their corresponding pixels across frames. Concretely, we design a two-branch prompts learning structure to construct the 3D point-box prompts pairs, combining with the bidirectional matching strategy for accurate point and proposal prompts generation. Then, we perform the iterative post-refinement adaptively when cooperated with different vision foundation models. Moreover, we design a affinity-aware merging algorithm to improve the final ensemble masks. PointSeg demonstrates impressive segmentation performance across various datasets, all without training. Specifically, our approach significantly surpasses the state-of-the-art specialist training-free model by 14.1$\%$, 12.3$\%$, and 12.6$\%$ mAP on ScanNet, ScanNet++, and KITTI-360 datasets, respectively. On top of that, PointSeg can incorporate with various foundation models and even surpasses the specialist training-based methods by 3.4$\%$-5.4$\%$ mAP across various datasets, serving as an effective generalist model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Test-Time Optimization for Domain Adaptive Open Vocabulary Segmentation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A plug-and-play test-time optimization method improves zero-shot open-vocabulary segmentation on specialized-domain datasets by jointly tuning per-category text embeddings and aggregating visual features.

  2. Foundational Models for 3D Point Clouds: A Survey and Outlook

    cs.CV 2025-01 conditional novelty 4.0 of 10

    A structured review of methods that build or adapt 2D foundation models and LLMs for 3D point cloud tasks, with a proposed taxonomy and curated paper list.

Pith tools