Pith. sign in

REVIEW 2 cited by

SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.16768 v1 pith:PUSLK7EA submitted 2024-08-29 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords sam2pointpromptablesegmentationhttpssegmentvideoszero-shotacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce SAM2Point, a preliminary exploration adapting Segment Anything Model 2 (SAM 2) for zero-shot and promptable 3D segmentation. SAM2Point interprets any 3D data as a series of multi-directional videos, and leverages SAM 2 for 3D-space segmentation, without further training or 2D-3D projection. Our framework supports various prompt types, including 3D points, boxes, and masks, and can generalize across diverse scenarios, such as 3D objects, indoor scenes, outdoor environments, and raw sparse LiDAR. Demonstrations on multiple 3D datasets, e.g., Objaverse, S3DIS, ScanNet, Semantic3D, and KITTI, highlight the robust generalization capabilities of SAM2Point. To our best knowledge, we present the most faithful implementation of SAM in 3D, which may serve as a starting point for future research in promptable 3D segmentation. Online Demo: https://huggingface.co/spaces/ZiyuG/SAM2Point . Code: https://github.com/ZiyuGuo99/SAM2Point .

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. 3DPipe: A Pipelined GPU Framework for Scalable Generalized Spatial Join over Polyhedral Objects

    cs.DB 2026-04 unverdicted novelty 5.0 of 10

    3DPipe delivers up to 9x faster 3D spatial joins on polyhedra via GPU pipelining, multi-level pruning, and chunked streaming compared to prior GPU methods.

  2. LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks

    cs.RO 2025-05 conditional novelty 4.0 of 10

    A unified vision-language-action model that emits a sub-task description followed by a discrete action token outperforms modular and action-only baselines on simulated long-horizon tabletop tasks.

Pith tools