Pith. sign in

REVIEW 7 cited by

SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.16768 v1 pith:PUSLK7EA submitted 2024-08-29 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords sam2pointpromptablesegmentationhttpssegmentvideoszero-shotacross
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce SAM2Point, a preliminary exploration adapting Segment Anything Model 2 (SAM 2) for zero-shot and promptable 3D segmentation. SAM2Point interprets any 3D data as a series of multi-directional videos, and leverages SAM 2 for 3D-space segmentation, without further training or 2D-3D projection. Our framework supports various prompt types, including 3D points, boxes, and masks, and can generalize across diverse scenarios, such as 3D objects, indoor scenes, outdoor environments, and raw sparse LiDAR. Demonstrations on multiple 3D datasets, e.g., Objaverse, S3DIS, ScanNet, Semantic3D, and KITTI, highlight the robust generalization capabilities of SAM2Point. To our best knowledge, we present the most faithful implementation of SAM in 3D, which may serve as a starting point for future research in promptable 3D segmentation. Online Demo: https://huggingface.co/spaces/ZiyuG/SAM2Point . Code: https://github.com/ZiyuGuo99/SAM2Point .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Applying test-time verifiers, DPO preference alignment, and a new adaptive reward model (PARM) to autoregressive image generators improves GenEval score from 53% to 77%.

  2. IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A new evaluation suite finds that CLIPScore, HPSv2, and Aesthetic Score misjudge challenging text-to-image outputs, while GPT-4o and human ratings favor FLUX.1 and Ideogram2.0.

  3. 3DPipe: A Pipelined GPU Framework for Scalable Generalized Spatial Join over Polyhedral Objects

    cs.DB 2026-04 unverdicted novelty 5.0 of 10

    3DPipe delivers up to 9x faster 3D spatial joins on polyhedra via GPU pipelining, multi-level pruning, and chunked streaming compared to prior GPU methods.

  4. InfiniteWorld: A Unified Scalable Simulation Framework for General Visual-Language Robot Interaction

    cs.RO 2024-12 conditional novelty 5.0 of 10

    InfiniteWorld presents an Isaac Sim based simulator with unified assets and four benchmarks, including scene graph exploration and social mobile manipulation, but reports zero success on the main social task.

  5. CellSeg1: Robust Cell Segmentation with One Training Image

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Training SAM with LoRA on one well-annotated cell image yields cell segmentation accuracy comparable to models trained on hundreds of images.

  6. PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A pipeline that infers object material with a multimodal model and optimizes material parameters with optical flow from video diffusion to simulate 4D dynamic scenes.

  7. LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks

    cs.RO 2025-05 conditional novelty 4.0 of 10

    A unified vision-language-action model that emits a sub-task description followed by a discrete action token outperforms modular and action-only baselines on simulated long-horizon tabletop tasks.

Pith tools