REVIEW 7 cited by
SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce SAM2Point, a preliminary exploration adapting Segment Anything Model 2 (SAM 2) for zero-shot and promptable 3D segmentation. SAM2Point interprets any 3D data as a series of multi-directional videos, and leverages SAM 2 for 3D-space segmentation, without further training or 2D-3D projection. Our framework supports various prompt types, including 3D points, boxes, and masks, and can generalize across diverse scenarios, such as 3D objects, indoor scenes, outdoor environments, and raw sparse LiDAR. Demonstrations on multiple 3D datasets, e.g., Objaverse, S3DIS, ScanNet, Semantic3D, and KITTI, highlight the robust generalization capabilities of SAM2Point. To our best knowledge, we present the most faithful implementation of SAM in 3D, which may serve as a starting point for future research in promptable 3D segmentation. Online Demo: https://huggingface.co/spaces/ZiyuG/SAM2Point . Code: https://github.com/ZiyuGuo99/SAM2Point .
Forward citations
Cited by 7 Pith papers
-
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
Applying test-time verifiers, DPO preference alignment, and a new adaptive reward model (PARM) to autoregressive image generators improves GenEval score from 53% to 77%.
-
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
A new evaluation suite finds that CLIPScore, HPSv2, and Aesthetic Score misjudge challenging text-to-image outputs, while GPT-4o and human ratings favor FLUX.1 and Ideogram2.0.
-
3DPipe: A Pipelined GPU Framework for Scalable Generalized Spatial Join over Polyhedral Objects
3DPipe delivers up to 9x faster 3D spatial joins on polyhedra via GPU pipelining, multi-level pruning, and chunked streaming compared to prior GPU methods.
-
InfiniteWorld: A Unified Scalable Simulation Framework for General Visual-Language Robot Interaction
InfiniteWorld presents an Isaac Sim based simulator with unified assets and four benchmarks, including scene graph exploration and social mobile manipulation, but reports zero success on the main social task.
-
CellSeg1: Robust Cell Segmentation with One Training Image
Training SAM with LoRA on one well-annotated cell image yields cell segmentation accuracy comparable to models trained on hundreds of images.
-
PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation
A pipeline that infers object material with a multimodal model and optimizes material parameters with optical flow from video diffusion to simulate 4D dynamic scenes.
-
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
A unified vision-language-action model that emits a sub-task description followed by a discrete action token outperforms modular and action-only baselines on simulated long-horizon tabletop tasks.
Discussion (0). Continue with ORCID to comment.