REVIEW 19 cited by
SAM3D: Segment Anything in 3D Scenes
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
SAM3D: Segment Anything in 3D Scenes
read the original abstract
In this work, we propose SAM3D, a novel framework that is able to predict masks in 3D point clouds by leveraging the Segment-Anything Model (SAM) in RGB images without further training or finetuning. For a point cloud of a 3D scene with posed RGB images, we first predict segmentation masks of RGB images with SAM, and then project the 2D masks into the 3D points. Later, we merge the 3D masks iteratively with a bottom-up merging approach. At each step, we merge the point cloud masks of two adjacent frames with the bidirectional merging approach. In this way, the 3D masks predicted from different frames are gradually merged into the 3D masks of the whole 3D scene. Finally, we can optionally ensemble the result from our SAM3D with the over-segmentation results based on the geometric information of the 3D scenes. Our approach is experimented with ScanNet dataset and qualitative results demonstrate that our SAM3D achieves reasonable and fine-grained 3D segmentation results without any training or finetuning of SAM.
Forward citations
Cited by 19 Pith papers
-
3AM: 3egment Anything with Geometric Consistency in Videos
3AM integrates MUSt3R 3D features into SAM2 via a Feature Merger and FOV-aware sampling to deliver geometry-consistent video object segmentation from RGB alone, with large gains on wide-baseline datasets.
-
GReFEM: Multimodal LLMs as Zero-Shot Semantic Assistants for Physics-Guided 3D Mesh Refinement
GReFEM shows MLLMs zero-shot isolate load-activated geometric features for volumetric mesh refinement with higher precision than matched-budget geometric heuristics.
-
JOPP-3D: Joint Open Vocabulary Semantic Segmentation on Point Clouds and Panoramas
A training-free pipeline jointly segments panoramic images and reconstructed point clouds with open-vocabulary language queries via tangential decomposition, instance proposals, CLIP alignment, and depth-based label t...
-
Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization from Unposed Views
A feed-forward framework learns instance-structured 3D token groups from unposed multi-view images via differentiable rendering, enabling native object-level segmentation, editing, and retrieval without 3D supervision.
-
Ilov3Splat: Instance-Level Open-Vocabulary 3D Scene Understanding in Gaussian Splatting
Ilov3Splat learns view-consistent CLIP and instance feature fields on 3D Gaussians to support open-vocabulary object selection and segmentation without category labels.
-
PanoSAMic: Panoramic Image Segmentation from SAM Feature Encoding and Dual View Fusion
PanoSAMic modifies SAM with multi-stage feature encoding, spatio-modal fusion, spherical attention, and dual-view fusion to achieve SOTA panoramic semantic segmentation on public RGB and RGB-D datasets.
-
ShelfGaussian: Shelf-Supervised Open-Vocabulary Gaussian-based 3D Scene Understanding
ShelfGaussian achieves state-of-the-art zero-shot semantic occupancy prediction on Occ3D-nuScenes by jointly supervising Gaussian representations with vision foundation model features at 2D image and 3D scene levels.
-
NeuMesh++: Towards Versatile and Efficient Volumetric Editing with Disentangled Neural Mesh-based Implicit Field
A disentangled mesh-vertex neural radiance field enables mesh-guided geometry edits, texture swap/fill/paint, and semantic-guided edits with claimed efficiency and quality gains.
-
GraspFoM: Towards Reconstruction-Driven Robotic Grasping with 3D Foundation Priors
GraspFoM creates a shared 3D latent from SAM3D priors, adds an anchor-initialized diffuser for multimodal grasps, and uses reconstruction-aware scoring plus residual updates to jointly achieve SOTA reconstruction and ...
-
ESAM++: Efficient Online 3D Perception on the Edge
ESAM++ introduces a 3D Sparse Feature Pyramid Network for efficient online 3D scene perception on edge devices, claiming competitive accuracy with up to 3x faster inference and 2x smaller model size than ESAM on four ...
-
AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models
AgentGrounder performs zero-shot 3D visual grounding on colored point clouds via an offline object lookup table and an online agent that selectively retrieves, scores geometrically, and renders images on demand, repor...
-
CAR-SAM: Cross-Attention Reconstruction for Post-Training Quantization of the Segment Anything Model
CAR-SAM introduces MatMul-Aware Compensation and Joint Cross-Attention Reconstruction to enable stable 4-bit post-training quantization of SAM, outperforming prior PTQ methods by 14.6% mAP on SAM-B and 6.6% on SAM-L.
-
Distill, Diffuse, and Semanticize (DDS): Annotation-Free 3D Scene Understanding Based on Multi-Granularity Distillation and Graph-Diffusion-Based Segmentation
DDS combines multi-granularity distillation from projected 2D features with graph diffusion on superpoints to deliver region-consistent semantic labels for 3D scenes without any dense annotations.
-
Distill, Diffuse, and Semanticize (DDS): Annotation-Free 3D Scene Understanding Based on Multi-Granularity Distillation and Graph-Diffusion-Based Segmentation
DDS performs annotation-free 3D semantic scene understanding by combining multi-granularity distillation from 2D visuals with graph-diffusion segmentation and cluster-to-category association, yielding up to 5.9% oAcc,...
-
Ilov3Splat: Instance-Level Open-Vocabulary 3D Scene Understanding in Gaussian Splatting
Ilov3Splat embeds multi-resolution hash-encoded CLIP features and contrastive instance fields into Gaussian splats, then uses query matching plus two-stage 3D clustering to enable instance-level open-vocabulary 3D sce...
-
MV3DIS: Multi-View Mask Matching via 3D Guides for Zero-Shot 3D Instance Segmentation
MV3DIS uses 3D-guided mask matching and depth consistency to produce more consistent multi-view 2D masks that refine into accurate zero-shot 3D instances.
-
Heterogeneous and Adept Snapshot Distillation for 3D Semantic Segmentation
HAS-KD combines information-oriented heterogeneous distillation from multi-modal models with adept snapshot distillation from training checkpoints to reach SOTA 3D semantic segmentation on ScanNetV2 and S3DIS without ...
-
T-FunS3D: Task-Driven Hierarchical Open-Vocabulary 3D Functionality Segmentation
T-FunS3D is a task-driven hierarchical method for open-vocabulary 3D functionality segmentation that constructs an open-vocabulary scene graph and applies vision-language models to achieve comparable accuracy with fas...
-
GraspSense: Physically Grounded Grasp and Grip Planning for a Dexterous Robotic Hand via Language-Guided Perception and Force Maps
GraspSense computes force maps from object geometry to select mechanically safe grasp regions and regulate grip forces for dexterous hands.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.