REVIEW 4 cited by
LUDVIG: Learning-Free Uplifting of 2D Visual Features to Gaussian Splatting Scenes
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We address the problem of extending the capabilities of vision foundation models such as DINO, SAM, and CLIP, to 3D tasks. Specifically, we introduce a novel method to uplift 2D image features into Gaussian Splatting representations of 3D scenes. Unlike traditional approaches that rely on minimizing a reconstruction loss, our method employs a simpler and more efficient feature aggregation technique, augmented by a graph diffusion mechanism. Graph diffusion refines 3D features, such as coarse segmentation masks, by leveraging 3D geometry and pairwise similarities induced by DINOv2. Our approach achieves performance comparable to the state of the art on multiple downstream tasks while delivering significant speed-ups. Notably, we obtain competitive segmentation results using only generic DINOv2 features, despite DINOv2 not being trained on millions of annotated segmentation masks like SAM. When applied to CLIP features, our method demonstrates strong performance in open-vocabulary object segmentation tasks, highlighting the versatility of our approach.
Forward citations
Cited by 4 Pith papers
-
CF3: Compact and Fast 3D Feature Fields
CF3 builds a compact 3D feature field from a pre-trained 3DGS by feature lifting, per-Gaussian autoencoding, and adaptive sparsification, matching baseline segmentation quality with roughly 5% of the Gaussians.
-
SeqAffordSplat: Scene-level Sequential Affordance Reasoning on 3D Gaussian Splatting
SeqAffordSplat introduces a 3DGS benchmark for long-horizon affordance tasks and SeqSplatNet, an LLM-driven model that predicts ordered sequences of 3D affordance masks from single instructions.
-
PanSt3R: Multi-view Consistent Panoptic Segmentation
A single network jointly reconstructs 3D scene geometry and predicts multi-view consistent panoptic segmentation from unposed images in one forward pass, without test-time optimization.
-
Hierarchical Scoring with 3D Gaussian Splatting for Instance Image-Goal Navigation
A two-stage scorer, CLIP semantic prefiltering plus DINOv2 geometric matching over a 3D Gaussian map, achieves a 0.784 success rate on HM3D instance image-goal navigation.
Discussion (0). Continue with ORCID to comment.