Pith. sign in

REVIEW 10 cited by

Gaga: Group Any Gaussians via 3D-aware Memory Bank

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.07977 v4 pith:HRR2Y6WG submitted 2024-04-11 cs.CV

classification cs.CV
keywords gagasegmentationmasksbankcameraclass-agnosticd-awaredemonstrates
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce Gaga, a framework that reconstructs and segments open-world 3D scenes by leveraging inconsistent 2D masks predicted by zero-shot class-agnostic segmentation models. Contrasted to prior 3D scene segmentation approaches that rely on video object tracking or contrastive learning methods, Gaga utilizes spatial information and effectively associates object masks across diverse camera poses through a novel 3D-aware memory bank. By eliminating the assumption of continuous view changes in training images, Gaga demonstrates robustness to variations in camera poses, particularly beneficial for sparsely sampled images, ensuring precise mask label consistency. Furthermore, Gaga accommodates 2D segmentation masks from diverse sources and demonstrates robust performance with different open-world zero-shot class-agnostic segmentation models, significantly enhancing its versatility. Extensive qualitative and quantitative evaluations demonstrate that Gaga performs favorably against state-of-the-art methods, emphasizing its potential for real-world applications such as 3D scene understanding and manipulation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Relation-Centric Open-Vocabulary 3D Gaussian Segmentation

    cs.CV 2026-07 unverdicted novelty 7.0 of 10

    PairGS builds a relation graph from sparse pairwise affinities on 3D Gaussians to achieve SOTA open-vocabulary segmentation with a 50x faster variant than optimization-based methods.

  2. LEGO: Leveled Language Gaussian Splatting

    cs.CV 2026-08 conditional novelty 6.0 of 10

    LEGO builds view-consistent, multi-level 3D semantic hierarchies from multi-view SAM masks by clustering their physical 3D scales, and grounds them with CLIP for open-vocabulary segmentation and LLM-driven spatial grounding.

  3. GaussianSelector: Lightweight Human-Guided Object Selection in 3D Gaussian Splatting with Graph Optimization

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A training-free graph-cut method selects 3D objects from Gaussian splatting scenes using sparse user scribbles, reaching 92.2 mIoU on NVOS with three interaction views.

  4. IGFuse: Interactive 3D Gaussian Scene Reconstruction via Multi-Scans Fusion

    cs.CV 2025-08 conditional novelty 6.0 of 10

    IGFuse jointly optimizes segmentation-aware Gaussian fields from multiple scans of rearranged scenes, producing complete, manipulable 3D reconstructions without inpainting.

  5. ObjectGS: Object-aware Scene Reconstruction and Scene Understanding via Gaussian Splatting

    cs.GR 2025-07 conditional novelty 6.0 of 10

    ObjectGS unifies 3D Gaussian scene reconstruction with object-level segmentation by binding each object to local anchors with fixed one-hot ID encodings, improving open-vocabulary and panoptic segmentation.

  6. DSG-World: Learning a 3D Gaussian World Model from Dual State Videos

    cs.CV 2025-06 conditional novelty 6.0 of 10

    DSG-World builds two segmented 3D Gaussian fields from two scene states and trains them with mutual consistency, enabling novel-state simulation without inpainting or dense capture.

  7. Tackling View-Dependent Semantics in 3D Language Gaussian Splatting

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new 3D language Gaussian Splatting method that clusters per-object multi-view CLIP features and reweights them to capture view-dependent semantics, improving direct 3D open-vocabulary segmentation.

  8. SLGaussian: Fast Language Gaussian Splatting in Sparse Views

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SLGaussian builds a 3D semantic field from two photos in a single forward pass, stores CLIP features in a memory bank for fast open-vocabulary queries, and reports higher IoU than LangSplat and LERF on the LERF and 3D...

  9. Lifting by Gaussians: A Simple, Fast and Flexible Method for 3D Instance Segmentation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    LBG segments 3D Gaussian Splatting scenes into objects, parts, and subparts by assigning each pixel's maximum-contributing Gaussian a 2D mask ID and merging fragments across frames using geometric and semantic similarity.

  10. The ALMA-QUARKS Survey: III. Clump-to-core fragmentation and search for high-mass starless cores

    astro-ph.GA 2025-08 unverdicted novelty 4.0 of 10

    In 139 infrared-bright massive protoclusters, ALMA resolves 1562 cores whose separations are much smaller than the Jeans length, and finds only two candidate high-mass starless cores.

Pith tools