Pith. sign in

REVIEW 5 cited by

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.04383 v2 pith:RIYUN5E7 submitted 2024-12-05 cs.CV cs.RO

classification cs.CVcs.RO
keywords descriptionsseegroundzero-shotdatagroundingimagesmethodsmodule
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

3D Visual Grounding (3DVG) aims to locate objects in 3D scenes based on textual descriptions, essential for applications like augmented reality and robotics. Traditional 3DVG approaches rely on annotated 3D datasets and predefined object categories, limiting scalability and adaptability. To overcome these limitations, we introduce SeeGround, a zero-shot 3DVG framework leveraging 2D Vision-Language Models (VLMs) trained on large-scale 2D data. SeeGround represents 3D scenes as a hybrid of query-aligned rendered images and spatially enriched text descriptions, bridging the gap between 3D data and 2D-VLMs input formats. We propose two modules: the Perspective Adaptation Module, which dynamically selects viewpoints for query-relevant image rendering, and the Fusion Alignment Module, which integrates 2D images with 3D spatial descriptions to enhance object localization. Extensive experiments on ScanRefer and Nr3D demonstrate that our approach outperforms existing zero-shot methods by large margins. Notably, we exceed weakly supervised methods and rival some fully supervised ones, outperforming previous SOTA by 7.7% on ScanRefer and 7.1% on Nr3D, showcasing its effectiveness in complex 3DVG tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A zero-training framework that adaptively selects spatial representation extractors per object and per task stage improves real-world robot manipulation success and efficiency over fixed-representation baselines.

  2. Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Video-grounded RL with two-stage 2D grounding and SAM2 lifting enables 3D object localization and QA without dense 3D instance supervision.

  3. AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making

    cs.RO 2025-06 conditional novelty 6.0 of 10

    AntiGrounding lifts candidate robot trajectories into the VLM's visual space via multi-view rendering and structured VQA, and reports 57.5% average success across eight manipulation tasks, beating three intermediate-r...

  4. SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Proposal-guided multi-view projection with iterative VLM selection achieves 55.6% and 53.2% Acc@0.25 on ScanRefer and Nr3D, a new zero-shot 3D visual grounding state of the art.

  5. Zero-Shot 3D Visual Grounding from Vision-Language Models

    cs.CV 2025-05 conditional novelty 5.0 of 10

    SeeGround localizes objects in 3D scenes from natural language without 3D-specific training, using query-aligned rendered views and spatially enriched text fed to a 2D vision-language model.

Pith tools