Pith. sign in

REVIEW 2 cited by

NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.22436 v2 pith:ISUEXMWJ submitted 2025-03-28 cs.CV

classification cs.CV
keywords groundingautonomousdrivinginstructionsmulti-viewnugroundingvisualabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-view 3D visual grounding is critical for autonomous driving vehicles to interpret natural languages and localize target objects in complex environments. However, existing datasets and methods suffer from coarse-grained language instructions, and inadequate integration of 3D geometric reasoning with linguistic comprehension. To this end, we introduce NuGrounding, the first large-scale benchmark for multi-view 3D visual grounding in autonomous driving. We present a Hierarchy of Grounding (HoG) method to construct NuGrounding to generate hierarchical multi-level instructions, ensuring comprehensive coverage of human instruction patterns. To tackle this challenging dataset, we propose a novel paradigm that seamlessly combines instruction comprehension abilities of multi-modal LLMs (MLLMs) with precise localization abilities of specialist detection models. Our approach introduces two decoupled task tokens and a context query to aggregate 3D geometric information and semantic instructions, followed by a fusion decoder to refine spatial-semantic feature fusion for precise localization. Extensive experiments demonstrate that our method significantly outperforms the baselines adapted from representative 3D scene understanding methods by a significant margin and achieves 0.59 in precision and 0.64 in recall, with improvements of 50.8% and 54.7%.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Talk2Sensors: 3D Visual Grounding in Autonomous Driving via Sensor-Adaptive Physical Cue Matching

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A tri-sensor camera-LiDAR-radar 3D visual grounding benchmark and a language-routed fusion model, TSFormer, that improves grounding accuracy over prior baselines.

  2. TDVR: Joint Text Disambiguation and Viewpoint Reasoning for Zero-Shot 3D Visual Grounding

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A training-free pipeline that disambiguates text queries and infers viewpoints improves zero-shot 3D visual grounding, reaching 64.06% Acc@0.5 on ScanRefer.

Pith tools