LSVG improves 3D visual grounding by constructing a task-specific scene graph from the text description and using CLIP-based 2D features to supervise and enrich 3D object encoding.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LSVG: Language-Guided Scene Graphs with 2D-Assisted Multi-Modal Encoding for 3D Visual Grounding
LSVG improves 3D visual grounding by constructing a task-specific scene graph from the text description and using CLIP-based 2D features to supervise and enrich 3D object encoding.