VGLD fuses CLIP image and text embeddings to predict the global scale and shift that convert relative monocular depth into metric depth, outperforming the text-only RSA baseline.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
VGLD: Visually-Guided Linguistic Disambiguation for Monocular Depth Scale Recovery
VGLD fuses CLIP image and text embeddings to predict the global scale and shift that convert relative monocular depth into metric depth, outperforming the text-only RSA baseline.