A CLIP-aligned transformer pretrained on multi-view RGB-Pointmap inputs with cross-view geometric and grounded view alignment yields unified 3D scene features that transfer to several scene-understanding tasks.
arXiv preprint arXiv:2406.11579 (2024) 2, 3
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2026 1verdicts
UNVERDICTED 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
RGB-Pointmap Pretraining for Unified 3D Scene Understanding
A CLIP-aligned transformer pretrained on multi-view RGB-Pointmap inputs with cross-view geometric and grounded view alignment yields unified 3D scene features that transfer to several scene-understanding tasks.