TriCLIP-3D encodes point clouds, images, and text with one frozen CLIP model plus adapters, reporting 6.5-point AP25 gains on EmbodiedScan 3D detection and grounding while cutting trainable parameters by 58%.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
TriCLIP-3D: A Unified Parameter-Efficient Framework for Tri-Modal 3D Visual Grounding based on CLIP
TriCLIP-3D encodes point clouds, images, and text with one frozen CLIP model plus adapters, reporting 6.5-point AP25 gains on EmbodiedScan 3D detection and grounding while cutting trainable parameters by 58%.