CL3DOR pairs 8,192-point inputs, GPT-4o-generated hard-negative response triplets, and an odds-ratio contrastive loss to achieve state-of-the-art results on 3D scene understanding benchmarks.
3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds
CL3DOR pairs 8,192-point inputs, GPT-4o-generated hard-negative response triplets, and an odds-ratio contrastive loss to achieve state-of-the-art results on 3D scene understanding benchmarks.