CL3DOR pairs 8,192-point inputs, GPT-4o-generated hard-negative response triplets, and an odds-ratio contrastive loss to achieve state-of-the-art results on 3D scene understanding benchmarks.
Grit-vlp: Grouped mini-batch sampling for efficient vision and language pre-training
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds
CL3DOR pairs 8,192-point inputs, GPT-4o-generated hard-negative response triplets, and an odds-ratio contrastive loss to achieve state-of-the-art results on 3D scene understanding benchmarks.