DC-Scene filters 3D scene-caption pairs by CLIP score and caption perplexity, trains on a top-75% subset with a curriculum, and reports higher CIDEr than full-data training at one-third of the epochs.
Referit3d: Neural listeners for fine-grained 3d object identification in real- world scenes
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
dataset 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
REJECT 1roles
dataset 1polarities
use dataset 1representative citing papers
citing papers explorer
-
DC-Scene: Data-Centric Learning for 3D Scene Understanding
DC-Scene filters 3D scene-caption pairs by CLIP score and caption perplexity, trains on a top-75% subset with a curriculum, and reports higher CIDEr than full-data training at one-third of the epochs.