A per-scene point cloud encoder trained with contrastive learning on one exploration sequence, ensembled with a vision-language model, associates relocated objects at 95.6% in AI2THOR and 100% in limited real-world tests.
Objects can move: 3d change detection by geometric transformation consistency,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.RO 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
contest 1representative citing papers
citing papers explorer
-
BYE: Build Your Encoder with One Sequence of Exploration Data for Long-Term Dynamic Scene Understanding
A per-scene point cloud encoder trained with contrastive learning on one exploration sequence, ensembled with a vision-language model, associates relocated objects at 95.6% in AI2THOR and 100% in limited real-world tests.