A per-scene point cloud encoder trained with contrastive learning on one exploration sequence, ensembled with a vision-language model, associates relocated objects at 95.6% in AI2THOR and 100% in limited real-world tests.
Mid-fusion: Octree-based object-level multi-instance dynamic slam,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.RO 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
BYE: Build Your Encoder with One Sequence of Exploration Data for Long-Term Dynamic Scene Understanding
A per-scene point cloud encoder trained with contrastive learning on one exploration sequence, ensembled with a vision-language model, associates relocated objects at 95.6% in AI2THOR and 100% in limited real-world tests.