A per-scene point cloud encoder trained with contrastive learning on one exploration sequence, ensembled with a vision-language model, associates relocated objects at 95.6% in AI2THOR and 100% in limited real-world tests.
Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collab- oration 0,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.RO 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
BYE: Build Your Encoder with One Sequence of Exploration Data for Long-Term Dynamic Scene Understanding
A per-scene point cloud encoder trained with contrastive learning on one exploration sequence, ensembled with a vision-language model, associates relocated objects at 95.6% in AI2THOR and 100% in limited real-world tests.