LAWM-3D learns 3D-aware latent actions from multi-view human videos by combining VGGT geometric alignment with RGB-depth reconstruction, improving robot world model prediction and generalization.
International Conference on Learning Representations , volume=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models
LAWM-3D learns 3D-aware latent actions from multi-view human videos by combining VGGT geometric alignment with RGB-depth reconstruction, improving robot world model prediction and generalization.