Imagine2Real enables zero-shot humanoid-object interaction by unifying motions as 4D point trajectories, tracking only base/hands/object keypoints inside a BFM latent space, and training with progressive simple rewards for mocap deployment.
Foundationpose: Unified 6d pose estimation and tracking of novel objects
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
other 1
citation-polarity summary
fields
cs.RO 2years
2026 2roles
other 1polarities
unclear 1representative citing papers
AssemLM fuses SO(3)-equivariant point-cloud features into a VLM to predict discrete 6D assembly poses, reaching ~89% success on AssemBench and improved real-robot multi-step assembly.
citing papers explorer
-
Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors
Imagine2Real enables zero-shot humanoid-object interaction by unifying motions as 4D point trajectories, tracking only base/hands/object keypoints inside a BFM latent space, and training with progressive simple rewards for mocap deployment.
-
AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly
AssemLM fuses SO(3)-equivariant point-cloud features into a VLM to predict discrete 6D assembly poses, reaching ~89% success on AssemBench and improved real-robot multi-step assembly.