An automated real-to-sim pipeline builds digital twins and affordance-preserving cousins from video, yielding sim evaluations that correlate with real robot policy success and zero-shot sim-to-real gains.
Point bridge: 3D representations for cross domain policy learning.arXiv preprint arXiv:2601.16212, 2026
4 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.RO 4years
2026 4roles
background 1polarities
background 1representative citing papers
Dexterous Point Policy learns dexterous hand policies from human videos using 3D keypoints of hands and objects, achieving 75% success on real-robot tasks compared to 1% for a VLA baseline.
HumanoidMimicGen automatically generates large loco-manipulation datasets from few source demonstrations using whole-body planning, enabling visuomotor policies that outperform real-data-only training by 20% on a new nine-task benchmark.
CoEnv introduces a compositional environment that integrates real and simulated spaces for multi-agent robotic collaboration, using real-to-sim reconstruction, VLM action synthesis, and validated sim-to-real transfer to achieve high success rates on multi-arm manipulation tasks.
citing papers explorer
-
SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation
An automated real-to-sim pipeline builds digital twins and affordance-preserving cousins from video, yielding sim evaluations that correlate with real robot policy success and zero-shot sim-to-real gains.
-
Dexterous Point Policy: Learning Point-based Dexterous Hand Policies from Human Demonstrations
Dexterous Point Policy learns dexterous hand policies from human videos using 3D keypoints of hands and objects, achieving 75% success on real-robot tasks compared to 1% for a VLA baseline.
-
HumanoidMimicGen: Data Generation for Loco-Manipulation via Whole-Body Planning
HumanoidMimicGen automatically generates large loco-manipulation datasets from few source demonstrations using whole-body planning, enabling visuomotor policies that outperform real-data-only training by 20% on a new nine-task benchmark.
-
CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment
CoEnv introduces a compositional environment that integrates real and simulated spaces for multi-agent robotic collaboration, using real-to-sim reconstruction, VLM action synthesis, and validated sim-to-real transfer to achieve high success rates on multi-arm manipulation tasks.