Pith. sign in

REVIEW 3 major objections 1 cited by

A single video yields interactive sim twins and cousins that train and rank robot policies like the real world.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 11:27 UTC pith:XMCZQHCP

load-bearing objection Strong modular real-to-sim system with solid ranking correlation and real sim-to-real transfer; abstract overstates real-world gains for scene/task cousins. the 3 major comments →

arxiv 2606.28276 v2 pith:XMCZQHCP submitted 2026-06-26 cs.RO

SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation

classification cs.RO
keywords real-to-simsim-to-realdigital twinsdigital cousinsrobot manipulationpolicy evaluationscene generationarticulated objects
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Training and testing robot manipulation policies in the real world is slow and expensive. This paper introduces SimFoundry, a modular pipeline that turns one real video into a physically interactive simulation of that scene, then automatically expands it into affordance-preserving variants called digital cousins that change objects, layouts, or tasks. The authors show that scores measured in these simulations track real-world policy rankings closely, and that policies trained only on the simulated cousins transfer zero-shot to hard real tasks, including multi-step, articulated, and two-arm work. Adding cousin diversity systematically raises real success rates, so one captured scene becomes a scalable source of training and evaluation data before physical deployment.

Core claim

SimFoundry can rebuild a real tabletop scene from a single video into a sim-ready digital twin, automatically generate object, scene, and task cousins that preserve affordances, and use those environments both to rank policies with mean Pearson correlation 0.911 to real results and to train policies that transfer zero-shot, with cousin training lifting average real success by 17%, 21%, and 40% respectively.

What carries the argument

Digital cousins: affordance-preserving simulated variants of a reconstructed twin along three axes (object instance, scene layout, task specification), produced after an Extraction-Generation-Augmentation pipeline that builds meshes, poses, articulations, and physics annotations from one video.

Load-bearing premise

The pipeline assumes objects rest on a single flat surface, so non-tabletop or multi-level scenes fall outside its current physics-stability and reconstruction design.

What would settle it

Rebuild the same seven tasks and five policies with SimFoundry, measure zero-shot real-world success and sim-real Pearson/MMRV without any sim finetuning; if mean Pearson falls well below ~0.9 or cousin-trained policies no longer beat twin-only training by the reported margins, the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. SimFoundry is a modular real-to-sim pipeline that reconstructs interactive, physics-ready digital twins from a single RGB video and automatically expands them into object, scene, and task “digital cousins.” The paper claims two main results: (i) zero-shot simulation evaluations of real policies strongly predict real-world success across 7 manipulation tasks and 5 policy architectures (mean Pearson 0.911, MMRV 0.018), outperforming PolaRiS under a matched zero-shot protocol; and (ii) policies trained on SimFoundry data transfer zero-shot to real multi-step, articulated, and bimanual tasks, with object/scene/task cousins improving average success by 17%/21%/40%. Supporting evidence includes reconstruction metrics vs SAM3D, multi-embodiment experiments (DROID and YAM), cousin ablations, and multi-task transfer results.

Significance. If the claims hold under corrected scoping, this is a substantial systems contribution for robot learning. Unifying automated real-to-sim reconstruction, articulated assets, multi-axis cousin generation, predictive policy evaluation, and sim-to-real training in one modular stack addresses a genuine bottleneck. The real-to-sim correlation result is particularly valuable: mean Pearson 0.911 and MMRV 0.018 across diverse policies and tasks, with a clear zero-shot comparison to PolaRiS, is stronger and more carefully controlled than much prior work. Demonstrations on multi-step, articulated, and bimanual tasks on two embodiments further raise the bar relative to pick-and-place-only real-to-sim systems. The modular foundation-model design and explicit reconstruction/throughput analysis are practical strengths that make the system reusable as components improve.

major comments (3)
  1. Abstract and §1 overstate the real-world impact of scene and task cousins. The abstract states that when evaluating sim-trained policies zero-shot in the real world, object/scene/task cousins yield average success improvements of 17%, 21%, and 40%. Object-cousin real-world gains are supported (Fig. 5A; Table G.4; App. G.2). Scene- and task-cousin numbers are not: Fig. 5B/C are labeled “(DROID, sim)”, Tables G.5–G.6 report simulation success only, and §5.2/App. G.2 attribute the large scene/task boosts to simulation (e.g., +13 task cousins: Store Marker 20→60, Throw Away Trash 8→68 in sim). Real multi-task transfer is a separate, smaller effect (Table 2: up to ~18% real on seen tasks). This mis-scoping is load-bearing for the dual claim that cousins both train transferable policies and deliver those large average real gains. Please either (a) rewrite the abstract/intro to attribute 21% an
  2. §5.1 and App. J: the real-to-sim evaluation claim is strong, but the manuscript should more clearly separate fully automatic reconstruction from the “few minutes of interactive pose tuning” used for robotics experiments (FAQ B.3; App. K.2). Table L.2 shows zero-shot F1 of 0.81–0.92 rising to 0.93–0.99 with 3 min/object tuning, and App. K.2 states interactive refinement is “primarily” used for real-to-sim experiments. For the headline Pearson 0.911 / MMRV 0.018 result, state explicitly whether evaluation scenes were zero-shot automatic or human-refined, and if refined, report a zero-shot (no pose edit) correlation ablation on at least a subset of tasks. Without that, it is hard to know how much of the correlation depends on manual alignment rather than the automated pipeline.
  3. §5.2 / Fig. 5 / Table G.4: object-cousin real-world gains are uneven and sometimes small on the twin object itself (e.g., Stack Dishware YAM Real Twin 39→43; Store Marker DROID Real Twin 4→20; Throw Away Trash YAM Real Twin 0→28). The abstract’s “average … 17%” aggregates twin and held-out cousin evals across tasks/embodiments. Please report the averaging protocol explicitly (which cells enter the 17% mean; twin-only vs held-out; absolute percentage points vs relative), and avoid implying uniform real-world gains. A short table of per-task real deltas for object cousins would make the claim falsifiable and proportionate.

Circularity Check

0 steps flagged

No circular derivation: headline Pearson/MMRV and cousin gains are measured against independent real (or sim) rollouts, not forced by construction from fitted inputs.

full rationale

SimFoundry is an empirical systems paper. The load-bearing claims are (i) sim vs. real policy success agreement (mean Pearson 0.911, MMRV 0.018 across 7 tasks / 5 architectures) and (ii) success-rate lifts from training with object/scene/task cousins. Both are computed from separate evaluation rollouts: real-world binary task success is measured on physical robots; simulation success is measured in reconstructed scenes; Pearson and MMRV then compare those two score vectors (Appendix J.2). Nothing in the pipeline fits a free parameter to the real success rates and then re-labels that fit as a prediction. Reconstruction F1/Chamfer metrics are likewise scored against quasi-ground-truth poses from FoundationPose on staged YCB scenes, not against quantities defined by the reconstruction itself. Digital cousins are generated by external VLMs/image models and physics predicates, then ablated by training policies with vs. without them; the reported percentage improvements are empirical deltas, not identities. Overlapping-author citations (e.g., ACDC digital cousins, MimicGen) supply prior tooling and terminology but do not force the measured correlations or transfer rates. The abstract’s possible over-scoping of scene/task-cousin gains to real-world zero-shot transfer is a presentation/correctness issue, not circularity: those numbers still come from measured success rates, not from a self-definitional reduction. No self-definitional loop, fitted-input-as-prediction, uniqueness import, or ansatz smuggling is present in the derivation chain.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 2 invented entities

Load-bearing content is an engineering pipeline plus empirical transfer/correlation claims. It rests on standard robotics assumptions (imitation learning, physics simulators, foundation models as black boxes) and domain restrictions (tabletop, video capture quality) rather than new physical laws or fitted universal constants. Free parameters are operational thresholds and training choices; invented terminology largely extends prior 'digital cousins' language.

free parameters (4)
  • Articulation critic acceptance threshold
    Actor-critic VLM loop accepts joint parameters when the critic score exceeds an unspecified threshold; affects articulated asset quality.
  • Per-object interactive tuning budget (~3 min)
    Optional human pose/scale edits improve F1 from 0.81–0.92 to 0.93–0.99; used for some robotics scenes.
  • Policy training hyperparameters (lr 1e-5, 10k steps DROID; 40k YAM; batch 256)
    Checkpoint selection via periodic sim eval; influences reported success rates.
  • Number of cousins and demos (e.g., +9 object cousins, 13 task cousins, ~10 human demos + MimicGen)
    Data-mix sizes chosen by authors; ablations show sensitivity of transfer gains.
axioms (5)
  • domain assumption Off-the-shelf foundation models for depth, segmentation, mesh generation, pose, and VLMs produce usable assets without task-specific training.
    Core of Extraction/Generation stages; Limitations notes inherited failure modes and nondeterminism.
  • domain assumption Objects rest on a single flat reference surface so PyBullet depenetration yields a stable tabletop configuration.
    Stated in Limitations and E.4; restricts multi-level scenes.
  • domain assumption End-to-end binary task success is a valid primary metric for real-to-sim fidelity.
    FAQ and preliminaries defend binary success over normalized reward; used for all Pearson/MMRV numbers.
  • ad hoc to paper Affordance-preserving object/scene/task edits (digital cousins) preserve task-relevant semantics while adding useful diversity.
    Operationalized via VLM prompts and spatial predicates; gains are empirical, not proven.
  • domain assumption Standard imitation learning / flow-matching / VLA finetuning transfers when visual and geometric gaps are small enough.
    Background assumption for all sim-to-real claims.
invented entities (2)
  • SimFoundry modular real-to-sim pipeline independent evidence
    purpose: Unify extraction, generation, physics annotation, and cousin augmentation for both policy eval and training.
    System name for the composed pipeline; not a physical entity. Independent evidence is the reported experiments and project page.
  • Object / scene / task digital cousins (as three automated axes) no independent evidence
    purpose: Scale diversity beyond a single twin while preserving affordances.
    Extends prior digital-cousins concept [17]; automation and task-cousin axis are paper-specific. Evidence is ablation gains, not external measurement of the construct itself.

pith-pipeline@v1.1.0-grok45 · 43342 in / 3281 out tokens · 36543 ms · 2026-07-12T11:27:52.335185+00:00 · methodology

0 comments
read the original abstract

Training and evaluating robot policies in the real world is costly and difficult to scale. We introduce SimFoundry, a modular and automated system for zero-shot real-to-sim scene construction from a video. SimFoundry generates sim-ready digital twins and supports object, scene, and task editing, enabling the automated generation of diverse digital cousins: affordance-preserving variations of reconstructed real-world scenes. Policies trained on SimFoundry data transfer zero-shot to challenging real tasks involving multi-step manipulation, articulated object interaction, and bimanual interaction, and its digital cousins (variations of the original scene, objects, and tasks) facilitate generalization to new real-world conditions. Across 7 manipulation tasks and 5 policy architectures, SimFoundry simulation evaluations strongly predict real-world performance, with mean Pearson correlation 0.911 and mean maximum ranking violation 0.018. When evaluating sim-trained policies zero-shot in the real world, policies trained with object, scene, and task cousins in simulation show average task success rate improvements of 17%, 21%, and 40%, respectively. Additional details at https://research.nvidia.com/labs/gear/simfoundry/ .

Figures

Figures reproduced from arXiv: 2606.28276 by Ajay Mandlekar, Bowen Wen, Brandon Huynh, Danfei Xu, Hang Yin, Josiah Wong, Li Fei-Fei, Linxi Fan, Masoud Moghani, Nadun Ranawaka, Ruohan Zhang, Tianyuan Dai, Wei-Lin Pai, Wei-Teng Chu, Wesley Durbano, Yu Fang, Yuke Zhu, Yunfan Jiang.

Figure 1
Figure 1. Figure 1: Overview. SimFoundry takes a single real-world input video and automatically reconstructs an interactive, sim-ready digital twin of the scene. Based on the reconstructed digital twin, SimFoundry can further generate an unlimited number of digital cousins — affordance-preserving variants of the original scene, spanning three different axes of variation, which we term object, scene, and task cousins, respect… view at source ↗
Figure 2
Figure 2. Figure 2: Method Overview. SimFoundry extracts per-object relevant information (segmentation masks, depth, etc.), generates 3D visual meshes via 2D-to-3D generation models, and compiles the final output scene by annotating relevant physical parameters and sanity checking the overall scene configuration in a physics simulator. SimFoundry additionally supports diverse simulated augmentations along these axes of variat… view at source ↗
Figure 3
Figure 3. Figure 3: SimFoundry Scene Generation Samples. We show real-world input images (top row), the corre￾sponding reconstructed digital twins generated by SimFoundry (middle), and sampled digital cousin scene variations (bottom). For instance, in the scene that is second from the left, the brown glass bottle becomes narrower for the cousin scene, and in the scene that is third from the right, the digital cousin of the wi… view at source ↗
Figure 4
Figure 4. Figure 4: Tasks and Real-to-Sim Policy Evaluation correlations. (Left) We apply SimFoundry to a DROID setup using a single Franka arm (top two rows), and a bimanual setup with two YAM arms (bottom row). Our tasks span multiple types of manipulation, including multi-step, articulated object interaction, and bimanual coordination (Clear Table not shown, more details in Appendix I). (Right) SimFoundry outperforms the s… view at source ↗
Figure 5
Figure 5. Figure 5: SimFoundry Data Diversity Improves Policy Performance. (A) Across multiple robot embodiments and multiple tasks, leveraging additional object cousins [17] improves direct Sim-to-Real policy transfer on the original target scene objects and additional held-out unseen objects. (B) Scene cousins improve policy performance on the original scene and allow policy transfer to cousin scenes. (C) Adding task cousin… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents

    cs.RO 2026-07 conditional novelty 6.0

    Agentic Real2Sim automates real-to-sim conversion of robot interaction episodes using vision-language agents, achieving a 48% VLM-judged replay success rate on DROID-100 with a 31B open-weight model.

Reference graph

Works this paper leans on

139 extracted references · 41 linked inside Pith · cited by 1 Pith paper

  1. [1]

    A bayesian treatment of real-to-sim for deformable object manipulation.IEEE Robotics and Automation Letters, 7(3):5819–5826, 2022

    Rika Antonova, Jingyun Yang, Priya Sundaresan, Dieter Fox, Fabio Ramos, and Jeannette Bohg. A bayesian treatment of real-to-sim for deformable object manipulation.IEEE Robotics and Automation Letters, 7(3):5819–5826, 2022. 3

  2. [2]

    Scan2cad: Learning cad model alignment in rgb-d scans

    Armen Avetisyan, Manuel Dahnert, Angela Dai, Manolis Savva, Angel X Chang, and Matthias Nießner. Scan2cad: Learning cad model alignment in rgb-d scans. InProceedings of the IEEE/CVF Conference on computer vision and pattern recognition, pages 2614–2623, 2019. 3

  3. [3]

    Reliable and scalable robot policy evaluation with imperfect simulators.arXiv preprint arXiv:2510.04354, 2025

    Apurva Badithela, David Snyder, Lihan Zha, Joseph Mikhail, Matthew O’Kelly, Anushri Dixit, and Anirudha Majumdar. Reliable and scalable robot policy evaluation with imperfect simulators.arXiv preprint arXiv:2510.04354, 2025. 2, 3, 21

  4. [4]

    A careful examination of large behavior models for multitask dexterous manipulation.arXiv preprint arXiv:2507.05331, 2025

    Jose Barreiros, Andrew Beaulieu, Aditya Bhat, Rick Cory, Eric Cousineau, Hongkai Dai, Ching-Hsin Fang, Kunimatsu Hashimoto, Muhammad Zubair Irshad, Masha Itkina, et al. A careful examination of large behavior models for multitask dexterous manipulation.arXiv preprint arXiv:2507.05331, 2025. 2

  5. [5]

    2, 7, 21

    Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, et al.𝜋0: A vision-language-action flow model for general robot control.arXiv preprint arXiv:2410.24164, 2024. 2, 7, 21

  6. [6]

    Rt-1: Robotics transformer for real- world control at scale.arXiv preprint arXiv:2212.06817, 2022

    Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, et al. Rt-1: Robotics transformer for real- world control at scale.arXiv preprint arXiv:2212.06817, 2022. 2, 21

  7. [7]

    Rt-2: Vision-language-action models transfer web knowledge to robotic control.arXiv preprint arXiv:2307.15818, 2023

    Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control.arXiv preprint arXiv:2307.15818, 2023. 21

  8. [8]

    Sauser, Darwin G

    Sylvain Calinon, Florent D’halluin, Eric L. Sauser, Darwin G. Caldwell, and Aude Billard. Learning and reproduction of gestures by imitation.IEEE Robotics and Automation Magazine, 17, 2010. 21

  9. [9]

    Berk Calli, Arjun Singh, Aaron Walsman, Siddhartha Srinivasa, Pieter Abbeel, and Aaron M. Dollar. The ycbobjectandmodelset: Towardscommonbenchmarksformanipulationresearch. In2015International Conference on Advanced Robotics (ICAR), pages 510–517, 2015. doi: 10.1109/ICAR.2015.7251504. 47

  10. [10]

    Sam 3: Segment anything with concepts, 2025

    Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Didac Suris, Chai- tanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, Jie Lei, Tengyu Ma, Baishan Guo, Arpit Kalla, Markus Marks, Joseph Greer, Meng Wang, Peize Sun, Roman Rädle, Triantafyllos Afouras, Effrosyni Mavroudi, Katherine Xu, Tsung-Han Wu, Yu Zhou, Lilian...

  11. [11]

    Freeart3d: Training-free articulated object generation using 3d diffusion

    Chuhao Chen, Isabella Liu, Xinyue Wei, Hao Su, and Minghua Liu. Freeart3d: Training-free articulated object generation using 3d diffusion. InProceedings of the SIGGRAPH Asia 2025 Conference Papers, pages 1–13, 2025. 3 10 SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation

  12. [12]

    Urdformer: A pipeline for constructing articulated simulation environ- ments from real-world images.arXiv preprint arXiv:2405.11656, 2024

    Zoey Chen, Aaron Walsman, Marius Memmel, Kaichun Mo, Alex Fang, Karthikeya Vemuri, Alan Wu, Dieter Fox, and Abhishek Gupta. Urdformer: A pipeline for constructing articulated simulation environ- ments from real-world images.arXiv preprint arXiv:2405.11656, 2024. 3

  13. [13]

    Generalizable domain adaptation for sim-and-real policy co-training.arXiv preprint arXiv:2509.18631, 2025

    Shuo Cheng, Liqian Ma, Zhenyang Chen, Ajay Mandlekar, Caelan Garrett, and Danfei Xu. Generalizable domain adaptation for sim-and-real policy co-training.arXiv preprint arXiv:2509.18631, 2025. 2, 21

  14. [14]

    Embodiedsplat: Personalized real-to-sim-to-real navigation with gaussian splats from a mobile device

    GunjanChhablani,XiaomengYe,MuhammadZubairIrshad,andZsoltKira. Embodiedsplat: Personalized real-to-sim-to-real navigation with gaussian splats from a mobile device. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 25431–25441, 2025. 3, 21

  15. [15]

    Diffusion policy: Visuomotor policy learning via action diffusion.The Int’l Journal of Robotics Research, 2023

    Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song. Diffusion policy: Visuomotor policy learning via action diffusion.The Int’l Journal of Robotics Research, 2023. 21

  16. [16]

    Pybullet, a python module for physics simulation for games, robotics and machine learning.http://pybullet.org, 2016–2021

    Erwin Coumans and Yunfei Bai. Pybullet, a python module for physics simulation for games, robotics and machine learning.http://pybullet.org, 2016–2021. 5

  17. [17]

    Automated creation of digital cousins for robust policy learning.arXiv preprint arXiv:2410.07408, 2024

    Tianyuan Dai, Josiah Wong, Yunfan Jiang, Chen Wang, Cem Gokmen, Ruohan Zhang, Jiajun Wu, and Li Fei-Fei. Automated creation of digital cousins for robust policy learning.arXiv preprint arXiv:2410.07408, 2024. 2, 3, 4, 8, 21, 36

  18. [18]

    Imitating task and motion planning with visuomotor transformers

    Murtaza Dalal, Ajay Mandlekar, Caelan Reed Garrett, Ankur Handa, Ruslan Salakhutdinov, and Dieter Fox. Imitating task and motion planning with visuomotor transformers. InConf on Robot Learning,

  19. [19]

    X-sim: Cross-embodiment learning via real-to-sim-to-real.arXiv preprint arXiv:2505.07096, 2025

    Prithwish Dan, Kushal Kedia, Angela Chao, Edward Weiyi Duan, Maximus Adrian Pace, Wei-Chiu Ma, and Sanjiban Choudhury. X-sim: Cross-embodiment learning via real-to-sim-to-real.arXiv preprint arXiv:2505.07096, 2025. 2, 3

  20. [20]

    Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets

    Frederik Ebert, Yanlai Yang, Karl Schmeckpeper, Bernadette Bucher, Georgios Georgakis, Kostas Dani- ilidis, Chelsea Finn, and Sergey Levine. Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets. InRobotics: Science and Systems, 2022. 2, 21

  21. [21]

    Gaussgym: An open-source real-to-sim framework for learning locomotion from pixels

    Alejandro Escontrela, Justin Kerr, Arthur Allshire, Jonas Frey, Rocky Duan, Carmelo Sferrazza, and Pieter Abbeel. Gaussgym: An open-source real-to-sim framework for learning locomotion from pixels. arXiv preprint arXiv:2510.15352, 2025. 3, 21

  22. [22]

    Diffcad: Weakly-supervised probabilistic cad model retrieval and alignment from an rgb image.ACM Transactions on Graphics (TOG), 43(4):1–15, 2024

    Daoyi Gao, Dávid Rozenberszki, Stefan Leutenegger, and Angela Dai. Diffcad: Weakly-supervised probabilistic cad model retrieval and alignment from an rgb image.ACM Transactions on Graphics (TOG), 43(4):1–15, 2024. 3

  23. [23]

    Skillmimicgen: Automated demonstration generation for efficient skill learning and deployment.arXiv preprint arXiv:2410.18907, 2024

    Caelan Garrett, Ajay Mandlekar, Bowen Wen, and Dieter Fox. Skillmimicgen: Automated demonstration generation for efficient skill learning and deployment.arXiv preprint arXiv:2410.18907, 2024. 2, 21

  24. [24]

    Igen: Scalable data generation for robot learning from open-world images.arXiv preprint arXiv:2512.01773, 2025

    Chenghao Gu, Haolan Kang, Junchao Lin, Jinghe Wang, Duo Wu, Shuzhao Xie, Fanding Huang, Junchen Ge, Ziyang Gong, Letian Li, et al. Igen: Scalable data generation for robot learning from open-world images.arXiv preprint arXiv:2512.01773, 2025. 3, 21

  25. [25]

    Roca: Robust cad model retrieval and alignment from a single image

    Can Gümeli, Angela Dai, and Matthias Nießner. Roca: Robust cad model retrieval and alignment from a single image. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4022–4031, 2022. 3

  26. [26]

    Point bridge: 3d representations for cross domain policy learning.arXiv preprint arXiv:2601.16212, 2026

    Siddhant Haldar, Lars Johannsmeier, Lerrel Pinto, Abhishek Gupta, Dieter Fox, Yashraj Narang, and Ajay Mandlekar. Point bridge: 3d representations for cross domain policy learning.arXiv preprint arXiv:2601.16212, 2026. 2, 21 11 SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation

  27. [27]

    Re3sim: Generating high-fidelity simulation data via 3d-photorealistic real-to-sim for robotic manipulation.arXiv preprint arXiv:2502.08645, 2025

    Xiaoshen Han, Minghuan Liu, Yilun Chen, Junqiu Yu, Xiaoyang Lyu, Yang Tian, Bolun Wang, Weinan Zhang, and Jiangmiao Pang. Re3sim: Generating high-fidelity simulation data via 3d-photorealistic real-to-sim for robotic manipulation.arXiv preprint arXiv:2502.08645, 2025. 2, 3, 21

  28. [28]

    Lrm: Large reconstruction model for single image to 3d.arXiv preprint arXiv:2311.04400, 2023

    Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan. Lrm: Large reconstruction model for single image to 3d.arXiv preprint arXiv:2311.04400, 2023. 3

  29. [29]

    Hunyuan3d 2.1: From images to high-fidelity 3d assets with production-ready pbr material, 2025

    Team Hunyuan3D, Shuhui Yang, Mingxin Yang, Yifei Feng, Xin Huang, Sheng Zhang, Zebin He, Di Luo, Haolin Liu, Yunfei Zhao, Qingxiang Lin, Zeqiang Lai, Xianghui Yang, Huiwen Shi, Zibo Zhao, Bowen Zhang, Hongyu Yan, Lifu Wang, Sicong Liu, Jihong Zhang, Meng Chen, Liang Dong, Yiwen Jia, Yulin Cai, Jiaao Yu, Yixuan Tang, Dongyuan Guo, Junlin Yu, Hao Zhang, Zhe...

  30. [30]

    Movement imitation with nonlinear dynamical systems in humanoid robots.Proceedings 2002 IEEE Int’l Conf on Robotics and Automation, 2, 2002

    Auke Jan Ijspeert, Jun Nakanishi, and Stefan Schaal. Movement imitation with nonlinear dynamical systems in humanoid robots.Proceedings 2002 IEEE Int’l Conf on Robotics and Automation, 2, 2002. 21

  31. [31]

    Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, et al.𝜋0.5: a vision-language-action model with open-world generalization.arXiv preprint arXiv:2504.16054, 2025. 7, 8

  32. [32]

    Polaris: Scalable real-to-sim evaluations for generalist robot policies.arXiv preprint arXiv:2512.16881, 2025

    Arhan Jain, Mingtong Zhang, Kanav Arora, William Chen, Marcel Torne, Muhammad Zubair Irshad, Sergey Zakharov, Yue Wang, Sergey Levine, Chelsea Finn, et al. Polaris: Scalable real-to-sim evaluations for generalist robot policies.arXiv preprint arXiv:2512.16881, 2025. 2, 3, 4, 7, 8, 21, 34, 42, 43, 44

  33. [33]

    RobotArena∞: Scalable robot benchmarking via real-to-sim translation

    Yash Jangir, Yidi Zhang, Kashu Yamazaki, Chenyu Zhang, Kuan-Hsun Tu, Tsung-Wei Ke, Lei Ke, Yonatan Bisk, and Katerina Fragkiadaki. RobotArena∞: Scalable robot benchmarking via real-to-sim translation. arXiv preprint arXiv:2510.23571, 2025. 2, 3, 21

  34. [34]

    Gsworld: Closed-loop photo-realistic simulation suite for robotic manipulation

    Guangqi Jiang, Haoran Chang, Ri-Zhao Qiu, Yutong Liang, Mazeyu Ji, Jiyue Zhu, Zhao Dong, Xueyan Zou, and Xiaolong Wang. Gsworld: Closed-loop photo-realistic simulation suite for robotic manipulation. arXiv preprint arXiv:2510.20813, 2025. 2, 3, 21

  35. [35]

    Behavior robot suite: Streamlining real-world whole-body manipulation for everyday household activities.arXiv preprint arXiv:2503.05652, 2025

    Yunfan Jiang, Ruohan Zhang, Josiah Wong, Chen Wang, Yanjie Ze, Hang Yin, Cem Gokmen, Shuran Song, Jiajun Wu, and Li Fei-Fei. Behavior robot suite: Streamlining real-world whole-body manipulation for everyday household activities.arXiv preprint arXiv:2503.05652, 2025. 39, 40

  36. [36]

    Ditto: Building digital twins of articulated objects from interaction

    Zhenyu Jiang, Cheng-Chun Hsu, and Yuke Zhu. Ditto: Building digital twins of articulated objects from interaction. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5616–5626, 2022. 3

  37. [37]

    Dexmimicgen: Automated data generation for bimanual dexterous manipulation via imitation learning

    Zhenyu Jiang, Yuqi Xie, Kevin Lin, Zhenjia Xu, Weikang Wan, Ajay Mandlekar, Linxi Fan, and Yuke Zhu. Dexmimicgen: Automated data generation for bimanual dexterous manipulation via imitation learning. arXiv preprint arXiv:2410.24185, 2024. 2, 4, 21

  38. [38]

    3d gaussian splatting for real-time radiance field rendering.ACM Trans

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George Drettakis, et al. 3d gaussian splatting for real-time radiance field rendering.ACM Trans. Graph., 42(4):139–1, 2023. 6, 24, 25

  39. [39]

    Droid: A large- scale in-the-wild robot manipulation dataset.arXiv preprint arXiv:2403.12945, 2024

    AlexanderKhazatsky, KarlPertsch, SurajNair, AshwinBalakrishna, SudeepDasari, SiddharthKaramcheti, Soroush Nasiriany, Mohan Kumar Srirama, Lawrence Yunliang Chen, Kirsty Ellis, et al. Droid: A large- scale in-the-wild robot manipulation dataset.arXiv preprint arXiv:2403.12945, 2024. 2, 6, 21, 40 12 SimFoundry: Modular and Automated Scene Generation for Pol...

  40. [40]

    Molmospaces: A large-scale open ecosystem for robot navigation and manipulation, 2026

    Yejin Kim, Wilbert Pumacay, Omar Rayyan, Max Argus, Winson Han, Eli VanderBilt, Jordi Salvador, Abhay Deshpande, Rose Hendrix, Snehal Jauhri, Shuo Liu, Nur Muhammad Mahi Shafiullah, Maya Guru, Ainaz Eftekhar, Karen Farley, Donovan Clay, Jiafei Duan, Arjun Guru, Piper Wolters, Alvaro Herrasti, Ying-Chun Lee, Georgia Chalvatzaki, Yuchen Cui, Ali Farhadi, Di...

  41. [41]

    Mask2cad: 3d shape prediction by learning to segment and retrieve

    Weicheng Kuo, Anelia Angelova, Tsung-Yi Lin, and Angela Dai. Mask2cad: 3d shape prediction by learning to segment and retrieve. InEuropean Conference on Computer Vision, pages 260–277. Springer,

  42. [42]

    Patch2cad: Patchwise embedding learning for in-the-wild shape retrieval from a single image

    Weicheng Kuo, Anelia Angelova, Tsung-Yi Lin, and Angela Dai. Patch2cad: Patchwise embedding learning for in-the-wild shape retrieval from a single image. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 12589–12599, 2021. 3

  43. [43]

    Articulate-anything: Automatic modeling of articulated objects via a vision-language foundation model.arXiv preprint arXiv:2410.13882, 2024

    Long Le, Jason Xie, William Liang, Hung-Ju Wang, Yue Yang, Yecheng Jason Ma, Kyle Vedder, Arjun Krishna, Dinesh Jayaraman, and Eric Eaton. Articulate-anything: Automatic modeling of articulated objects via a vision-language foundation model.arXiv preprint arXiv:2410.13882, 2024. 3, 5, 23

  44. [44]

    Any6d: Model-free 6d pose estimation of novel objects.CVPR, 2025

    Taeyeop Lee, Bowen Wen, Minjun Kang, Gyuree Kang, In So Kweon, and Kuk-Jin Yoon. Any6d: Model-free 6d pose estimation of novel objects.CVPR, 2025. 3

  45. [45]

    Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation

    Chengshu Li, Ruohan Zhang, Josiah Wong, Cem Gokmen, Sanjana Srivastava, Roberto Martín-Martín, Chen Wang, Gabrael Levine, Michael Lingelbach, Jiankai Sun, et al. Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation. InConference on Robot Learning, pages 80–93. PMLR, 2023. 29

  46. [46]

    Momagen: Generating demonstrations under soft and hard constraints for multi-step bimanual mobile manipulation

    Chengshu Li, Mengdi Xu, Arpit Bahety, Hang Yin, Yunfan Jiang, Huang Huang, Josiah Wong, Sujay Garlanka, Cem Gokmen, Ruohan Zhang, Weiyu Liu, Jiajun Wu, Roberto Martín-Martín, and Li Fei-Fei. Momagen: Generating demonstrations under soft and hard constraints for multi-step bimanual mobile manipulation. InRSS 2025 Workshop on Whole-body Control and Bimanual...

  47. [47]

    Evaluatingreal-worldrobotmanipulationpoliciesinsimulation

    Xuanlin Li, Kyle Hsu, Jiayuan Gu, Karl Pertsch, Oier Mees, Homer Rich Walke, Chuyuan Fu, Ishikaa Lunawat,IsabelSieh,SeanKirmani,etal. Evaluatingreal-worldrobotmanipulationpoliciesinsimulation. arXiv preprint arXiv:2405.05941, 2024. 2, 3, 4, 21, 42

  48. [48]

    Art: Articulated reconstruction transformer.arXiv preprint arXiv:2512.14671, 2025

    Zizhang Li, Cheng Zhang, Zhengqin Li, Henry Howard-Jenkins, Zhaoyang Lv, Chen Geng, Jiajun Wu, Richard Newcombe, Jakob Engel, and Zhao Dong. Art: Articulated reconstruction transformer.arXiv preprint arXiv:2512.14671, 2025. 3

  49. [49]

    Planar robot casting with real2sim2real self-supervised learning

    Vincent Lim, Huang Huang, Lawrence Yunliang Chen, Jonathan Wang, Jeffrey Ichnowski, Daniel Seita, Michael Laskey, and Ken Goldberg. Planar robot casting with real2sim2real self-supervised learning. arXiv preprint arXiv:2111.04814, 2021. 3

  50. [50]

    Depth anything 3: Recovering the visual space from any views.arXiv preprint arXiv:2511.10647,

    Haotong Lin, Sili Chen, Junhao Liew, Donny Y Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang. Depth anything 3: Recovering the visual space from any views.arXiv preprint arXiv:2511.10647,

  51. [51]

    Singapo: Single image controlled generation of articulated parts in objects.arXiv preprint arXiv:2410.16499, 2024

    Jiayi Liu, Denys Iliash, Angel X Chang, Manolis Savva, and Ali Mahdavi-Amiri. Singapo: Single image controlled generation of articulated parts in objects.arXiv preprint arXiv:2410.16499, 2024. 3

  52. [52]

    Partfield: Learning 3d feature fields for part segmentation and beyond

    Minghua Liu, Mikaela Angelina Uy, Donglai Xiang, Hao Su, Sanja Fidler, Nicholas Sharp, and Jun Gao. Partfield: Learning 3d feature fields for part segmentation and beyond. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 9704–9715, 2025. 22

  53. [53]

    Grounding dino: Marrying dino with grounded pre-training for open-set object detection

    Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. InEuropean conference on computer vision, pages 38–55. Springer, 2024. 3 13 SimFoundry: Modular and Automated Scene Generation for Policy Learni...

  54. [54]

    P3-sam: Native 3d part segmentation.arXiv preprint arXiv:2509.06784,

    Changfeng Ma, Yang Li, Xinhao Yan, Jiachen Xu, Yunhan Yang, Chunshi Wang, Zibo Zhao, Yanwen Guo, Zhuo Chen, and Chunchao Guo. P3-sam: Native 3d part segmentation.arXiv preprint arXiv:2509.06784,

  55. [55]

    Sim-and-real co-training: A simple recipe for vision-based robotic manipulation.arXiv preprint arXiv:2503.24361, 2025

    Abhiram Maddukuri, Zhenyu Jiang, Lawrence Yunliang Chen, Soroush Nasiriany, Yuqi Xie, Yu Fang, Wenqi Huang, Zu Wang, Zhenjia Xu, Nikita Chernyadev, et al. Sim-and-real co-training: A simple recipe for vision-based robotic manipulation.arXiv preprint arXiv:2503.24361, 2025. 2, 4, 21

  56. [56]

    Roboturk: A crowdsourcing platform for robotic skill learning through imitation

    Ajay Mandlekar, Yuke Zhu, Animesh Garg, Jonathan Booher, Max Spero, Albert Tung, Julian Gao, John Emmons, Anchit Gupta, Emre Orbay, et al. Roboturk: A crowdsourcing platform for robotic skill learning through imitation. InConf on Robot Learning, 2018. 21

  57. [57]

    What matters in learning from offline human demonstrations for robot manipulation.arXiv preprint arXiv:2108.03298, 2021

    Ajay Mandlekar, Danfei Xu, Josiah Wong, Soroush Nasiriany, Chen Wang, Rohun Kulkarni, Li Fei-Fei, Silvio Savarese, Yuke Zhu, and Roberto Martín-Martín. What matters in learning from offline human demonstrations for robot manipulation.arXiv preprint arXiv:2108.03298, 2021. 21

  58. [58]

    Mimicgen: A data generation system for scalable robot learning using human demonstrations.arXiv preprint arXiv:2310.17596, 2023

    Ajay Mandlekar, Soroush Nasiriany, Bowen Wen, Iretiayo Akinola, Yashraj Narang, Linxi Fan, Yuke Zhu, and Dieter Fox. Mimicgen: A data generation system for scalable robot learning using human demonstrations.arXiv preprint arXiv:2310.17596, 2023. 2, 21, 29, 39

  59. [59]

    Isaaclab: Agpu-acceleratedsimulationframework for multi-modal robot learning.arXiv preprint arXiv:2511.04831, 2025

    Mayank Mittal, Pascal Roth, James Tigue, Antoine Richard, Octi Zhang, Peter Du, Antonio Serrano- Munoz,XinjieYao,RenéZurbrügg,NikitaRudin,etal. Isaaclab: Agpu-acceleratedsimulationframework for multi-modal robot learning.arXiv preprint arXiv:2511.04831, 2025. 5, 29

  60. [60]

    Void: Video object and interaction deletion.arXiv preprint arXiv:2604.02296, 2026

    Saman Motamed, William Harvey, Benjamin Klein, Luc Van Gool, Zhuoning Yuan, and Ta-Ying Cheng. Void: Video object and interaction deletion.arXiv preprint arXiv:2604.02296, 2026. 23

  61. [61]

    Gr00t n1: An open foundation model for generalist humanoid robots.arXiv preprint arXiv:2503.14734, 2025

    NVIDIA, Johan Bjorck, Fernando Castañeda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, et al. Gr00t n1: An open foundation model for generalist humanoid robots.arXiv preprint arXiv:2503.14734, 2025. 2, 7, 21, 39

  62. [62]

    Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

    Abby O’Neill, Abdul Rehman, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, et al. Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0. In2024 IEEE Int’l Conf on Robotics and Automation (ICRA), 2024. 21

  63. [63]

    Scenesmith: Agentic generation of simulation-ready indoor scenes, 2026

    Nicholas Pfaff, Thomas Cohn, Sergey Zakharov, Rick Cory, and Russ Tedrake. Scenesmith: Agentic generation of simulation-ready indoor scenes, 2026. 2, 3

  64. [64]

    Alvinn: An autonomous land vehicle in a neural network

    Dean A Pomerleau. Alvinn: An autonomous land vehicle in a neural network. InAdvances in neural information processing systems, 1989. 21

  65. [65]

    Articulate anymesh: Open-vocabulary 3d articulated objects modeling.arXiv preprint arXiv:2502.02590, 2025

    Xiaowen Qiu, Jincheng Yang, Yian Wang, Zhehuan Chen, Yufei Wang, Tsun-Hsuan Wang, Zhou Xian, and Chuang Gan. Articulate anymesh: Open-vocabulary 3d articulated objects modeling.arXiv preprint arXiv:2502.02590, 2025. 5, 23

  66. [66]

    Splatsim: Zero-shot sim2real transfer of rgb manipulation policies using gaussian splatting

    M Nomaan Qureshi, Sparsh Garg, Francisco Yandun, David Held, George Kantor, and Abhisesh Silwal. Splatsim: Zero-shot sim2real transfer of rgb manipulation policies using gaussian splatting. In2025 IEEE International Conference on Robotics and Automation (ICRA), pages 6502–6509. IEEE, 2025. 3

  67. [67]

    Sam 2: Segment anything in images and videos.arXiv preprint arXiv:2408.00714, 2024

    Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, et al. Sam 2: Segment anything in images and videos.arXiv preprint arXiv:2408.00714, 2024. 3

  68. [68]

    Sam 2: Segment anything in images and videos

    Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, et al. Sam 2: Segment anything in images and videos. InInternational Conference on Learning Representations, volume 2025, pages 28085–28128,

  69. [69]

    23 14 SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation

  70. [70]

    Yam robot arm, 2025

    I2RT Robotics. Yam robot arm, 2025. URLhttps://i2rt.com/collections/yam-arm. 6, 40

  71. [71]

    Is imitation learning the route to humanoid robots?Trends in cognitive sciences, 3, 1999

    Stefan Schaal. Is imitation learning the route to humanoid robots?Trends in cognitive sciences, 3, 1999. 21

  72. [72]

    Structure-from-Motion Revisited

    Johannes Lutz Schönberger and Jan-Michael Frahm. Structure-from-Motion Revisited. InConference on Computer Vision and Pattern Recognition (CVPR), 2016. 25

  73. [73]

    Nerfstudio: A modular framework for neural radiance field development

    Matthew Tancik, Ethan Weber, Evonne Ng, Ruilong Li, Brent Yi, Terrance Wang, Alexander Kristoffersen, Jake Austin, Kamyar Salahi, Abhik Ahuja, et al. Nerfstudio: A modular framework for neural radiance field development. InACM SIGGRAPH 2023 conference proceedings, pages 1–12, 2023. 24, 25

  74. [74]

    Segment any mesh.arXiv preprint arXiv:2408.13679, 2024

    George Tang, William Zhao, Logan Ford, David Benhaim, and Paul Zhang. Segment any mesh.arXiv preprint arXiv:2408.13679, 2024. 22, 23

  75. [75]

    Sam 3d: 3dfy anything in images, 2025

    SAM 3D Team, Xingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang, Alexander Sax, Hao Tang, Weiyao Wang, Michelle Guo, Thibaut Hardin, Xiang Li, Aohan Lin, Jiawei Liu, Ziqi Ma, Anushka Sagar, Bowen Song, Xiaodong Wang, Jianing Yang, Bowen Zhang, Piotr Dollár, Georgia Gkioxari, Matt Feiszli, and Jitendra Malik. Sam 3d: 3dfy anything in images, 2025. URLht...

  76. [76]

    Hunyuan3d 2.5: Towards high-fidelity 3d assets generation with ultimate details, 2025

    Tencent Hunyuan3D Team. Hunyuan3d 2.5: Towards high-fidelity 3d assets generation with ultimate details, 2025. URLhttps://arxiv.org/abs/2506.16504. 3

  77. [77]

    Interndata-a1: Pioneering high-fidelity synthetic data for pre-training generalist policy

    Yang Tian, Yuyin Yang, Yiman Xie, Zetao Cai, Xu Shi, Ning Gao, Hangxu Liu, Xuekun Jiang, Zherui Qiu, Feng Yuan, et al. Interndata-a1: Pioneering high-fidelity synthetic data for pre-training generalist policy. arXiv preprint arXiv:2511.16651, 2025. 21

  78. [78]

    Triposr: Fast 3d object reconstruction from a single image.arXiv preprint arXiv:2403.02151, 2024

    Dmitry Tochilkin, David Pankratz, Zexiang Liu, Zixuan Huang, Adam Letts, Yangguang Li, Ding Liang, Christian Laforte, Varun Jampani, and Yan-Pei Cao. Triposr: Fast 3d object reconstruction from a single image.arXiv preprint arXiv:2403.02151, 2024. 3

  79. [79]

    Robot learning with super-linear scaling.arXiv preprint arXiv:2412.01770, 2024

    MarcelTorne, ArhanJain, JiayiYuan, VidaaranyaMacha, LarsAnkile, AnthonySimeonov, PulkitAgrawal, and Abhishek Gupta. Robot learning with super-linear scaling.arXiv preprint arXiv:2412.01770, 2024. 3, 21

  80. [80]

    Reconciling reality through simulation: A real-to-sim-to-real approach for robust manipulation.arXiv preprint arXiv:2403.03949, 2024

    Marcel Torne, Anthony Simeonov, Zechu Li, April Chan, Tao Chen, Abhishek Gupta, and Pulkit Agrawal. Reconciling reality through simulation: A real-to-sim-to-real approach for robust manipulation.arXiv preprint arXiv:2403.03949, 2024. 2, 3, 21

Showing first 80 references.