REVIEW 3 major objections 1 cited by
A single video yields interactive sim twins and cousins that train and rank robot policies like the real world.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 11:27 UTC pith:XMCZQHCP
load-bearing objection Strong modular real-to-sim system with solid ranking correlation and real sim-to-real transfer; abstract overstates real-world gains for scene/task cousins. the 3 major comments →
SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
SimFoundry can rebuild a real tabletop scene from a single video into a sim-ready digital twin, automatically generate object, scene, and task cousins that preserve affordances, and use those environments both to rank policies with mean Pearson correlation 0.911 to real results and to train policies that transfer zero-shot, with cousin training lifting average real success by 17%, 21%, and 40% respectively.
What carries the argument
Digital cousins: affordance-preserving simulated variants of a reconstructed twin along three axes (object instance, scene layout, task specification), produced after an Extraction-Generation-Augmentation pipeline that builds meshes, poses, articulations, and physics annotations from one video.
Load-bearing premise
The pipeline assumes objects rest on a single flat surface, so non-tabletop or multi-level scenes fall outside its current physics-stability and reconstruction design.
What would settle it
Rebuild the same seven tasks and five policies with SimFoundry, measure zero-shot real-world success and sim-real Pearson/MMRV without any sim finetuning; if mean Pearson falls well below ~0.9 or cousin-trained policies no longer beat twin-only training by the reported margins, the central claim fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. SimFoundry is a modular real-to-sim pipeline that reconstructs interactive, physics-ready digital twins from a single RGB video and automatically expands them into object, scene, and task “digital cousins.” The paper claims two main results: (i) zero-shot simulation evaluations of real policies strongly predict real-world success across 7 manipulation tasks and 5 policy architectures (mean Pearson 0.911, MMRV 0.018), outperforming PolaRiS under a matched zero-shot protocol; and (ii) policies trained on SimFoundry data transfer zero-shot to real multi-step, articulated, and bimanual tasks, with object/scene/task cousins improving average success by 17%/21%/40%. Supporting evidence includes reconstruction metrics vs SAM3D, multi-embodiment experiments (DROID and YAM), cousin ablations, and multi-task transfer results.
Significance. If the claims hold under corrected scoping, this is a substantial systems contribution for robot learning. Unifying automated real-to-sim reconstruction, articulated assets, multi-axis cousin generation, predictive policy evaluation, and sim-to-real training in one modular stack addresses a genuine bottleneck. The real-to-sim correlation result is particularly valuable: mean Pearson 0.911 and MMRV 0.018 across diverse policies and tasks, with a clear zero-shot comparison to PolaRiS, is stronger and more carefully controlled than much prior work. Demonstrations on multi-step, articulated, and bimanual tasks on two embodiments further raise the bar relative to pick-and-place-only real-to-sim systems. The modular foundation-model design and explicit reconstruction/throughput analysis are practical strengths that make the system reusable as components improve.
major comments (3)
- Abstract and §1 overstate the real-world impact of scene and task cousins. The abstract states that when evaluating sim-trained policies zero-shot in the real world, object/scene/task cousins yield average success improvements of 17%, 21%, and 40%. Object-cousin real-world gains are supported (Fig. 5A; Table G.4; App. G.2). Scene- and task-cousin numbers are not: Fig. 5B/C are labeled “(DROID, sim)”, Tables G.5–G.6 report simulation success only, and §5.2/App. G.2 attribute the large scene/task boosts to simulation (e.g., +13 task cousins: Store Marker 20→60, Throw Away Trash 8→68 in sim). Real multi-task transfer is a separate, smaller effect (Table 2: up to ~18% real on seen tasks). This mis-scoping is load-bearing for the dual claim that cousins both train transferable policies and deliver those large average real gains. Please either (a) rewrite the abstract/intro to attribute 21% an
- §5.1 and App. J: the real-to-sim evaluation claim is strong, but the manuscript should more clearly separate fully automatic reconstruction from the “few minutes of interactive pose tuning” used for robotics experiments (FAQ B.3; App. K.2). Table L.2 shows zero-shot F1 of 0.81–0.92 rising to 0.93–0.99 with 3 min/object tuning, and App. K.2 states interactive refinement is “primarily” used for real-to-sim experiments. For the headline Pearson 0.911 / MMRV 0.018 result, state explicitly whether evaluation scenes were zero-shot automatic or human-refined, and if refined, report a zero-shot (no pose edit) correlation ablation on at least a subset of tasks. Without that, it is hard to know how much of the correlation depends on manual alignment rather than the automated pipeline.
- §5.2 / Fig. 5 / Table G.4: object-cousin real-world gains are uneven and sometimes small on the twin object itself (e.g., Stack Dishware YAM Real Twin 39→43; Store Marker DROID Real Twin 4→20; Throw Away Trash YAM Real Twin 0→28). The abstract’s “average … 17%” aggregates twin and held-out cousin evals across tasks/embodiments. Please report the averaging protocol explicitly (which cells enter the 17% mean; twin-only vs held-out; absolute percentage points vs relative), and avoid implying uniform real-world gains. A short table of per-task real deltas for object cousins would make the claim falsifiable and proportionate.
Circularity Check
No circular derivation: headline Pearson/MMRV and cousin gains are measured against independent real (or sim) rollouts, not forced by construction from fitted inputs.
full rationale
SimFoundry is an empirical systems paper. The load-bearing claims are (i) sim vs. real policy success agreement (mean Pearson 0.911, MMRV 0.018 across 7 tasks / 5 architectures) and (ii) success-rate lifts from training with object/scene/task cousins. Both are computed from separate evaluation rollouts: real-world binary task success is measured on physical robots; simulation success is measured in reconstructed scenes; Pearson and MMRV then compare those two score vectors (Appendix J.2). Nothing in the pipeline fits a free parameter to the real success rates and then re-labels that fit as a prediction. Reconstruction F1/Chamfer metrics are likewise scored against quasi-ground-truth poses from FoundationPose on staged YCB scenes, not against quantities defined by the reconstruction itself. Digital cousins are generated by external VLMs/image models and physics predicates, then ablated by training policies with vs. without them; the reported percentage improvements are empirical deltas, not identities. Overlapping-author citations (e.g., ACDC digital cousins, MimicGen) supply prior tooling and terminology but do not force the measured correlations or transfer rates. The abstract’s possible over-scoping of scene/task-cousin gains to real-world zero-shot transfer is a presentation/correctness issue, not circularity: those numbers still come from measured success rates, not from a self-definitional reduction. No self-definitional loop, fitted-input-as-prediction, uniqueness import, or ansatz smuggling is present in the derivation chain.
Axiom & Free-Parameter Ledger
free parameters (4)
- Articulation critic acceptance threshold
- Per-object interactive tuning budget (~3 min)
- Policy training hyperparameters (lr 1e-5, 10k steps DROID; 40k YAM; batch 256)
- Number of cousins and demos (e.g., +9 object cousins, 13 task cousins, ~10 human demos + MimicGen)
axioms (5)
- domain assumption Off-the-shelf foundation models for depth, segmentation, mesh generation, pose, and VLMs produce usable assets without task-specific training.
- domain assumption Objects rest on a single flat reference surface so PyBullet depenetration yields a stable tabletop configuration.
- domain assumption End-to-end binary task success is a valid primary metric for real-to-sim fidelity.
- ad hoc to paper Affordance-preserving object/scene/task edits (digital cousins) preserve task-relevant semantics while adding useful diversity.
- domain assumption Standard imitation learning / flow-matching / VLA finetuning transfers when visual and geometric gaps are small enough.
invented entities (2)
-
SimFoundry modular real-to-sim pipeline
independent evidence
-
Object / scene / task digital cousins (as three automated axes)
no independent evidence
read the original abstract
Training and evaluating robot policies in the real world is costly and difficult to scale. We introduce SimFoundry, a modular and automated system for zero-shot real-to-sim scene construction from a video. SimFoundry generates sim-ready digital twins and supports object, scene, and task editing, enabling the automated generation of diverse digital cousins: affordance-preserving variations of reconstructed real-world scenes. Policies trained on SimFoundry data transfer zero-shot to challenging real tasks involving multi-step manipulation, articulated object interaction, and bimanual interaction, and its digital cousins (variations of the original scene, objects, and tasks) facilitate generalization to new real-world conditions. Across 7 manipulation tasks and 5 policy architectures, SimFoundry simulation evaluations strongly predict real-world performance, with mean Pearson correlation 0.911 and mean maximum ranking violation 0.018. When evaluating sim-trained policies zero-shot in the real world, policies trained with object, scene, and task cousins in simulation show average task success rate improvements of 17%, 21%, and 40%, respectively. Additional details at https://research.nvidia.com/labs/gear/simfoundry/ .
Figures
Forward citations
Cited by 1 Pith paper
-
Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents
Agentic Real2Sim automates real-to-sim conversion of robot interaction episodes using vision-language agents, achieving a 48% VLM-judged replay success rate on DROID-100 with a 31B open-weight model.
Reference graph
Works this paper leans on
-
[1]
A bayesian treatment of real-to-sim for deformable object manipulation.IEEE Robotics and Automation Letters, 7(3):5819–5826, 2022
Rika Antonova, Jingyun Yang, Priya Sundaresan, Dieter Fox, Fabio Ramos, and Jeannette Bohg. A bayesian treatment of real-to-sim for deformable object manipulation.IEEE Robotics and Automation Letters, 7(3):5819–5826, 2022. 3
2022
-
[2]
Scan2cad: Learning cad model alignment in rgb-d scans
Armen Avetisyan, Manuel Dahnert, Angela Dai, Manolis Savva, Angel X Chang, and Matthias Nießner. Scan2cad: Learning cad model alignment in rgb-d scans. InProceedings of the IEEE/CVF Conference on computer vision and pattern recognition, pages 2614–2623, 2019. 3
2019
-
[3]
Apurva Badithela, David Snyder, Lihan Zha, Joseph Mikhail, Matthew O’Kelly, Anushri Dixit, and Anirudha Majumdar. Reliable and scalable robot policy evaluation with imperfect simulators.arXiv preprint arXiv:2510.04354, 2025. 2, 3, 21
arXiv 2025
-
[4]
Jose Barreiros, Andrew Beaulieu, Aditya Bhat, Rick Cory, Eric Cousineau, Hongkai Dai, Ching-Hsin Fang, Kunimatsu Hashimoto, Muhammad Zubair Irshad, Masha Itkina, et al. A careful examination of large behavior models for multitask dexterous manipulation.arXiv preprint arXiv:2507.05331, 2025. 2
Pith/arXiv arXiv 2025
-
[5]
Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, et al.𝜋0: A vision-language-action flow model for general robot control.arXiv preprint arXiv:2410.24164, 2024. 2, 7, 21
Pith/arXiv arXiv 2024
-
[6]
Rt-1: Robotics transformer for real- world control at scale.arXiv preprint arXiv:2212.06817, 2022
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, et al. Rt-1: Robotics transformer for real- world control at scale.arXiv preprint arXiv:2212.06817, 2022. 2, 21
Pith/arXiv arXiv 2022
-
[7]
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control.arXiv preprint arXiv:2307.15818, 2023. 21
Pith/arXiv arXiv 2023
-
[8]
Sauser, Darwin G
Sylvain Calinon, Florent D’halluin, Eric L. Sauser, Darwin G. Caldwell, and Aude Billard. Learning and reproduction of gestures by imitation.IEEE Robotics and Automation Magazine, 17, 2010. 21
2010
-
[9]
Berk Calli, Arjun Singh, Aaron Walsman, Siddhartha Srinivasa, Pieter Abbeel, and Aaron M. Dollar. The ycbobjectandmodelset: Towardscommonbenchmarksformanipulationresearch. In2015International Conference on Advanced Robotics (ICAR), pages 510–517, 2015. doi: 10.1109/ICAR.2015.7251504. 47
-
[10]
Sam 3: Segment anything with concepts, 2025
Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Didac Suris, Chai- tanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, Jie Lei, Tengyu Ma, Baishan Guo, Arpit Kalla, Markus Marks, Joseph Greer, Meng Wang, Peize Sun, Roman Rädle, Triantafyllos Afouras, Effrosyni Mavroudi, Katherine Xu, Tsung-Han Wu, Yu Zhou, Lilian...
2025
-
[11]
Freeart3d: Training-free articulated object generation using 3d diffusion
Chuhao Chen, Isabella Liu, Xinyue Wei, Hao Su, and Minghua Liu. Freeart3d: Training-free articulated object generation using 3d diffusion. InProceedings of the SIGGRAPH Asia 2025 Conference Papers, pages 1–13, 2025. 3 10 SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation
2025
-
[12]
Zoey Chen, Aaron Walsman, Marius Memmel, Kaichun Mo, Alex Fang, Karthikeya Vemuri, Alan Wu, Dieter Fox, and Abhishek Gupta. Urdformer: A pipeline for constructing articulated simulation environ- ments from real-world images.arXiv preprint arXiv:2405.11656, 2024. 3
Pith/arXiv arXiv 2024
-
[13]
Shuo Cheng, Liqian Ma, Zhenyang Chen, Ajay Mandlekar, Caelan Garrett, and Danfei Xu. Generalizable domain adaptation for sim-and-real policy co-training.arXiv preprint arXiv:2509.18631, 2025. 2, 21
arXiv 2025
-
[14]
Embodiedsplat: Personalized real-to-sim-to-real navigation with gaussian splats from a mobile device
GunjanChhablani,XiaomengYe,MuhammadZubairIrshad,andZsoltKira. Embodiedsplat: Personalized real-to-sim-to-real navigation with gaussian splats from a mobile device. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 25431–25441, 2025. 3, 21
2025
-
[15]
Diffusion policy: Visuomotor policy learning via action diffusion.The Int’l Journal of Robotics Research, 2023
Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song. Diffusion policy: Visuomotor policy learning via action diffusion.The Int’l Journal of Robotics Research, 2023. 21
2023
-
[16]
Pybullet, a python module for physics simulation for games, robotics and machine learning.http://pybullet.org, 2016–2021
Erwin Coumans and Yunfei Bai. Pybullet, a python module for physics simulation for games, robotics and machine learning.http://pybullet.org, 2016–2021. 5
2016
-
[17]
Tianyuan Dai, Josiah Wong, Yunfan Jiang, Chen Wang, Cem Gokmen, Ruohan Zhang, Jiajun Wu, and Li Fei-Fei. Automated creation of digital cousins for robust policy learning.arXiv preprint arXiv:2410.07408, 2024. 2, 3, 4, 8, 21, 36
Pith/arXiv arXiv 2024
-
[18]
Imitating task and motion planning with visuomotor transformers
Murtaza Dalal, Ajay Mandlekar, Caelan Reed Garrett, Ankur Handa, Ruslan Salakhutdinov, and Dieter Fox. Imitating task and motion planning with visuomotor transformers. InConf on Robot Learning,
-
[19]
X-sim: Cross-embodiment learning via real-to-sim-to-real.arXiv preprint arXiv:2505.07096, 2025
Prithwish Dan, Kushal Kedia, Angela Chao, Edward Weiyi Duan, Maximus Adrian Pace, Wei-Chiu Ma, and Sanjiban Choudhury. X-sim: Cross-embodiment learning via real-to-sim-to-real.arXiv preprint arXiv:2505.07096, 2025. 2, 3
arXiv 2025
-
[20]
Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets
Frederik Ebert, Yanlai Yang, Karl Schmeckpeper, Bernadette Bucher, Georgios Georgakis, Kostas Dani- ilidis, Chelsea Finn, and Sergey Levine. Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets. InRobotics: Science and Systems, 2022. 2, 21
2022
-
[21]
Gaussgym: An open-source real-to-sim framework for learning locomotion from pixels
Alejandro Escontrela, Justin Kerr, Arthur Allshire, Jonas Frey, Rocky Duan, Carmelo Sferrazza, and Pieter Abbeel. Gaussgym: An open-source real-to-sim framework for learning locomotion from pixels. arXiv preprint arXiv:2510.15352, 2025. 3, 21
arXiv 2025
-
[22]
Diffcad: Weakly-supervised probabilistic cad model retrieval and alignment from an rgb image.ACM Transactions on Graphics (TOG), 43(4):1–15, 2024
Daoyi Gao, Dávid Rozenberszki, Stefan Leutenegger, and Angela Dai. Diffcad: Weakly-supervised probabilistic cad model retrieval and alignment from an rgb image.ACM Transactions on Graphics (TOG), 43(4):1–15, 2024. 3
2024
-
[23]
Caelan Garrett, Ajay Mandlekar, Bowen Wen, and Dieter Fox. Skillmimicgen: Automated demonstration generation for efficient skill learning and deployment.arXiv preprint arXiv:2410.18907, 2024. 2, 21
Pith/arXiv arXiv 2024
-
[24]
Chenghao Gu, Haolan Kang, Junchao Lin, Jinghe Wang, Duo Wu, Shuzhao Xie, Fanding Huang, Junchen Ge, Ziyang Gong, Letian Li, et al. Igen: Scalable data generation for robot learning from open-world images.arXiv preprint arXiv:2512.01773, 2025. 3, 21
Pith/arXiv arXiv 2025
-
[25]
Roca: Robust cad model retrieval and alignment from a single image
Can Gümeli, Angela Dai, and Matthias Nießner. Roca: Robust cad model retrieval and alignment from a single image. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4022–4031, 2022. 3
2022
-
[26]
Siddhant Haldar, Lars Johannsmeier, Lerrel Pinto, Abhishek Gupta, Dieter Fox, Yashraj Narang, and Ajay Mandlekar. Point bridge: 3d representations for cross domain policy learning.arXiv preprint arXiv:2601.16212, 2026. 2, 21 11 SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation
arXiv 2026
-
[27]
Xiaoshen Han, Minghuan Liu, Yilun Chen, Junqiu Yu, Xiaoyang Lyu, Yang Tian, Bolun Wang, Weinan Zhang, and Jiangmiao Pang. Re3sim: Generating high-fidelity simulation data via 3d-photorealistic real-to-sim for robotic manipulation.arXiv preprint arXiv:2502.08645, 2025. 2, 3, 21
Pith/arXiv arXiv 2025
-
[28]
Lrm: Large reconstruction model for single image to 3d.arXiv preprint arXiv:2311.04400, 2023
Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan. Lrm: Large reconstruction model for single image to 3d.arXiv preprint arXiv:2311.04400, 2023. 3
Pith/arXiv arXiv 2023
-
[29]
Hunyuan3d 2.1: From images to high-fidelity 3d assets with production-ready pbr material, 2025
Team Hunyuan3D, Shuhui Yang, Mingxin Yang, Yifei Feng, Xin Huang, Sheng Zhang, Zebin He, Di Luo, Haolin Liu, Yunfei Zhao, Qingxiang Lin, Zeqiang Lai, Xianghui Yang, Huiwen Shi, Zibo Zhao, Bowen Zhang, Hongyu Yan, Lifu Wang, Sicong Liu, Jihong Zhang, Meng Chen, Liang Dong, Yiwen Jia, Yulin Cai, Jiaao Yu, Yixuan Tang, Dongyuan Guo, Junlin Yu, Hao Zhang, Zhe...
Pith/arXiv arXiv 2025
-
[30]
Movement imitation with nonlinear dynamical systems in humanoid robots.Proceedings 2002 IEEE Int’l Conf on Robotics and Automation, 2, 2002
Auke Jan Ijspeert, Jun Nakanishi, and Stefan Schaal. Movement imitation with nonlinear dynamical systems in humanoid robots.Proceedings 2002 IEEE Int’l Conf on Robotics and Automation, 2, 2002. 21
2002
-
[31]
Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, et al.𝜋0.5: a vision-language-action model with open-world generalization.arXiv preprint arXiv:2504.16054, 2025. 7, 8
Pith/arXiv arXiv 2025
-
[32]
Arhan Jain, Mingtong Zhang, Kanav Arora, William Chen, Marcel Torne, Muhammad Zubair Irshad, Sergey Zakharov, Yue Wang, Sergey Levine, Chelsea Finn, et al. Polaris: Scalable real-to-sim evaluations for generalist robot policies.arXiv preprint arXiv:2512.16881, 2025. 2, 3, 4, 7, 8, 21, 34, 42, 43, 44
arXiv 2025
-
[33]
RobotArena∞: Scalable robot benchmarking via real-to-sim translation
Yash Jangir, Yidi Zhang, Kashu Yamazaki, Chenyu Zhang, Kuan-Hsun Tu, Tsung-Wei Ke, Lei Ke, Yonatan Bisk, and Katerina Fragkiadaki. RobotArena∞: Scalable robot benchmarking via real-to-sim translation. arXiv preprint arXiv:2510.23571, 2025. 2, 3, 21
arXiv 2025
-
[34]
Gsworld: Closed-loop photo-realistic simulation suite for robotic manipulation
Guangqi Jiang, Haoran Chang, Ri-Zhao Qiu, Yutong Liang, Mazeyu Ji, Jiyue Zhu, Zhao Dong, Xueyan Zou, and Xiaolong Wang. Gsworld: Closed-loop photo-realistic simulation suite for robotic manipulation. arXiv preprint arXiv:2510.20813, 2025. 2, 3, 21
arXiv 2025
-
[35]
Yunfan Jiang, Ruohan Zhang, Josiah Wong, Chen Wang, Yanjie Ze, Hang Yin, Cem Gokmen, Shuran Song, Jiajun Wu, and Li Fei-Fei. Behavior robot suite: Streamlining real-world whole-body manipulation for everyday household activities.arXiv preprint arXiv:2503.05652, 2025. 39, 40
Pith/arXiv arXiv 2025
-
[36]
Ditto: Building digital twins of articulated objects from interaction
Zhenyu Jiang, Cheng-Chun Hsu, and Yuke Zhu. Ditto: Building digital twins of articulated objects from interaction. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5616–5626, 2022. 3
2022
-
[37]
Dexmimicgen: Automated data generation for bimanual dexterous manipulation via imitation learning
Zhenyu Jiang, Yuqi Xie, Kevin Lin, Zhenjia Xu, Weikang Wan, Ajay Mandlekar, Linxi Fan, and Yuke Zhu. Dexmimicgen: Automated data generation for bimanual dexterous manipulation via imitation learning. arXiv preprint arXiv:2410.24185, 2024. 2, 4, 21
Pith/arXiv arXiv 2024
-
[38]
3d gaussian splatting for real-time radiance field rendering.ACM Trans
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George Drettakis, et al. 3d gaussian splatting for real-time radiance field rendering.ACM Trans. Graph., 42(4):139–1, 2023. 6, 24, 25
2023
-
[39]
Droid: A large- scale in-the-wild robot manipulation dataset.arXiv preprint arXiv:2403.12945, 2024
AlexanderKhazatsky, KarlPertsch, SurajNair, AshwinBalakrishna, SudeepDasari, SiddharthKaramcheti, Soroush Nasiriany, Mohan Kumar Srirama, Lawrence Yunliang Chen, Kirsty Ellis, et al. Droid: A large- scale in-the-wild robot manipulation dataset.arXiv preprint arXiv:2403.12945, 2024. 2, 6, 21, 40 12 SimFoundry: Modular and Automated Scene Generation for Pol...
Pith/arXiv arXiv 2024
-
[40]
Molmospaces: A large-scale open ecosystem for robot navigation and manipulation, 2026
Yejin Kim, Wilbert Pumacay, Omar Rayyan, Max Argus, Winson Han, Eli VanderBilt, Jordi Salvador, Abhay Deshpande, Rose Hendrix, Snehal Jauhri, Shuo Liu, Nur Muhammad Mahi Shafiullah, Maya Guru, Ainaz Eftekhar, Karen Farley, Donovan Clay, Jiafei Duan, Arjun Guru, Piper Wolters, Alvaro Herrasti, Ying-Chun Lee, Georgia Chalvatzaki, Yuchen Cui, Ali Farhadi, Di...
2026
-
[41]
Mask2cad: 3d shape prediction by learning to segment and retrieve
Weicheng Kuo, Anelia Angelova, Tsung-Yi Lin, and Angela Dai. Mask2cad: 3d shape prediction by learning to segment and retrieve. InEuropean Conference on Computer Vision, pages 260–277. Springer,
-
[42]
Patch2cad: Patchwise embedding learning for in-the-wild shape retrieval from a single image
Weicheng Kuo, Anelia Angelova, Tsung-Yi Lin, and Angela Dai. Patch2cad: Patchwise embedding learning for in-the-wild shape retrieval from a single image. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 12589–12599, 2021. 3
2021
-
[43]
Long Le, Jason Xie, William Liang, Hung-Ju Wang, Yue Yang, Yecheng Jason Ma, Kyle Vedder, Arjun Krishna, Dinesh Jayaraman, and Eric Eaton. Articulate-anything: Automatic modeling of articulated objects via a vision-language foundation model.arXiv preprint arXiv:2410.13882, 2024. 3, 5, 23
Pith/arXiv arXiv 2024
-
[44]
Any6d: Model-free 6d pose estimation of novel objects.CVPR, 2025
Taeyeop Lee, Bowen Wen, Minjun Kang, Gyuree Kang, In So Kweon, and Kuk-Jin Yoon. Any6d: Model-free 6d pose estimation of novel objects.CVPR, 2025. 3
2025
-
[45]
Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation
Chengshu Li, Ruohan Zhang, Josiah Wong, Cem Gokmen, Sanjana Srivastava, Roberto Martín-Martín, Chen Wang, Gabrael Levine, Michael Lingelbach, Jiankai Sun, et al. Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation. InConference on Robot Learning, pages 80–93. PMLR, 2023. 29
2023
-
[46]
Momagen: Generating demonstrations under soft and hard constraints for multi-step bimanual mobile manipulation
Chengshu Li, Mengdi Xu, Arpit Bahety, Hang Yin, Yunfan Jiang, Huang Huang, Josiah Wong, Sujay Garlanka, Cem Gokmen, Ruohan Zhang, Weiyu Liu, Jiajun Wu, Roberto Martín-Martín, and Li Fei-Fei. Momagen: Generating demonstrations under soft and hard constraints for multi-step bimanual mobile manipulation. InRSS 2025 Workshop on Whole-body Control and Bimanual...
2025
-
[47]
Evaluatingreal-worldrobotmanipulationpoliciesinsimulation
Xuanlin Li, Kyle Hsu, Jiayuan Gu, Karl Pertsch, Oier Mees, Homer Rich Walke, Chuyuan Fu, Ishikaa Lunawat,IsabelSieh,SeanKirmani,etal. Evaluatingreal-worldrobotmanipulationpoliciesinsimulation. arXiv preprint arXiv:2405.05941, 2024. 2, 3, 4, 21, 42
Pith/arXiv arXiv 2024
-
[48]
Art: Articulated reconstruction transformer.arXiv preprint arXiv:2512.14671, 2025
Zizhang Li, Cheng Zhang, Zhengqin Li, Henry Howard-Jenkins, Zhaoyang Lv, Chen Geng, Jiajun Wu, Richard Newcombe, Jakob Engel, and Zhao Dong. Art: Articulated reconstruction transformer.arXiv preprint arXiv:2512.14671, 2025. 3
Pith/arXiv arXiv 2025
-
[49]
Planar robot casting with real2sim2real self-supervised learning
Vincent Lim, Huang Huang, Lawrence Yunliang Chen, Jonathan Wang, Jeffrey Ichnowski, Daniel Seita, Michael Laskey, and Ken Goldberg. Planar robot casting with real2sim2real self-supervised learning. arXiv preprint arXiv:2111.04814, 2021. 3
Pith/arXiv arXiv 2021
-
[50]
Depth anything 3: Recovering the visual space from any views.arXiv preprint arXiv:2511.10647,
Haotong Lin, Sili Chen, Junhao Liew, Donny Y Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang. Depth anything 3: Recovering the visual space from any views.arXiv preprint arXiv:2511.10647,
-
[51]
Jiayi Liu, Denys Iliash, Angel X Chang, Manolis Savva, and Ali Mahdavi-Amiri. Singapo: Single image controlled generation of articulated parts in objects.arXiv preprint arXiv:2410.16499, 2024. 3
Pith/arXiv arXiv 2024
-
[52]
Partfield: Learning 3d feature fields for part segmentation and beyond
Minghua Liu, Mikaela Angelina Uy, Donglai Xiang, Hao Su, Sanja Fidler, Nicholas Sharp, and Jun Gao. Partfield: Learning 3d feature fields for part segmentation and beyond. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 9704–9715, 2025. 22
2025
-
[53]
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. InEuropean conference on computer vision, pages 38–55. Springer, 2024. 3 13 SimFoundry: Modular and Automated Scene Generation for Policy Learni...
2024
-
[54]
P3-sam: Native 3d part segmentation.arXiv preprint arXiv:2509.06784,
Changfeng Ma, Yang Li, Xinhao Yan, Jiachen Xu, Yunhan Yang, Chunshi Wang, Zibo Zhao, Yanwen Guo, Zhuo Chen, and Chunchao Guo. P3-sam: Native 3d part segmentation.arXiv preprint arXiv:2509.06784,
-
[55]
Abhiram Maddukuri, Zhenyu Jiang, Lawrence Yunliang Chen, Soroush Nasiriany, Yuqi Xie, Yu Fang, Wenqi Huang, Zu Wang, Zhenjia Xu, Nikita Chernyadev, et al. Sim-and-real co-training: A simple recipe for vision-based robotic manipulation.arXiv preprint arXiv:2503.24361, 2025. 2, 4, 21
Pith/arXiv arXiv 2025
-
[56]
Roboturk: A crowdsourcing platform for robotic skill learning through imitation
Ajay Mandlekar, Yuke Zhu, Animesh Garg, Jonathan Booher, Max Spero, Albert Tung, Julian Gao, John Emmons, Anchit Gupta, Emre Orbay, et al. Roboturk: A crowdsourcing platform for robotic skill learning through imitation. InConf on Robot Learning, 2018. 21
2018
-
[57]
Ajay Mandlekar, Danfei Xu, Josiah Wong, Soroush Nasiriany, Chen Wang, Rohun Kulkarni, Li Fei-Fei, Silvio Savarese, Yuke Zhu, and Roberto Martín-Martín. What matters in learning from offline human demonstrations for robot manipulation.arXiv preprint arXiv:2108.03298, 2021. 21
Pith/arXiv arXiv 2021
-
[58]
Ajay Mandlekar, Soroush Nasiriany, Bowen Wen, Iretiayo Akinola, Yashraj Narang, Linxi Fan, Yuke Zhu, and Dieter Fox. Mimicgen: A data generation system for scalable robot learning using human demonstrations.arXiv preprint arXiv:2310.17596, 2023. 2, 21, 29, 39
Pith/arXiv arXiv 2023
-
[59]
Mayank Mittal, Pascal Roth, James Tigue, Antoine Richard, Octi Zhang, Peter Du, Antonio Serrano- Munoz,XinjieYao,RenéZurbrügg,NikitaRudin,etal. Isaaclab: Agpu-acceleratedsimulationframework for multi-modal robot learning.arXiv preprint arXiv:2511.04831, 2025. 5, 29
Pith/arXiv arXiv 2025
-
[60]
Void: Video object and interaction deletion.arXiv preprint arXiv:2604.02296, 2026
Saman Motamed, William Harvey, Benjamin Klein, Luc Van Gool, Zhuoning Yuan, and Ta-Ying Cheng. Void: Video object and interaction deletion.arXiv preprint arXiv:2604.02296, 2026. 23
arXiv 2026
-
[61]
NVIDIA, Johan Bjorck, Fernando Castañeda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, et al. Gr00t n1: An open foundation model for generalist humanoid robots.arXiv preprint arXiv:2503.14734, 2025. 2, 7, 21, 39
Pith/arXiv arXiv 2025
-
[62]
Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
Abby O’Neill, Abdul Rehman, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, et al. Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0. In2024 IEEE Int’l Conf on Robotics and Automation (ICRA), 2024. 21
2024
-
[63]
Scenesmith: Agentic generation of simulation-ready indoor scenes, 2026
Nicholas Pfaff, Thomas Cohn, Sergey Zakharov, Rick Cory, and Russ Tedrake. Scenesmith: Agentic generation of simulation-ready indoor scenes, 2026. 2, 3
2026
-
[64]
Alvinn: An autonomous land vehicle in a neural network
Dean A Pomerleau. Alvinn: An autonomous land vehicle in a neural network. InAdvances in neural information processing systems, 1989. 21
1989
-
[65]
Xiaowen Qiu, Jincheng Yang, Yian Wang, Zhehuan Chen, Yufei Wang, Tsun-Hsuan Wang, Zhou Xian, and Chuang Gan. Articulate anymesh: Open-vocabulary 3d articulated objects modeling.arXiv preprint arXiv:2502.02590, 2025. 5, 23
Pith/arXiv arXiv 2025
-
[66]
Splatsim: Zero-shot sim2real transfer of rgb manipulation policies using gaussian splatting
M Nomaan Qureshi, Sparsh Garg, Francisco Yandun, David Held, George Kantor, and Abhisesh Silwal. Splatsim: Zero-shot sim2real transfer of rgb manipulation policies using gaussian splatting. In2025 IEEE International Conference on Robotics and Automation (ICRA), pages 6502–6509. IEEE, 2025. 3
2025
-
[67]
Sam 2: Segment anything in images and videos.arXiv preprint arXiv:2408.00714, 2024
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, et al. Sam 2: Segment anything in images and videos.arXiv preprint arXiv:2408.00714, 2024. 3
Pith/arXiv arXiv 2024
-
[68]
Sam 2: Segment anything in images and videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, et al. Sam 2: Segment anything in images and videos. InInternational Conference on Learning Representations, volume 2025, pages 28085–28128,
2025
-
[69]
23 14 SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation
-
[70]
Yam robot arm, 2025
I2RT Robotics. Yam robot arm, 2025. URLhttps://i2rt.com/collections/yam-arm. 6, 40
2025
-
[71]
Is imitation learning the route to humanoid robots?Trends in cognitive sciences, 3, 1999
Stefan Schaal. Is imitation learning the route to humanoid robots?Trends in cognitive sciences, 3, 1999. 21
1999
-
[72]
Structure-from-Motion Revisited
Johannes Lutz Schönberger and Jan-Michael Frahm. Structure-from-Motion Revisited. InConference on Computer Vision and Pattern Recognition (CVPR), 2016. 25
2016
-
[73]
Nerfstudio: A modular framework for neural radiance field development
Matthew Tancik, Ethan Weber, Evonne Ng, Ruilong Li, Brent Yi, Terrance Wang, Alexander Kristoffersen, Jake Austin, Kamyar Salahi, Abhik Ahuja, et al. Nerfstudio: A modular framework for neural radiance field development. InACM SIGGRAPH 2023 conference proceedings, pages 1–12, 2023. 24, 25
2023
-
[74]
Segment any mesh.arXiv preprint arXiv:2408.13679, 2024
George Tang, William Zhao, Logan Ford, David Benhaim, and Paul Zhang. Segment any mesh.arXiv preprint arXiv:2408.13679, 2024. 22, 23
Pith/arXiv arXiv 2024
-
[75]
Sam 3d: 3dfy anything in images, 2025
SAM 3D Team, Xingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang, Alexander Sax, Hao Tang, Weiyao Wang, Michelle Guo, Thibaut Hardin, Xiang Li, Aohan Lin, Jiawei Liu, Ziqi Ma, Anushka Sagar, Bowen Song, Xiaodong Wang, Jianing Yang, Bowen Zhang, Piotr Dollár, Georgia Gkioxari, Matt Feiszli, and Jitendra Malik. Sam 3d: 3dfy anything in images, 2025. URLht...
Pith/arXiv arXiv 2025
-
[76]
Hunyuan3d 2.5: Towards high-fidelity 3d assets generation with ultimate details, 2025
Tencent Hunyuan3D Team. Hunyuan3d 2.5: Towards high-fidelity 3d assets generation with ultimate details, 2025. URLhttps://arxiv.org/abs/2506.16504. 3
Pith/arXiv arXiv 2025
-
[77]
Interndata-a1: Pioneering high-fidelity synthetic data for pre-training generalist policy
Yang Tian, Yuyin Yang, Yiman Xie, Zetao Cai, Xu Shi, Ning Gao, Hangxu Liu, Xuekun Jiang, Zherui Qiu, Feng Yuan, et al. Interndata-a1: Pioneering high-fidelity synthetic data for pre-training generalist policy. arXiv preprint arXiv:2511.16651, 2025. 21
arXiv 2025
-
[78]
Triposr: Fast 3d object reconstruction from a single image.arXiv preprint arXiv:2403.02151, 2024
Dmitry Tochilkin, David Pankratz, Zexiang Liu, Zixuan Huang, Adam Letts, Yangguang Li, Ding Liang, Christian Laforte, Varun Jampani, and Yan-Pei Cao. Triposr: Fast 3d object reconstruction from a single image.arXiv preprint arXiv:2403.02151, 2024. 3
Pith/arXiv arXiv 2024
-
[79]
Robot learning with super-linear scaling.arXiv preprint arXiv:2412.01770, 2024
MarcelTorne, ArhanJain, JiayiYuan, VidaaranyaMacha, LarsAnkile, AnthonySimeonov, PulkitAgrawal, and Abhishek Gupta. Robot learning with super-linear scaling.arXiv preprint arXiv:2412.01770, 2024. 3, 21
arXiv 2024
-
[80]
Marcel Torne, Anthony Simeonov, Zechu Li, April Chan, Tao Chen, Abhishek Gupta, and Pulkit Agrawal. Reconciling reality through simulation: A real-to-sim-to-real approach for robust manipulation.arXiv preprint arXiv:2403.03949, 2024. 2, 3, 21
Pith/arXiv arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.