REVIEW 3 major objections 5 minor 68 references
RoboSnap: One-Shot Real-to-Sim Scene Generation for Generalizable Robot Learning and Evaluation
T0 review · 3 major / 5 minor · reviewed 2026-07-10 · grok-4.5
Pith's one-line read One RGB photo becomes a reusable, physics-stable simulation scene for robot learning and evaluation.
desk verdict Solid systems paper: single-image layered real-to-sim that actually ships robot-learning evidence and a 564-scene DROID companion, with the usual monocular-physics caveats. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Layered scene reconstruction plus alternating SDF–physics refinement: monocular geometry and asset models build collision-aware foreground objects and a Gaussian-splat background; a VLM-derived support/contact graph then drives residual pose optimization that alternates geometric contact losses with gravity settling until the layout is simulation-stable.
What would settle it
On a held-out set of real scenes and tasks, run the same real-only policies in RoboSnap scenes and check whether open-loop replay fails on most contact-critical demos or whether sim success rates no longer rank-order real success rates (Pearson r collapses and rank violations rise).
Extended reading notes
Core claim
From one RGB image, a layered real-to-sim pipeline can produce a simulation-ready scene whose physical layer is stable enough for open-loop trajectory replay, task-specific data generation that improves real-world policy fine-tuning, and closed-loop evaluation whose success rates correlate with real-world performance (Pearson r = 0.887, low rank violation).
Load-bearing premise
The method assumes that monocular object meshes, automatic support/contact relations, and generic mass/friction priors are accurate enough for contact events and policy ranking without extra real trajectories to repair the scene.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. RoboSnap reconstructs a simulation-ready manipulation scene from a single RGB image via a layered design: monocular instance reconstruction (SAM 3 / SAM 3D / VGGT registration) yields collision-aware foreground objects and support surfaces that are gravity-aligned and refined by a VLM-inferred support/contact graph plus alternating SDF–physics optimization, while a completed Gaussian-splat background supplies novel-view visual context. The recovered scenes are claimed to support open-loop replay of real DROID trajectories, task-specific synthetic demonstration generation for π0/π0.5 fine-tuning, robustness under real perturbations, and closed-loop sim-real evaluation correlation. The authors also release DROID-Sim, a companion set of 564 reconstructed DROID scenes, and report multi-axis experiments including stability after 300 Isaac steps, visual metrics vs RoLA, 5/5 replay on a small subset, real Franka success rates under three data mixtures, perturbation degradation, and r/MMRV on ten tasks.
Significance. If the claims hold, the paper offers a practical one-shot real-to-sim pipeline that turns casual RGB captures into reusable interactive environments for data generation and evaluation—addressing a genuine bottleneck for generalist robot policies. Strengths include a clear layered formulation, an explicit simulation-readiness refinement procedure with reported losses and algorithm, systematic multi-axis validation on real Franka hardware (Tables 2–3, Fig. 6), and the DROID-Sim companion dataset as community infrastructure. The work is systems-empirical rather than theoretical; its value is in demonstrating that single-image recovered scenes can be more than visual digital twins and can enter training/evaluation loops with measurable real-world gains.
major comments (3)
- §4.2 and the success definition for trajectory replay: success is open-loop geometric (grasp intended object and move to target without interpenetration/collision) on only 5 scenes (5/5 vs RoLA 2/5). This does not isolate whether monocular contact geometry and VLM physics priors match real contact dynamics. Because Q2–Q5 rest on the premise that refined single-image assets are accurate enough for closed-loop data generation and ranking without demonstration-driven pose repair (§3.1–3.2; Limitations §6), the manuscript needs either a larger replay set with quantitative contact/pose error, or an ablation that freezes layout and varies only recovered contact/physics parameters.
- Table 2 (R3 vs Real vs R2): pure-sim fine-tuning (R3) yields ~15–17% average real success, below real-only (~29–33%), while mixed R2 reaches ~42%. This pattern is consistent with generic visual diversity / domain randomization rather than faithful recovered contact dynamics. The central claim that RoboSnap scenes are reusable evaluation and data infrastructure would be stronger if the authors ablated (i) recovered layout vs random/procedural layout in the same visual layer, and (ii) VLM-prior mass/friction vs randomized physics, and reported whether ranking or failure modes flip under those controls.
- §4.5 / Fig. 6: sim-real correlation (r=0.887, MMRV=0.0066) is reported for N=10 tasks on real-only π0.5 only. With ten points and no confidence intervals or leave-one-out sensitivity, the correlation is fragile as evidence that monocular scale/contact and prior friction preserve task-relevant dynamics. Please report uncertainty (bootstrap CI), results for π0 and for mixed-trained policies, and at least one stress case where wrong contact or friction would reverse relative ranking if the recovered physics were inaccurate.
minor comments (5)
- §3.2 / Appendix B: loss weights w_pen, w_sup, w_con, w_reg and the full hyperparameter schedule (N_round, N_sdf, N_sim, N_damp, ε) are only partially specified in the main text; a single table of defaults would improve reproducibility.
- Fig. 3: PSNR is slightly worse than RoLA while other metrics improve; a short discussion of why pixel metrics are secondary for interaction-area reconstruction would help readers interpret the visual comparison.
- §4.1: detailed quantitative evaluation uses a fixed 10-scene subset of DROID-Sim; state selection criteria and whether the 564-scene release will include the same quality filters.
- Limitations §6 correctly notes no dedicated physical-parameter estimation and rigid/articulated scope; cross-reference these limits more explicitly when interpreting Table 1 stability and Fig. 6 correlation.
- Minor presentation: repeated “Homepage:” in the header; ensure equation numbering and Appendix B.1.1 prompt formatting are consistent in the camera-ready version.
Circularity Check
Empirical systems paper: success rates, replay, and sim-real correlation are measured against real robots and held-out rollouts, not forced by construction from inputs.
full rationale
RoboSnap is a real-to-sim engineering pipeline (monocular assets, VLM scene graph, SDF–physics refinement, layered Gaussian background), not a first-principles derivation that defines its answers into its premises. Load-bearing claims—physical stability after 300 Isaac steps (Table 1), open-loop DROID trajectory replay (5/5 vs RoLA 2/5), real-world π0/π0.5 success under real/sim mixtures (Table 2), robustness under perturbations (Table 3), and Pearson r=0.887 / MMRV=0.0066 on real-only policies (Fig. 6)—are all external measurements against real hardware, real trajectories, or independent baselines. Refinement is designed to reduce floating/interpenetration and is then scored with separate stability metrics versus ablations; that is method evaluation, not self-definitional circularity. Use of InternDataEngine, SAPIEN, Isaac Sim, SAM 3D, and VGGT is ordinary tooling; none of those citations supply a uniqueness theorem or fitted parameter that is renamed as the reported r or success gains. No fitted-input-called-prediction, no load-bearing self-citation chain, and no renaming of a known empirical law as a derived result. Score 0 is the correct honest finding.
Assumptions & free parameters
free parameters (5)
- SDF–physics refinement schedule (N_round, N_sdf, N_sim, N_damp, ε)
- Loss weights w_pen, w_sup, w_con, w_reg and λ_r=5
- Simulation material priors (density 3000 kg/m³, friction 0.5/5.0, restitution 0)
- Data mixture ratios R1–R3 for real/sim/sim-aug streaming
- Fixed 15k-step fine-tuning checkpoint
assumptions (5)
- domain assumption Rigid-body contact dynamics in Isaac/SAPIEN with convex V-HACD hulls adequately model the tabletop manipulation contacts of interest.
- domain assumption A single RGB view plus monocular geometry (VGGT) and image-conditioned meshes (SAM 3D) suffice to recover interaction layout up to residual SE(3) refinement.
- domain assumption VLM Set-of-Mark majority vote yields a usable support/contact scene graph for SDF constraints.
- domain assumption Depth compositing of physical-layer render and Gaussian-splat background preserves task-relevant appearance for VLA policies.
- standard math Standard SE(3) pose composition, ICP registration, and RANSAC plane fitting behave as usual.
invented entities (2)
-
RoboSnap layered scene S* (physical layer + Gaussian visual layer + refined poses)
independent evidence
-
DROID-Sim companion dataset (564 scenes)
Cite this review
Pith. "Pith review of RoboSnap: One-Shot Real-to-Sim Scene Generation for Generalizable Robot Learning and Evaluation." pith.science (2026). https://pith.science/paper/I73TBNG6
@misc{pith2026260706699,
author = {Pith},
title = {Pith review of: RoboSnap: One-Shot Real-to-Sim Scene Generation for Generalizable Robot Learning and Evaluation},
year = {2026},
howpublished = {\url{https://pith.science/paper/I73TBNG6}},
note = {Machine review of arXiv:2607.06699}
}
read the original abstract
Recovering real-world scenes as interactive simulation environments can enable generalizable robot learning and reproducible policy evaluation. However, constructing scenes that are both physically stable and visually faithful remains slow and expensive. In this work, we present RoboSnap, a real-to-sim framework that turns a single RGB image into a simulation-ready scene. The key idea is a layered design that separates the physics-critical interaction area from the surrounding visual context: collision-aware foreground assets are refined for stable robot interaction, while a 3D Gaussian splatting visual layer preserves faithful background appearance under novel views. Experiments on DROID scenes and real-robot tasks show that RoboSnap achieves reliable trajectory replay in the recovered scenes, supports task-specific synthetic data generation for policy training, and yields meaningful sim-real correlation for policy evaluation. To further support real-to-sim research, we introduce DROID-Sim, a real-to-sim companion dataset constructed from 564 real-world scenes in DROID. Extensive experiments suggest that the value of real-to-sim methods lies not only in high-fidelity visual reconstruction, but in turning real environments into reusable infrastructure for robot learning and evaluation.
Figures
Figures from the paper (13 more)
Reference graph
Works this paper leans on
-
[1]
RT-1: Robotics Transformer for Real-World Control at Scale
A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Haus- man, A. Herzog, et al. Rt-1: Robotics transformer for real-world control at scale. InarXiv preprint arXiv:2212.06817, 2022
work page Pith review arXiv 2022
-
[2]
Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, C. Xu, J. Luo, T. Kreiman, Y . Tan, L. Y . Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine. Octo: An open-source generalist robot policy. InProceedings of Robotics: Science and Systems, Delft, Netherlands, 2024
work page 2024
-
[3]
M. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn. Openvla: An open-source vision-language-action model.arXiv preprint arXiv:2406.09246, 2024
work page Pith review arXiv 2024
-
[4]
$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, et al.π0: A vision-language- action flow model for general robot control.arXiv preprint arXiv:2410.24164, 2024
work page Pith review arXiv 2024
-
[5]
$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
P. Intelligence.π0.5: A vision-language-action model with open-world generalization.arXiv preprint arXiv:2504.16054, 2025
work page Pith review arXiv 2025
-
[6]
DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, et al. Droid: A large-scale in-the-wild robot manipulation dataset.arXiv preprint arXiv:2403.12945, 2024
work page Pith review arXiv 2024
- [7]
-
[8]
O. X.-E. Collaboration, A. O’Neill, A. Rehman, and et al. Open X-Embodiment: Robotic learning datasets and RT-X models.https://arxiv.org/abs/2310.08864, 2023
work page Pith review arXiv 2023
Show all 68 references
-
[9]
J. Gu, F. Xiang, X. Li, Z. Ling, X. Liu, T. Mu, Y . Tang, S. Tao, X. Wei, Y . Yao, X. Yuan, P. Xie, Z. Huang, R. Chen, and H. Su. Maniskill2: A unified benchmark for generalizable manipulation skills. InInternational Conference on Learning Representations, 2023
2023
-
[10]
Raistrick, L
A. Raistrick, L. Lipson, Z. Ma, L. Mei, M. Wang, Y . Zuo, K. Kayan, H. Wen, B. Han, Y . Wang, A. Newell, H. Law, A. Goyal, K. Yang, and J. Deng. Infinite photorealistic worlds using procedural generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern ...
2023
-
[11]
J. Hao, N. Liang, Z. Luo, X. Xu, W. Zhong, R. Yi, Y . Jin, Z. Lyu, F. Zheng, L. Ma, and J. Pang. Mesatask: Towards task-driven tabletop scene generation via 3d spatial reasoning, 2025
2025
-
[12]
Y . Yang, B. Jia, P. Zhi, and S. Huang. Physcene: Physically interactable 3d scene synthesis for embodied ai. InProceedings of Conference on Computer Vision and Pattern Recognition (CVPR), 2024
2024
-
[13]
Z. Wang, Y . He, L. Yang, W. Zou, H. Ma, L. Liu, W. Sui, Y . Guo, and H. Su. Tabletopgen: Instance-level interactive 3d tabletop scene generation from text or single image.arXiv preprint arXiv:2512.01204, 2025
2025 arXiv
-
[14]
A. Choi, X. Wang, Z. Su, and W. Xu. Scaling sim-to-real reinforcement learning for robot vlas with generative 3d worlds, 2026. URLhttps://arxiv.org/abs/2603.18532
2026
-
[15]
Torne, A
M. Torne, A. Simeonov, Z. Li, A. Chan, T. Chen, A. Gupta, and P. Agrawal. Reconciling reality through simulation: A real-to-sim-to-real approach for robust manipulation.Robotics: Science and Systems (RSS), 2024
2024
-
[16]
X. Han, J. Yu, M. Liu, Y . Chen, X. Lyu, Y . Tian, B. Wang, W. Zhang, W. Zhang, and J. Pang. Re3sim: Generating high-fidelity simulation data via 3d-photorealistic real-to-sim for robotic manipulation. InIEEE International Conference on Robotics and Automation (ICRA), 2026
2026
-
[17]
M. N. Qureshi, S. Garg, F. Yandun, D. Held, G. Kantor, and A. Silwal. Splatsim: Zero-shot sim2real transfer of rgb manipulation policies using gaussian splatting, 2024. URLhttps: //arxiv.org/abs/2409.10161
2024 arXiv
-
[18]
H. Yu, B. Jia, Y . Chen, Y . Yang, P. Li, R. Su, J. Li, Q. Li, W. Liang, Z. Song-Chun, T. Liu, and S. Huang. Metascenes: Towards automated replica creation for real-world 3d scans. In Conference on Computer Vision and Pattern Recognition(CVPR), 2025
2025
-
[19]
A. Jain, M. Zhang, K. Arora, W. Chen, M. Torne, M. Z. Irshad, S. Zakharov, Y . Wang, S. Levine, C. Finn, W.-C. Ma, D. Shah, A. Gupta, and K. Pertsch. Polaris: Scalable real-to- sim evaluations for generalist robot policies, 2025. URLhttps://arxiv.org/abs/2512. 16881
2025
-
[20]
T. Dai, J. Wong, Y . Jiang, C. Wang, C. Gokmen, R. Zhang, J. Wu, and L. Fei-Fei. Automated creation of digital cousins for robust policy learning. InConference on Robot Learning (CoRL), 2024
2024
-
[21]
K. Yao, L. Zhang, X. Yan, Y . Zeng, Q. Zhang, L. Xu, W. Yang, J. Gu, and J. Yu. Cast: Component-aligned 3d scene reconstruction from an rgb image.ACM Transactions on Graph- ics (TOG), 44(4):1–19, 2025
2025
-
[22]
S. Zhao, J. Mao, W. Chow, Z. Shangguan, T. Shi, R. Xue, Y . Zheng, Y . Weng, Y . You, D. Seita, et al. Robot learning from any images. InConference on Robot Learning, pages 4226–4245. PMLR, 2025
2025
-
[23]
J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny. Vggt: Visual geometry grounded transformer. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025
2025
-
[24]
J. Wang, M. Chen, S. Zhang, N. Karaev, J. Sch¨onberger, P. Labatut, P. Bojanowski, D. Novotny, A. Vedaldi, and C. Rupprecht. VGGT-Ω. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
2026
-
[25]
S. D. Team, X. Chen, F.-J. Chu, P. Gleize, K. J. Liang, A. Sax, H. Tang, W. Wang, M. Guo, T. Hardin, X. Li, A. Lin, J. Liu, Z. Ma, A. Sagar, B. Song, X. Wang, J. Yang, B. Zhang, P. Doll´ar, G. Gkioxari, M. Feiszli, and J. Malik. Sam 3d: 3dfy anything in images. 2025. URL https...
2025 arXiv
-
[26]
Xiang, Z
J. Xiang, Z. Lv, S. Xu, Y . Deng, R. Wang, B. Zhang, D. Chen, X. Tong, and J. Yang. Structured 3d latents for scalable and versatile 3d generation.arXiv preprint arXiv:2412.01506, 2024
2024 arXiv
-
[27]
Nasiriany, A
S. Nasiriany, A. Maddukuri, L. Zhang, A. Parikh, A. Lo, A. Joshi, A. Mandlekar, and Y . Zhu. Robocasa: Large-scale simulation of everyday tasks for generalist robots. InRobotics: Science and Systems (RSS), 2024
2024
-
[28]
Nasiriany, S
S. Nasiriany, S. Nasiriany, A. Maddukuri, and Y . Zhu. Robocasa365: A large-scale simulation framework for training and benchmarking generalist robots. InInternational Conference on Learning Representations (ICLR), 2026
2026
-
[29]
Y . Yang, B. Jia, S. Zhang, and S. Huang. Sceneweaver: All-in-one 3d scene synthesis with an extensible and self-reflective agent. InAdvances in Neural Information Processing Systems (NeurIPS), 2025
2025
-
[30]
Zhong, P
W. Zhong, P. Cao, Y . Jin, L. Luo, W. Cai, J. Lin, H. Wang, Z. Lyu, T. Wang, B. Dai, X. Xu, and J. Pang. Internscenes: A large-scale simulatable indoor scene dataset with realistic layouts,
-
[31]
URLhttps://arxiv.org/abs/2509.10813
-
[32]
Bochkovskii, A
A. Bochkovskii, A. Delaunoy, H. Germain, M. Santos, Y . Zhou, S. R. Richter, and V . Koltun. Depth pro: Sharp monocular metric depth in less than a second. InInternational Conference on Learning Representations, 2025. URLhttps://arxiv.org/abs/2410.02073
2025 arXiv
-
[33]
T. H. Team. Hunyuan3d 2.0: Scaling diffusion models for high resolution textured 3d assets generation, 2025
2025
-
[34]
H. Lou, Y . Liu, Y . Pan, Y . Geng, J. Chen, W. Ma, C. Li, L. Wang, H. Feng, L. Shi, L. Luo, and Y . Shi. Robo-gs: A physics consistent spatial-temporal model for robotic arm with hybrid representation, 2024. URLhttps://arxiv.org/abs/2408.14873
2024 arXiv
-
[35]
H. Fan, H. Dai, J. Zhang, J. Li, Q. Yan, Y . Zhao, M. Gao, J. Wu, H. Tang, and H. Dong. Twinaligner: Visual-dynamic alignment empowers physics-aware real2sim2real for robotic manipulation. 2025. URLhttps://arxiv.org/abs/2512.19390
2025
-
[36]
Pfaff, E
N. Pfaff, E. Fu, J. Binagia, P. Isola, and R. Tedrake. Scalable real2sim: Physics-aware asset generation via robotic pick-and-place setups. 2025. URLhttps://arxiv.org/abs/2503. 00370
2025
-
[37]
Zook, F.-Y
A. Zook, F.-Y . Sun, J. Spjut, V . Blukis, S. Birchfield, and J. Tremblay. GRS: Generating robotic simulation tasks from real-world images. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 594–603, 2025
2025
-
[38]
Y . Fang, Y . Yang, X. Zhu, K. Zheng, G. Bertasius, D. Szafir, and M. Ding. Rebot: Scaling robot learning with real-to-sim-to-real robotic video synthesis.arXiv preprint arXiv:2503.14526, 2025
2025 arXiv
-
[39]
C. Yuan, S. Joshi, S. Zhu, H. Su, H. Zhao, and Y . Gao. Roboengine: Plug-and-play robot data augmentation with semantic robot segmentation and background generation.arXiv preprint arXiv:2503.18738, 2025
2025 arXiv
-
[40]
J. Yu, L. Fu, H. Huang, K. El-Refai, R. A. Ambrus, R. Cheng, M. Z. Irshad, and K. Goldberg. Real2render2real: Scaling robot data without dynamics simulation or robot hardware, 2025. URLhttps://arxiv.org/abs/2505.09601
2025 arXiv
-
[41]
S. Yang, W. Yu, J. Zeng, J. Lv, K. Ren, C. Lu, D. Lin, and J. Pang. Novel demonstra- tion generation with gaussian splatting enables robust one-shot manipulation.arXiv preprint arXiv:2504.13175, 2025. 11
2025 arXiv
-
[42]
B. Wang, H. Zhang, S. Zhang, J. Hao, M. Jia, Q. Lv, Y . Mao, Z. Lyu, J. Zeng, X. Xu, et al. Robovip: Multi-view video generation with visual identity prompting augments robot manip- ulation.arXiv preprint arXiv:2601.05241, 2026
2026
-
[43]
Y . Zhao, H. Fan, D. Chen, S. Chen, L. Chen, X. Li, G. Ren, and H. Dong. Real2edit2real: Generating robotic demonstrations via a 3d control interface. 2025. URLhttps://arxiv. org/abs/2512.19402
2025
-
[44]
Mandlekar, S
A. Mandlekar, S. Nasiriany, B. Wen, I. Akinola, Y . Narang, L. Fan, Y . Zhu, and D. Fox. Mimicgen: A data generation system for scalable robot learning using human demonstrations. In7th Annual Conference on Robot Learning, 2023
2023
-
[45]
Jiang, Y
Z. Jiang, Y . Xie, K. Lin, Z. Xu, W. Wan, A. Mandlekar, L. Fan, and Y . Zhu. Dexmimicgen: Automated data generation for bimanual dexterous manipulation via imitation learning. In 2025 IEEE International Conference on Robotics and Automation (ICRA), 2025
2025
-
[46]
Y . Wang, Z. Xian, F. Chen, T.-H. Wang, Y . Wang, K. Fragkiadaki, Z. Erickson, D. Held, and C. Gan. Robogen: Towards unleashing infinite data for automated robot learning via generative simulation.arXiv preprint arXiv:2311.01455, 2023
2023 arXiv
-
[47]
Y . Tian, Y . Yang, Y . Xie, Z. Cai, X. Shi, N. Gao, H. Liu, X. Jiang, Z. Qiu, F. Yuan, et al. Interndata-a1: Pioneering high-fidelity synthetic data for pre-training generalist policy.arXiv preprint arXiv:2511.16651, 2025
2025
-
[48]
B. Liu, Y . Zhu, C. Gao, Y . Feng, Q. Liu, Y . Zhu, and P. Stone. Libero: Benchmarking knowl- edge transfer for lifelong robot learning.arXiv preprint arXiv:2306.03310, 2023
2023 arXiv
-
[49]
C. Li, R. Zhang, J. Wong, C. Gokmen, S. Srivastava, R. Mart ´ın-Mart´ın, C. Wang, G. Levine, W. Ai, B. Martinez, H. Yin, M. Lingelbach, M. Hwang, A. Hiranaka, S. Garlanka, A. Ay- din, S. Lee, J. Sun, M. Anvari, M. Sharma, D. Bansal, S. Hunter, K.-Y . Kim, A. Lou, C. R. Matthew...
2024 arXiv
-
[50]
T. Chen, Z. Chen, B. Chen, Z. Cai, Y . Liu, Z. Li, Q. Liang, X. Lin, Y . Ge, Z. Gu, et al. Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation.arXiv preprint arXiv:2506.18088, 2025
2025 arXiv
-
[51]
X. Li, K. Hsu, J. Gu, K. Pertsch, O. Mees, H. R. Walke, C. Fu, I. Lunawat, I. Sieh, S. Kir- mani, S. Levine, J. Wu, C. Finn, H. Su, Q. Vuong, and T. Xiao. Evaluating real-world robot manipulation policies in simulation.arXiv preprint arXiv:2405.05941, 2024
2024 arXiv
-
[52]
Achiam, S
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. Gpt-4 technical report.arXiv preprint arXiv:2303.08774, 2023
2023 arXiv
-
[53]
Carion, L
N. Carion, L. Gustafson, Y .-T. Hu, S. Debnath, R. Hu, D. Suris, C. Ryali, K. V . Alwala, H. Khedr, A. Huang, J. Lei, T. Ma, B. Guo, A. Kalla, M. Marks, J. Greer, M. Wang, P. Sun, R. R¨adle, T. Afouras, E. Mavroudi, K. Xu, T.-H. Wu, Y . Zhou, L. Momeni, R. Hazra, S. Ding, S. V...
2025 arXiv
-
[54]
Marble.https://docs.worldlabs.ai/, 2026
World Labs. Marble.https://docs.worldlabs.ai/, 2026
2026
-
[55]
C. Ma, Y . Li, X. Yan, J. Xu, Y . Yang, C. Wang, Z. Zhao, Y . Guo, Z. Chen, and C. Guo. P3-sam: Native 3d part segmentation, 2025. URLhttps://arxiv.org/abs/2509.06784. 12
2025
-
[56]
Xiang, Y
F. Xiang, Y . Qin, K. Mo, Y . Xia, H. Zhu, F. Liu, M. Liu, H. Jiang, Y . Yuan, H. Wang, L. Yi, A. X. Chang, L. J. Guibas, and H. Su. SAPIEN: A simulated part-based interactive environ- ment. InThe IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2020
2020
-
[57]
J. Yang, H. Zhang, F. Li, X. Zou, C. Li, and J. Gao. Set-of-mark prompting unleashes extraor- dinary visual grounding in gpt-4v.arXiv preprint arXiv:2310.11441, 2023
2023 arXiv
-
[58]
Mamou, E
K. Mamou, E. Lengyel, and A. Peters. V olumetric hierarchical approximate convex decompo- sition.Game engine gems, 3:141–158, 2016
2016
-
[59]
Isaac Sim
NVIDIA. Isaac Sim. URLhttps://github.com/isaac-sim/IsaacSim
-
[60]
H.-S. Fang, C. Wang, H. Fang, M. Gou, J. Liu, H. Yan, W. Liu, Y . Xie, and C. Lu. Anygrasp: Robust and efficient grasp perception in spatial and temporal domains.IEEE Transactions on Robotics (T-RO), 2023
2023
-
[61]
Sundaralingam, A
B. Sundaralingam, A. Murali, and S. Birchfield. curobov2: Dynamics-aware motion generation with depth-fused distance fields for high-dof robots, 2026
2026
-
[62]
Zhang, P
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang. The unreasonable effectiveness of deep features as a perceptual metric.arXiv preprint arXiv:1801.03924, 2018
2018 arXiv
-
[63]
B. Wen, W. Yang, J. Kautz, and S. Birchfield. FoundationPose: Unified 6d pose estimation and tracking of novel objects. InCVPR, 2024
2024
-
[64]
Gemini 2.5 flash image (nano banana).https://ai.google
Google AI for Developers. Gemini 2.5 flash image (nano banana).https://ai.google. dev/gemini-api/docs/models/gemini-2.5-flash-image, 2026. Model documenta- tion. Accessed: 2026-05-16
2026
-
[65]
Q.-Y . Zhou, J. Park, and V . Koltun. Open3D: A modern library for 3D data processing. arXiv:1801.09847, 2018
2018 arXiv
-
[66]
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection.arXiv preprint arXiv:2303.05499, 2023
2023 arXiv
-
[67]
B. Xiao, H. Wu, W. Xu, X. Dai, H. Hu, Y . Lu, M. Zeng, C. Liu, and L. Yuan. Florence-2: Ad- vancing a unified representation for a variety of vision tasks.arXiv preprint arXiv:2311.06242, 2023
2023 arXiv
-
[68]
B. Li, D. Wu, J. Li, S. Zhou, Z. Zeng, L. Li, and H. Zha. Mv-sam3d: Adaptive multi-view fusion for layout-aware 3d generation.arXiv preprint arXiv:2603.11633, 2026. 13 A Experiment A.1 Metrics A.1.1 Visual-alignment metrics We evaluate visual alignment by comparing each method...
2026 arXiv
Reviewed July 10, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.