Pith. sign in

REVIEW 4 major objections 4 minor 68 references

Adding 50K diverse community demos before fine-tuning lifts a robot policy's benchmark success by 5.8 points.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 06:58 UTC pith:ETSPQAYG

load-bearing objection Real infrastructure and a useful dataset, but the headline numbers rest on single runs and a collapsed RoboCasa baseline — worth reviewing, not ready to trust. the 4 major comments →

arxiv 2607.21588 v1 pith:ETSPQAYG submitted 2026-07-23 cs.RO

AXIS: A Growable Community-Driven Data Engine for Scalable Robot Manipulation

classification cs.RO
keywords robot manipulationdata enginecommunity-driven data collectionbrowser teleoperationvision-language-action modelscontinual pretrainingdata augmentationbenchmark scaling
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that robot-manipulation datasets should be growable—continuously expandable through community teleoperation and automated task generation—rather than fixed static collections, and that such growth can measurably improve downstream policy robustness. To show this, the authors build an end-to-end data engine that lets people teleoperate a simulated robot arm in a web browser, automatically validates and cleans the collected demonstrations, smooths them, and replays them under randomized visual and physical conditions to multiply their variety. The central result is that adding a snapshot of roughly 50,000 such demonstrations as continual pretraining for a vision-language-action policy improves its success rate on a seven-axis manipulation robustness benchmark by 5.8 percentage points over the baseline, while a same-sized alternative corpus does not produce the same gain. The authors also report consistent scaling: success rises as the dataset grows from a quarter to the full snapshot, suggesting the benefit has not saturated. If true, this points to a low-cost, participatory path to the diverse data that general-purpose robot policies need.

Core claim

The paper's central claim is that the path to scalable robot manipulation runs through data infrastructure that grows with community participation, not through larger one-time collections. Concretely, it claims that continual pretraining of a vision-language-action flow model on AXIS-100%—50,129 browser-collected, automatically validated, and physically and visually augmented demonstrations spanning 207 tabletop tasks—improves downstream success on a seven-axis manipulation robustness benchmark by 5.8 percentage points over the vanilla checkpoint, and outperforms a same-volume baseline drawn from another large simulated dataset by the reported margin of 37.3%. The authors attribute this gap

What carries the argument

The AXIS data engine itself: a three-layer pipeline that (1) generates new manipulation tasks from text instructions, with automated layout validation and per-task success checkers; (2) collects demonstrations in a browser through a lightweight physics engine compiled to run on commodity machines; and (3) cleans trajectories—removing static segments, applying polynomial smoothing, and resampling to a fixed control frequency—then augments them by replaying verified state sequences under randomized camera viewpoints, lighting, textures, object poses, and dynamics. The load-bearing pieces are the task-specific success checker reused across collection and backend validation, and the state-author

Load-bearing premise

The quantitative claims rest on simulated benchmark success; if the measured robustness gains do not transfer to physical robot manipulation—the paper's only real-world evidence is a qualitative demonstration—the central promise of scalable real-world manipulation is not yet supported.

What would settle it

Run the same continual-pretraining and fine-tuning recipe on a physical Franka robot across the same perturbation types and measure success; if real-robot success does not track the simulated benchmark gains, or if a same-volume generic corpus matches the AXIS gain on real hardware, the central claim fails. A cheaper simulation-side falsifier: hold the AXIS snapshot fixed, remove the augmentation stages, and check whether the 5.8-point gain disappears; if it does, the claim reduces from 'community data engine' to 'augmentation pipeline'.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Additional AXIS data beyond the current snapshot should keep improving downstream robustness; the 25%/50%/100% slope shows no sign of saturation.
  • A same-size simulation corpus does not substitute for the diversity produced by community task generation and augmentation; dataset composition matters as much as volume.
  • Continual pretraining on diverse simulated manipulation data transfers to visual and geometric perturbation axes such as camera, sensor noise, background, and layout, and appears to help even on language and robot-pose variations that were not explicitly augmented.
  • Cleaned, smoothed, resampled trajectories preserve task semantics while sharply reducing acceleration and jerk, so browser-quality demonstrations can be turned into training-ready data at scale.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the gain is driven by augmentation diversity rather than trajectory count, a natural stress test is to hold the data fixed and ablate the augmentation stages; the paper's per-axis analysis predicts the largest drops on camera, sensor-noise, and background axes.
  • The same growth loop could be pointed at other embodiments or physical-robot interfaces; if browser-simulated data remains useful, decoupling collection from robot availability would make long-tail skill coverage far cheaper.
  • The growable principle suggests a failure-driven collection loop: after evaluating on held-out perturbations, the hardest failures could be converted into new tasks and new demonstrations, closing the loop between evaluation and data collection.
  • A caveat the paper leaves implicit: the matched-volume baseline was drawn from a different task distribution, so a stronger control would be an equal-diversity, equal-volume dataset built with the same augmentation pipeline but without community teleoperation, isolating community effects from augmentation effects.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces AXIS, a community-driven data engine for robot manipulation that combines browser-based MuJoCo teleoperation, LLM-assisted task generation, automated validation/cleaning/augmentation, and versioned dataset snapshots. The authors release a dataset of 207 tasks and 50K+ trajectories, and they evaluate continual pretraining of π0.5 on AXIS snapshots before fine-tuning on LIBERO. The headline results are that AXIS-100% pretraining improves LIBERO-Plus overall success from 83.9% (vanilla) to 88.8%, outperforms a volume-matched RoboCasa365 control (57.5%), and scales monotonically from AXIS-25% (84.7%) to AXIS-50% (85.7%) to AXIS-100%.

Significance. The infrastructure contribution is potentially strong: a growable, web-based collection platform, automated success checking, trajectory cleaning, IsaacSim-based augmentation, and snapshot-based evaluation protocol are useful resources for the manipulation community. The 207-task, 50K-trajectory dataset and the fixed-protocol scaling study are timely. However, the current empirical support for the central claims is weakened by missing variance estimates, an untuned and collapsed control baseline, and an internally inconsistent headline percentage. If the experimental issues are addressed, the paper could be an important community resource; as written, the quantitative conclusions are not yet reliable.

major comments (4)
  1. [§5.1, Table 2] All success rates in Table 2 are single point estimates: no standard deviations, no number of seeds, and the rollout budget K is never stated. The headline scaling trend (84.7/85.7/88.8) and the 5.8% improvement over vanilla are therefore not statistically grounded; differences of this size could easily be within rollout noise. Report means and variances over multiple seeds and state K for every (task, perturbation) cell, or provide confidence intervals.
  2. [§5.1 and Table 2, RoboCasa-matched row] The matched control collapses to 57.5% overall versus 83.9% for vanilla π0.5 with no pretraining. Since the paper states that all pretraining conditions use 'identical optimizer settings, learning rate, batch size, and gradient-step budget' (§5.1) and no hyperparameter search or seed variation is reported for RoboCasa365, this row does not establish a fair baseline. A control that is 26.4 points below vanilla signals an ill-suited pretraining recipe, not AXIS-specific data quality. The claim that gains are 'not explained by dataset size alone' depends entirely on this broken control. Re-run the control with per-corpus hyperparameter tuning, report the best (or multiple) settings, and verify that the tuned control does not collapse.
  3. [Abstract, §1, Tables 2–3] The headline 'outperforms the model pretrained on RoboCasa365 by 37.3%' is not reproducible from the reported numbers. Table 2 gives 88.8% versus 57.5%, an absolute gap of 31.3 percentage points; relative to RoboCasa the gain is 54.4%, and 37.3% is only obtained by dividing the gap by the vanilla baseline (83.9%). The paper uses '5.8%' relative for vanilla but '37.3%' absolute/relative mix. Define one comparison convention and report both absolute and relative gains consistently; the current formulation makes the central result appear larger than supported.
  4. [§5.3, Table 3; Appendix E, Table 8] The LIBERO-Plus perturbation axes on which AXIS shows the largest gains (Camera, Sensor Noise, Background, Layout) are the same axes randomized by the AXIS IsaacSim augmentation pipeline (camera position deltas, lighting intensities, full-scene mode, etc., Table 8). This is to some extent a self-fulfilling evaluation: the experiment tests whether AXIS augmentation transfers to a downstream benchmark that probes the same axes. The paper acknowledges this in question (iii), but then interprets the gains as evidence of broad robustness improvement. The claim should be restricted to transfer of trained augmentation axes, and supported by at least one held-out perturbation family not explicitly augmented (e.g., the Language and Robot axes are useful but were not tuned).
minor comments (4)
  1. [Table 1] Smoothed+resampled trajectories reduce replay success from 100% to 86.2%; since the pipeline uses replay as verification, explain whether this reflects accepted episodes and how the 14% failure is handled.
  2. [Appendix H] The real-world evidence is qualitative video only, with no success rates or quantitative sim-to-real comparison. The figure captions and Fig. 1 label this 'Real-World Validation'; please label as qualitative illustration or add metrics.
  3. [§5.1] The rollout budget K is never specified, although the protocol says it is fixed. State the exact value, the number of tasks per perturbation axis, and the total number of rollouts per condition.
  4. [§4/Appendix A] Dataset release details (URL, license, schema documentation, and exact snapshot version) should be included for reproducibility.

Circularity Check

0 steps flagged

No significant circularity: AXIS is an empirical data-engine study whose headline result is benchmarked on external LIBERO-Plus, not derived from fitted inputs or self-citations.

full rationale

The paper's central claim is a controlled empirical comparison: π0.5 is continually pretrained on AXIS snapshots and evaluated on LIBERO-Plus, an externally defined robustness suite, against a vanilla checkpoint and a RoboCasa365-matched control. No parameter is fitted to the LIBERO-Plus metric, and no 'prediction' is constructed from the evaluation definition; the success criteria, rollouts, and perturbation axes come from LIBERO-Plus, not from AXIS. The fact that AXIS's IsaacSim augmentation randomizes camera, lighting, texture, and layout while LIBERO-Plus evaluates those same axes is a design alignment that makes the transfer claim testable, not tautological: vanilla π0.5 and the RoboCasa control receive no such augmentation, so the comparison has independent content. The RoboCasa control's collapse (57.5% vs 83.9% vanilla) and the unreproducible 37.3% figure are serious experimental-validity and reporting problems, but they do not make the derivation circular. The paper contains no load-bearing self-citations, no imported uniqueness theorem, and no fitted-input-called-prediction step. The Limitations section explicitly concedes that sim-to-real transfer remains an open challenge, which further confirms that the simulation results are not being dressed up as a derivation. Overall: no significant circularity, score 0.

Axiom & Free-Parameter Ledger

7 free parameters · 4 axioms · 0 invented entities

AXIS introduces a system and a dataset, not new physical entities. Its central claims rest on domain assumptions about (1) the validity of crowdsourced browser teleoperation as a source of useful robot data, (2) the accuracy of automated success checkers, (3) the semantic preservation under large-scale simulation augmentation, and (4) the transferability of simulated benchmark results to physical robots. These assumptions are plausible but not independently verified in the paper.

free parameters (7)
  • static_segment_threshold = 5e-3
    Threshold for classifying robot joint variation as static in trajectory cleaning (Appendix D, Algorithm 1); chosen by hand and directly affects which segments are removed and thus the training data.
  • Savitzky-Golay_window = 15
    Window size for smoothing robot motion (Appendix D); manually set, influences trajectory shape and smoothness.
  • Savitzky-Golay_polynomial_order = 3
    Polynomial order for the smoothing filter (Appendix D); a design choice.
  • target_control_frequency = 20 Hz
    Resampling frequency after cleaning (Appendix D); affects the action chunking and policy training interface.
  • pretraining_steps = 100,000
    Number of continual pretraining steps on AXIS (Appendix F, Table 10); part of the training recipe that defines the optimization budget.
  • LIBERO_finetune_steps = 30,000
    Number of fine-tuning steps on LIBERO (Appendix G, Table 11); a hyperparameter that constrains the comparison.
  • rollout_budget_K = not stated in text
    The paper repeatedly says the rollout budget is fixed to K but never gives its value (Section 5.1). Without K, the reported success rates are impossible to interpret with confidence.
axioms (4)
  • domain assumption Browser-based MuJoCo teleoperation produces demonstrations that reflect real-world manipulation behavior
    The entire data collection method relies on this equivalence; no comparison of browser-teleoperated trajectories with physical-or high-fidelity-simulated teleoperation is provided.
  • domain assumption Task-specific success checkers accurately determine task completion
    Success checkers are used both during collection and backend validation; if they are inaccurate, the dataset contains mislabeled trajectories. No manual verification of checker accuracy is reported.
  • domain assumption IsaacSim augmentation preserves task semantics and success conditions
    Randomizing scenes, physics, and visuals is assumed to generate valid variations that still satisfy the original task goals. If augmentation creates impossible or inconsistent states, the augmented data could be harmful.
  • domain assumption LIBERO-Plus robustness axes are a good proxy for real-world robustness
    All quantitative claims about 'robustness' and 'generalization' are measured on the simulated LIBERO-Plus benchmark. The paper does not establish a correlation with physical robot performance; its own real-world validation is qualitative.

pith-pipeline@v1.3.0-alltime-deepseek · 21564 in / 12844 out tokens · 107009 ms · 2026-08-01T06:58:11.203989+00:00 · methodology

0 comments
read the original abstract

Learning effective robot manipulation policies requires diverse, high-quality demonstrations, yet existing data pipelines are often difficult to scale because they rely on specialized hardware, centralized operators, or fixed task suites. We present AXIS, a growable community-driven data engine and benchmark for scalable robot learning, which enables browser-based teleoperation for large-scale demonstration collection, automatically generates and validates new manipulation tasks, and transforms community-collected demonstrations into training-ready data through automated success checking, quality filtering, trajectory smoothing, and visual and physics-based augmentation. The AXIS dataset currently contains 207 diverse tasks and 50K+ trajectories. Meanwhile, AXIS organizes data into task snapshots and evaluates policies with a systematic held-out protocol. We compare vision-language-action (VLA) policies under a unified AXIS evaluation suite and analyze scaling behavior across different data volumes. Continual pretraining on AXIS substantially improves the overall success rate of $\pi_{0.5}$ by 5.8%, outperforms the model pretrained on RoboCasa365 by 37.3%, and exhibits consistent scaling with increasing data volume, with the largest gains observed under layout, sensor-noise, and camera perturbations.

Figures

Figures reproduced from arXiv: 2607.21588 by Dihong Huang, Hai Zhai, Jiachen Li, Jianfei Yang, Jie Wang, Mengfei Zhao, Mingxuan Yan, Peihao Li, Ruiqi Zhuang, Rui Zhang, Tony Zhou, Yanjia Huang, Yikai Tang, Yuchen Huang, Zhexi Luo.

Figure 1
Figure 1. Figure 1: AXIS: A growable community-driven data engine unifying task generation, web teleoper [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Visualization of augmentation diversity for three example tasks. For each task (row), we [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Overview of the AXIS dataset. Validation and refinement. To ensure data quality, all demonstrations undergo a standardized val￾idation pipeline that verifies trajectory completeness, task success, and physical consistency. Valid trajectories are subsequently smoothed and resampled to a unified control frequency, improving temporal consistency and reducing teleoperation noise. As shown in [PITH_FULL_IMAGE:… view at source ↗
Figure 4
Figure 4. Figure 4: TaskGen pipeline overview. A language instruction is converted into task, scene, and [PITH_FULL_IMAGE:figures/full_fig_p017_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Real-world Franka rollouts. The first row shows a grasping task, and the second row shows [PITH_FULL_IMAGE:figures/full_fig_p024_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Real-time pick-and-place rollout. The two rows show the global and wrist views. [PITH_FULL_IMAGE:figures/full_fig_p025_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Pick-and-place replay. The three rows show the simulation replay, real-world global view, [PITH_FULL_IMAGE:figures/full_fig_p025_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

68 extracted references · 25 linked inside Pith

  1. [1]

    Brohan, N

    A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Haus- man, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. Joshi, R. Ju- lian, D. Kalashnikov, Y . Kuang, I. Leal, K.-H. Lee, S. Levine, Y . Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, ...

  2. [2]

    H. Geng, F. Wang, S. Wei, Y . Li, B. Wang, B. An, C. T. Cheng, H. Lou, P. Li, Y .-J. Wang, et al. Roboverse: Towards a unified platform, dataset and benchmark for scalable and generalizable robot learning. InRobotics: Science and Systems (RSS), 2025

  3. [3]

    J. Duan, S. Yu, H. L. Tan, H. Zhu, and C. Tan. A survey of embodied ai: From simulators to research tasks.IEEE Transactions on Emerging Topics in Computational Intelligence, 6(2): 230–244, 2022

  4. [4]

    H. Choi, C. Crump, C. Duriez, A. Elmquist, G. Hager, D. Han, F. Hearl, J. Hodgins, A. Jain, F. Leve, et al. On the use of simulation in robotics: Opportunities, challenges, and suggestions for moving forward.Proceedings of the National Academy of Sciences, 118(1):e1907856118, 2021

  5. [5]

    J. Eßer, N. Bach, C. Jestel, O. Urbann, and S. Kerner. Guided reinforcement learning: A review and evaluation for efficient and effective real-world robotics [survey].IEEE Robotics & Automation Magazine, 30(2):67–85, 2022

  6. [6]

    Dasari, F

    S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn. Robonet: Large-scale multi-robot learning. InConference on Robot Learning, pages 885–897. PMLR, 2020

  7. [7]

    Ebert, Y

    F. Ebert, Y . Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine. Bridge data: Boosting generalization of robotic skills with cross-domain datasets. arXiv preprint arXiv:2109.13396, 2021

  8. [8]

    H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, et al. Bridgedata v2: A dataset for robot learning at scale. In J. Tan, M. Toussaint, and K. Darvish, editors, Proceedings of The 7th Conference on Robot Learning, volume 229 ofProceedings of Machine Learning Research, pages 1723–1736. PMLR, 06–09 Nov 2023. 9

  9. [9]

    Khazatsky, K

    A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, et al. Droid: A large-scale in-the-wild robot manipulation dataset. InRobotics: Science and Systems, 2024

  10. [10]

    K. Wu, C. Hou, J. Liu, Z. Che, X. Ju, Z. Yang, M. Li, Y . Zhao, Z. Xu, G. Yang, et al. Robo- mind: Benchmark on multi-embodiment intelligence normative data for robot manipulation. InRobotics: Science and Systems (RSS) 2025, 2025

  11. [11]

    O’Neill, A

    A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, et al. Open X-Embodiment: Robotic learning datasets and RT-X models. In2024 IEEE International Conference on Robotics and Automation (ICRA), pages 6892–6903, 2024

  12. [12]

    Zitkovich, T

    B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. In J. Tan, M. Toussaint, and K. Darvish, editors,Proceedings of The 7th Conference on Robot Learning, volume 229 ofProceedings of Machine Learning Research, pages 2165–2183. PMLR, 06–09 Nov 2023

  13. [13]

    B. Liu, Y . Zhu, C. Gao, Y . Feng, Q. Liu, Y . Zhu, and P. Stone. Libero: Benchmarking knowl- edge transfer for lifelong robot learning. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors,Advances in Neural Information Processing Systems, vol- ume 36, pages 44776–44791. Curran Associates, Inc., 2023

  14. [14]

    Nasiriany, A

    S. Nasiriany, A. Maddukuri, L. Zhang, A. Parikh, A. Lo, A. Joshi, A. Mandlekar, and Y . Zhu. Robocasa: Large-scale simulation of everyday tasks for generalist robots. InRobotics: Science and Systems (RSS), 2024

  15. [15]

    Nasiriany, S

    S. Nasiriany, S. Nasiriany, A. Maddukuri, and Y . Zhu. Robocasa365: A large-scale simulation framework for training and benchmarking generalist robots. InInternational Conference on Learning Representations (ICLR), 2026

  16. [16]

    Z. Chen, C. Gao, L. Shao, J. Shi, J. Huo, and Y . Gao. Manilong-shot: Interaction-aware one- shot imitation learning for long-horizon manipulation. InProceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 18189–18197, 2026

  17. [17]

    Y . Dai, H. Fu, J. Lee, Y . Liu, H. Zhang, J. Yang, C. Finn, N. Fazeli, and J. Chai. Robomme: Benchmarking and understanding memory for robotic generalist policies.arXiv preprint arXiv:2603.04639, 2026

  18. [18]

    S. K. Srinivas, Y . Shukla, A. Arnold, and S. Chitta. Graspfactory: A large object-centric grasping dataset.arXiv preprint arXiv:2509.20550, 2025

  19. [19]

    P. Li, H. Geng, J. Crate, Y . Han, J. Zhang, F. Wang, C. T. Cheng, R. Dong, Y .-J. Wang, H. Lou, et al. Rose: Reconstructing objects, scenes, and trajectories from casual videos for robotic manipulation. InNeurIPS 2025 Workshop on Bridging Language, Agent, and World Models for Reasoning and Planning, 2025

  20. [20]

    Y . Yin, Z. Han, S. Aarya, S. Xu, J. Wang, J. Peng, A. Wang, A. Yuille, and T. Shu. Partin- struct: Part-level instruction following for fine-grained robot manipulation. InProceedings of Robotics: Science and Systems (RSS), 2025

  21. [21]

    Mandlekar, Y

    A. Mandlekar, Y . Zhu, A. Garg, J. Booher, M. Spero, A. Tung, J. Gao, J. Emmons, A. Gupta, E. Orbay, S. Savarese, and L. Fei-Fei. RoboTurk: A crowdsourcing platform for robotic skill learning through imitation. InConference on Robot Learning (CoRL), 2018

  22. [22]

    Mirchandani, M

    S. Mirchandani, M. Tang, J. Duan, J. I. Hamid, M. Cho, and D. Sadigh. Robocade: Gamifying robot data collection.arXiv preprint arXiv:2512.21235, 2025

  23. [23]

    Kormushev, S

    P. Kormushev, S. Calinon, and D. G. Caldwell. Imitation learning of positional and force skills demonstrated via kinesthetic teaching and haptic input.Advanced Robotics, 25(5):581–603, 2011. 10

  24. [24]

    P. Wu, Y . Shentu, Z. Yi, X. Lin, and P. Abbeel. Gello: A general, low-cost, and intuitive teleoperation framework for robot manipulators.arXiv preprint arXiv:2309.13037, 2023

  25. [25]

    T. Z. Zhao, V . Kumar, S. Levine, and C. Finn. Learning fine-grained bimanual manipulation with low-cost hardware. InRobotics: Science and Systems, 2023

  26. [26]

    Y . Zou, Z. Zhou, C. Shi, Z. Ye, J. Huang, Y . Ding, and B. Zhao. U-arm: Ultra low-cost general teleoperation interface for robot manipulation.arXiv preprint arXiv:2509.02437, 2025

  27. [27]

    Honerkamp, H

    D. Honerkamp, H. Mahesheka, J. O. von Hartz, T. Welschehold, and A. Valada. Zero-cost whole-body teleoperation for mobile manipulation.IEEE Robotics and Automation Letters, 2025

  28. [28]

    C. Zhou, C. Peers, Y . Wan, R. Richardson, and D. Kanoulas. Teleman: Teleoperation for legged robot loco-manipulation using wearable imu-based motion capture.arXiv preprint arXiv:2209.10314, 2022

  29. [29]

    George, A

    A. George, A. Bartsch, and A. B. Farimani. Openvr: Teleoperation for manipulation.Soft- wareX, 29:102054, 2025

  30. [30]

    K. Lu, Y . He, C. Lu, and P. Li. I know kung fu: Synthetic dexterous hand demonstration collection via vr teleoperation. InNeurIPS 2025 Workshop on Space in Vision, Language, and Embodied AI, 2025

  31. [31]

    Wang, C.-C

    J. Wang, C.-C. Chang, J. Duan, D. Fox, and R. Krishna. Eve: Enabling anyone to train robot using augmented reality.arXiv preprint arXiv:2404.06089, 2024

  32. [32]

    C. Wang, H. Shi, W. Wang, R. Zhang, L. Fei-Fei, and C. K. Liu. Dexcap: Scalable and portable mocap data collection system for dexterous manipulation.arXiv preprint arXiv:2403.07788, 2024

  33. [33]

    M. Xu, H. Zhang, Y . Hou, Z. Xu, L. Fan, M. Veloso, and S. Song. Dexumi: Using hu- man hand as the universal manipulation interface for dexterous manipulation.arXiv preprint arXiv:2505.21864, 2025

  34. [34]

    C.-L. Fok, F. Sun, M. Mangum, A. K. Mok, B. He, and L. Sentis. Web-based teleoperation of a humanoid robot.arXiv preprint arXiv:1607.05402, 2016

  35. [35]

    Isaac sim, 2025

    NVIDIA. Isaac sim, 2025. Robotics simulator, accessed 2026-03-06

  36. [36]

    Xiang, Y

    F. Xiang, Y . Qin, K. Mo, Y . Xia, H. Zhu, F. Liu, M. Liu, H. Jiang, Y . Yuan, H. Wang, L. Yi, A. X. Chang, L. J. Guibas, and H. Su. Sapien: A simulated part-based interactive environment. arXiv preprint arXiv:2003.08515, 2020

  37. [37]

    Todorov, T

    E. Todorov, T. Erez, and Y . Tassa. Mujoco: A physics engine for model-based control. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, pages 5026–

  38. [38]

    Koenig and A

    N. Koenig and A. Howard. Design and use paradigms for gazebo, an open-source multi-robot simulator. InIEEE/RSJ International Conference on Intelligent Robots and Systems, pages 2149–2154, Sendai, Japan, Sept. 2004

  39. [39]

    Rohmer, S

    E. Rohmer, S. P. N. Singh, and M. Freese. Coppeliasim (formerly v-rep): A versatile and scal- able robot simulation framework. In2013 IEEE/RSJ International Conference on Intelligent Robots and Systems, 2013

  40. [40]

    Tedrake and the Drake Development Team

    R. Tedrake and the Drake Development Team. Drake: Model-based design and verification for robotics, 2019. 11

  41. [41]

    Coumans and Y

    E. Coumans and Y . Bai. Pybullet, a python module for physics simulation for games, robotics and machine learning, 2016. Accessed: 2026-03-06

  42. [42]

    Makoviychuk, L

    V . Makoviychuk, L. Wawrzyniak, Y . Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, and G. State. Isaac gym: High performance gpu-based physics simu- lation for robot learning. InProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, 2021. NeurIPS 2021 Datasets and Benchmarks Track

  43. [43]

    Zakka, B

    K. Zakka, B. Tabanpour, Q. Liao, M. Haiderbhai, S. Holt, J. Y . Luo, A. Allshire, E. Frey, K. Sreenath, L. A. Kahrs, C. Sferrazza, Y . Tassa, and P. Abbeel. Mujoco playground: An open- source framework for gpu-accelerated robot learning and sim-to-real transfer.arXiv preprint arXiv:2502.08844, 2025

  44. [44]

    C. D. Freeman, E. Frey, A. Raichuk, S. Girgin, I. Mordatch, and O. Bachem. Brax: A dif- ferentiable physics engine for large scale rigid body simulation. InProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, 2021. NeurIPS 2021 Datasets and Benchmarks Track

  45. [45]

    Genesis: A universal and generative physics engine for robotics and beyond,

    Genesis Authors. Genesis: A universal and generative physics engine for robotics and beyond,

  46. [46]

    J. Gu, F. Xiang, X. Li, Z. Ling, X. Liu, T. Mu, Y . Tang, S. Tao, X. Wei, Y . Yao, X. Yuan, P. Xie, Z. Huang, R. Chen, and H. Su. ManiSkill2: A unified benchmark for generalizable manipulation skills. InInternational Conference on Learning Representations (ICLR), 2023

  47. [47]

    S. Tao, F. Xiang, A. Shukla, Y . Qin, X. Hinrichsen, X. Yuan, C. Bao, X. Lin, Y . Liu, T. kai Chan, Y . Gao, X. Li, T. Mu, N. Xiao, A. Gurha, Z. Huang, R. Calandra, R. Chen, S. Luo, and H. Su. ManiSkill3: Gpu parallelized robotics simulation and rendering for generalizable embodied AI.arXiv preprint arXiv:2410.00425, 2024

  48. [48]

    Y . Zhu, J. Wong, A. Mandlekar, R. Mart ´ın-Mart´ın, A. Joshi, S. Nasiriany, and Y . Zhu. ro- bosuite: A modular simulation framework and benchmark for robot learning.arXiv preprint arXiv:2009.12293, 2020

  49. [49]

    Mandlekar, S

    A. Mandlekar, S. Nasiriany, B. Wen, I. Akinola, Y . Narang, L. Fan, Y . Zhu, and D. Fox. MimicGen: A data generation system for scalable robot learning using human demonstrations. InConference on Robot Learning (CoRL), 2023

  50. [50]

    Jiang, Y

    Z. Jiang, Y . Xie, K. Lin, Z. Xu, W. Wan, A. Mandlekar, L. Fan, and Y . Zhu. DexMimicGen: Automated data generation for bimanual dexterous manipulation via imitation learning.arXiv preprint arXiv:2410.24185, 2024

  51. [51]

    L. Wang, Y . Ling, Z. Yuan, M. Shridhar, C. Bao, Y . Qin, B. Wang, H. Xu, and X. Wang. GenSim: Generating robotic simulation tasks via large language models. InInternational Conference on Learning Representations (ICLR), 2024

  52. [52]

    P. Hua, M. Liu, A. Macaluso, Y . Lin, W. Zhang, H. Xu, and L. Wang. GenSim2: Scaling robot data generation with multi-modal and reasoning LLMs. InConference on Robot Learning (CoRL), 2024

  53. [53]

    Y . Wang, Z. Xian, F. Chen, T.-H. Wang, Y . Wang, K. Fragkiadaki, Z. Erickson, D. Held, and C. Gan. RoboGen: Towards unleashing infinite data for automated robot learning via generative simulation. InInternational Conference on Machine Learning (ICML), 2024

  54. [54]

    Y . He, X. Wang, P. Li, Y . Huang, J. Masterjohn, J. Wu, L. Guibas, Y . Yang, Y . Jiang, and C. Jiang. Fishbone: From one 3d asset to a million controllable edits.arXiv preprint arXiv:2605.24805, 2026. 12

  55. [55]

    Ghosh, H

    Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y . L. Tan, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine. Octo: An open-source generalist robot policy. InRobotics: Science and Systems (RSS), 2024

  56. [56]

    M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Fos- ter, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn. OpenVLA: An open-source vision-language-action model. InConfer- ence on Robot Learning (CoRL), 2024

  57. [57]

    M. J. Kim, C. Finn, and P. Liang. Fine-tuning vision-language-action models: Optimizing speed and success.arXiv preprint arXiv:2502.19645, 2025

  58. [58]

    S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu. RDT-1B: a diffusion foundation model for bimanual manipulation.arXiv preprint arXiv:2410.07864, 2024

  59. [59]

    Q. Li, Y . Liang, Z. Wang, L. Luo, X. Chen, M. Liao, F. Wei, Y . Deng, S. Xu, Y . Zhang, X. Wang, B. Liu, J. Fu, J. Bao, D. Chen, Y . Shi, J. Yang, and B. Guo. CogACT: A foundational vision-language-action model for synergizing cognition and action in robotic manipulation. arXiv preprint arXiv:2411.19650, 2024

  60. [60]

    H. Shi, B. Xie, Y . Liu, L. Sun, F. Liu, T. Wang, E. Zhou, H. Fan, X. Zhang, and G. Huang. Memoryvla: Perceptual-cognitive memory in vision-language-action models for robotic ma- nipulation.arXiv preprint arXiv:2508.19236, 2025

  61. [61]

    Black, N

    K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Haus- man, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky.π 0: A vision-language-action flow model for general robot control.arXiv preprint arXiv:2410....

  62. [62]

    Black, N

    Physical Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, M. Y . Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner...

  63. [63]

    Bjorck, F

    NVIDIA, J. Bjorck, F. C. neda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y . Fang, D. Fox, F. Hu, S. Huang, J. Jang, Z. Jiang, J. Kautz, K. Kundalia, L. Lao, Z. Li, Z. Lin, K. Lin, G. Liu, E. Llontop, L. Magne, A. Mandlekar, A. Narayan, S. Nasiriany, S. Reed, Y . L. Tan, G. Wang, Z. Wang, J. Wang, Q. Wang, J. Xiang, Y . Xie, Y . Xu, Z. Xu, S. Ye, Z. Yu, A. Z...

  64. [64]

    B. Chen, T. Zhang, H. Geng, C. Zhang, P. Li, K. Song, W. T. Freeman, J. Malik, P. Abbeel, R. Tedrake, et al. Large video planner enables generalizable robot control.arXiv preprint arXiv:2512.15840, 2025

  65. [65]

    S. Fei, S. Wang, J. Shi, Z. Dai, J. Cai, P. Qian, L. Ji, X. He, S. Zhang, Z. Fei, J. Fu, J. Gong, and X. Qiu. LIBERO-Plus: In-depth robustness analysis of vision-language-action models.arXiv preprint arXiv:2510.13626, 2025

  66. [66]

    X. Li, K. Hsu, J. Gu, K. Pertsch, O. Mees, H. R. Walke, C. Fu, I. Lunawat, I. Sieh, S. Kir- mani, S. Levine, J. Wu, C. Finn, H. Su, Q. Vuong, and T. Xiao. Evaluating real-world robot manipulation policies in simulation.arXiv preprint arXiv:2405.05941, 2024

  67. [67]

    Pick up the toy car on the table

    W. Pumacay, I. Singh, J. Duan, R. Krishna, J. Thomason, and D. Fox. THE COLOSSEUM: A benchmark for evaluating generalization for robotic manipulation. InRobotics: Science and Systems (RSS), 2024. 13 Appendix Contents A Dataset Comparison 15 B Web-Infra: Browser-Based Infrastructure 16 C TaskGen: Task Generation and Deployment 17 D Data Cleaning and Trajec...

  68. [2024]

    Software project, accessed 2026-03-06