Pith. sign in

REVIEW 3 major objections 6 minor 6 cited by

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

T0 review · 3 major / 6 minor · reviewed 2026-07-11 · grok-4.5

Pith's one-line read A unified sim-and-real benchmark shows generalist robot policies still fail most hard manipulation tasks.

desk verdict Solid systems/benchmark paper: real shared sim+real infrastructure and a 30-policy leaderboard that honestly shows how far generalist manip is from human teleop; novelty is the package, not any single axis. read the letter →

arxiv 2607.04434 v1 pith:UV7JDS4H submitted 2026-07-05 cs.RO cs.AIcs.CVcs.GR

classification cs.ROcs.AIcs.CVcs.GR
keywords robotmanipulationgeneralistpoliciessim-and-realbenchmarkvision-language-actionmodelslong-horizonprecisionheterogeneousparallelsimulationleaderboard
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

RoboDojo argues that existing robot-manipulation benchmarks are too narrow, too short-horizon, or confined to either simulation or the real world alone, so they cannot systematically diagnose what generalist policies can and cannot do. The authors introduce a single evaluation loop with 42 simulation tasks organized into five capability dimensions—generalization, memory, precision, long-horizon execution, and open-vocabulary instruction following—plus 18 real-world tasks on three bimanual robot platforms under a standardized, reproducible physical setup. Shared infrastructure lets a policy be integrated once and then run at scale in heterogeneous parallel simulation and in remote real-robot trials. After evaluating 30 policies, the paper reports that even the strongest models succeed on only a small fraction of episodes and remain far below expert human teleoperation, with open-semantic, memory, and precision tasks especially weak and with simulation rank only partly predicting real-world deployability.

What carries the argument

RoboDojo—a unified sim-and-real evaluation system whose simulation half stresses five complementary capability dimensions via heterogeneous parallel Isaac Sim, whose real half (RoboDojo-RealEval) standardizes hardware, layout replay, scoring, and remote access across three embodiments, and whose XPolicyLab interface lets policies be integrated once for both settings.

What would settle it

A new policy that, after the same training protocol, reaches high success on all five simulation dimensions and on the full 18-task multi-embodiment real suite, with ranks that stay consistent under hidden layouts and independent re-runs, would falsify the claimed large, systematic gap.

Watch

Extended reading notes

Core claim

Current generalist robot manipulation policies remain far from reliable, balanced performance: the best of 30 evaluated models reaches only about 9 percent average success in simulation and about 13 percent in the real world, versus human teleoperation near or at 100 percent under the same protocols, and strengths on one capability dimension rarely transfer to the others.

Load-bearing premise

That these five hand-designed simulation dimensions plus a compact, non-paired real-world suite under standardized but human-scored trials are fair and sufficient proxies for comprehensive generalist capability and deployability.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. RoboDojo proposes a unified sim-and-real benchmark for generalist robot manipulation: 42 Isaac Sim tasks organized into five capability dimensions (Generalization, Memory, Precision, Long-Horizon, Open) and 18 real-world tasks on three bimanual embodiments (ARX X5, Piper, Piper X). The system contribution includes heterogeneous parallel simulation, RoboDojo-RealEval (standardized hardware, layout replay, remote evaluation), and XPolicyLab (shared policy interface). The authors integrate and evaluate 30 policies in simulation and 10 in the real world, report multi-seed/multi-rater protocols, efficiency and stability studies, and a public leaderboard showing large gaps to expert teleoperation (e.g., best sim average success ~8.8% vs human ~76%; best real overall success ~12.8% vs human 100%).

Significance. If accepted as a community evaluation standard, this is a substantial systems and benchmarking contribution for embodied manipulation. Strengths that should be credited explicitly include: (i) scale and breadth of the evaluation (42 sim + 18 real tasks; 30 integrated policies via XPolicyLab); (ii) multi-seed simulation reporting, double-blind multi-rater real scoring, and documented anti-gaming/hidden-layout rules (§3.3, App. A.2); (iii) measured evaluation efficiency for heterogeneous parallelism (Table 4) and real-world wall-clock cost (Table 5); and (iv) GPU and repeated-run stability checks (§6.4). The diagnostic findings—especially Standard vs Random collapse (Table 3), precision/memory/open bottlenecks, and partial sim–real rank misalignment—are useful for the field even if the suite is not a universal proxy for all deployment settings.

major comments (3)
  1. [Appendix K / Table 1] Appendix K and Table 1: cross-policy rankings are difficult to interpret under highly heterogeneous fine-tuning budgets and training regimes. Policies differ by orders of magnitude in steps/epochs, batch size, and initialization (e.g., Hy-Embodied-0.5-VLA 200K steps vs LingBot-VLA 15K vs ACT 6K; single-task ACT/SmolVLA mixed with multi-task VLAs; some models not trained with three seeds). Because the central empirical claim is comparative leaderboard performance and capability diagnosis, the paper should either (a) standardize a primary training protocol for official ranking, or (b) clearly demote Table 1 to a protocol-conditioned snapshot and report compute-matched or budget-normalized ablations for the leading group.
  2. [§3.2.1 / §6.2 Finding 2] Sections 3.2.1 and 6.2 Finding 2: the manuscript correctly states that sim and real tasks are not one-to-one matched and are not a sim-to-real transfer benchmark, yet the abstract/introduction still frame RoboDojo as comprehensively diagnosing generalist capability and deployability. With only 18 compact real tasks, human-scored partial credit, and unpaired distributions, the paper should bound what can and cannot be concluded (e.g., that real rankings measure protocol-specific physical executability, not transfer of the five sim dimensions). A short limitations subsection quantifying this scope would make the central claim proportionate to the evidence.
  3. [§6.4.2 / Table 7 / Table 2] Section 6.4.2 and Table 7 (with App. Table 9): real-world aggregate stability is good (overall SR std ≤1.3 pp), but several tasks show large trial-to-trial variance (e.g., store_in_safe SR std 23.1 pp for π0.5). With only 10 trials per task and embodiment-specific averages used for ranking (Table 2), embodiment-level and task-level orderings may be noise-sensitive. Please report confidence intervals or bootstrap uncertainty for Table 2 aggregates, and clarify whether leaderboard ranks use only overall average or also per-embodiment thresholds.
minor comments (6)
  1. [Table 8 / §5] Table 8 claims 60 tasks and 35+ policies, while the main text consistently states 42+18=60 tasks and 30 integrated policies (10 real). Align the comparison table with the main claims and freeze date.
  2. [§3.1.1 / Table 1 / App. I.2] The Open dimension is near floor for all policies (Table 1). Briefly discuss whether this primarily diagnoses open-semantic grounding failure or under-specified/out-of-distribution task design relative to the training skill set (§3.1.1 Open; App. I.2).
  3. [§3.1.3 / §3.2.3 / App. E.2] Clarify how average score is computed for multi-step tasks (partial progress weights) in both sim and real; App. E.2 describes multi-rater averaging but not the sub-step rubric weights used for the reported scores.
  4. [Figure 2 / Figure 3] Figure 2 and task lists are dense; a compact table of task→dimension→skill mapping (beyond Fig. 3) would improve navigability for readers using the suite.
  5. [Front matter / §5] Minor consistency: dates appear as 2026 throughout (arXiv stamp, leaderboard freeze). Ensure camera-ready metadata matches the intended publication year and that website/code links remain stable.
  6. [Table 3] In Table 3, OpenVLA-OFT shows a −100% relative drop due to near-zero Standard score; either exclude relative drop when Standard is near zero or mark it as undefined to avoid misleading comparison.

Circularity Check

0 steps flagged · score 0.0 of 10

Empirical benchmark paper with no derivation chain that recovers its own inputs; leaderboard gaps are measured outcomes under stated protocols, not fitted-or-self-defined predictions.

full rationale

RoboDojo is a constructive systems/benchmark paper, not a first-principles derivation. Its central claims are that the suite (42 sim tasks across five hand-designed dimensions, 18 real tasks on three embodiments), infrastructure (heterogeneous Isaac Sim, RoboDojo-RealEval, XPolicyLab), and 30-policy leaderboard exist and that evaluated policies score far below human teleoperation under the stated protocols (Tables 1–2; Secs. 3–6). Policies are trained on released demos and scored on held-out episodes/layouts; the Open dimension is explicitly train-excluded (3.1.2); public layouts plus hidden verification and multi-seed/multi-embodiment rules are stated anti-gaming measures (3.3, A.2). Self-citations to prior author work (RoboTwin, RMBench, etc.) appear only as related-work context and design inspiration, not as uniqueness theorems or load-bearing premises that force the reported gaps. There are no fitted constants renamed as predictions, no self-definitional equations, and no ansatz smuggled in as external fact. Residual community self-evaluation (authors integrate models and maintain the leaderboard) is normal benchmark practice and does not reduce the measured success rates to inputs by construction. Score 0 is therefore appropriate.

Assumptions & free parameters 5 free parameters · 5 assumptions · 4 invented entities

Load-bearing content is definitional and operational rather than mathematical: task taxonomies, success/score rubrics, hardware standardization, and evaluation protocols. Free parameters are design choices (episode counts, horizon multipliers, clutter caps, training steps) that shape difficulty and rankings. No new physical entities are postulated; invented entities are named infrastructure components.

free parameters (5)
  • Evaluation episode counts (50 sim / 10 real per task)
    Chosen sample sizes that determine ranking stability; real task-level variance remains large under 10 trials.
  • Horizon multipliers (1.2× or 1.5× of 90th-percentile demo length)
    Hand-set execution budgets that can truncate slow but correct policies or allow excessive recovery time.
  • Generalization clutter/randomization scale (up to 25 clutter objects; Standard/Random split)
    Difficulty knobs that strongly drive the reported generalization collapse relative to prior suites like RoboTwin.
  • Per-policy fine-tuning schedules (batch size, steps/epochs in Appendix K)
    Training budgets differ across models and can confound architecture comparisons on the leaderboard.
  • Aggregate metric as mean over five dimensions (not over all tasks)
    Aggregation choice that equalizes dimensions with unequal task counts and affects overall ranking.
assumptions (5)
  • domain assumption Capability-oriented task dimensions (Generalization, Memory, Precision, Long-Horizon, Open) are distinct and jointly sufficient for comprehensive diagnosis of generalist manipulation.
    Section 3.1 organizes the sim suite around these five axes; conclusions about 'balanced generalist' progress depend on this taxonomy.
  • domain assumption Standardized RealEval hardware, lighting, layout replay, and multi-rater scoring make real-world scores comparable across policies and sessions.
    Sections 3.2–4.2 and 6.4 treat protocol standardization as the basis for reproducible physical evaluation.
  • ad hoc to paper Unpaired sim and real tasks still jointly diagnose deployability even without matched sim-to-real transfer pairs.
    Explicit design choice in 3.2.1 and Finding 2 of 6.2; ranking mismatches are interpreted as complementary stress tests rather than transfer failures on matched tasks.
  • domain assumption Human teleoperation under the same horizons/success criteria is a valid near-ceiling reference for task feasibility.
    Sections 5.1–5.2 and Tables 1–2 use expert teleop as the human reference excluded from policy ranking.
  • domain assumption Isaac Sim / Isaac Lab physics and rendering are adequate for capability diagnosis despite residual nondeterminism.
    Platform choice in Section 4.1; stability study (Table 6) assumes residual GPU variance is small enough for fair leaderboards.
invented entities (4)
  • RoboDojo benchmark suite (42 sim + 18 real tasks) independent evidence
    purpose: Provide multi-axis sim diagnosis and multi-embodiment real evaluation under one protocol.
    Core contribution; existence is definitional to the paper's measurements.
  • RoboDojo-RealEval platform independent evidence
    purpose: Standardize real hardware, reset, remote access, and scoring for reproducible physical tests.
    Named physical/software system; independent evidence is the described hardware and protocol, not an external physics prediction.
  • XPolicyLab independent evidence
    purpose: Unify data conversion, policy interface, and deployment so models integrate once for sim and real.
    Software infrastructure enabling the 30-policy leaderboard.
  • Heterogeneous parallel simulation mode independent evidence
    purpose: Run diverse scenes/tasks concurrently under a vectorized interface for higher evaluation throughput.
    Implementation claim supported by efficiency experiments in Section 6.3.1.

how reviews work

0 comments
Cite this review

Pith. "Pith review of RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies." pith.science (2026). https://pith.science/paper/UV7JDS4H

@misc{pith2026260704434,
  author       = {Pith},
  title        = {Pith review of: RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UV7JDS4H}},
  note         = {Machine review of arXiv:2607.04434}
}
read the original abstract

Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their capabilities. Many rely on simple, short-horizon, or skill-narrow tasks with limited capability coverage, and are often conducted only in simulation or only in the real world. Simulation enables scalable feedback but misses physical deployment challenges, while real-world evaluation is costly, time-consuming, and difficult to reproduce. We introduce RoboDojo, a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies. RoboDojo includes 42 simulation tasks and 18 real-world tasks covering diverse and complementary manipulation capabilities. The simulation benchmark evaluates five dimensions: generalization, memory, precision, long-horizon execution, and open-vocabulary instruction following, while the real-world benchmark exposes policies to challenging physical-world deployment conditions. RoboDojo supports scalable evaluation through heterogeneous parallel simulation in Isaac Sim and provides RoboDojo-RealEval, a reproducible real-world evaluation system with remote cloud access, standardized hardware, scene reset, evaluation protocol, and deployment interface. Together with XPolicyLab, policies can be integrated once and evaluated across simulation and real-world settings with minimal adaptation. We integrate 30 policies into XPolicyLab and evaluate them on RoboDojo, establishing a public leaderboard and systematic analysis of current policy performance. The website is available at http://robodojo-benchmark.com/.

Figures

Figures reproduced from arXiv: 2607.04434 by the authors.

Figure 1
Figure 1. Overview of RoboDojo. RoboDojo unifies efficient simulation evaluation and reproducible real-world testing for generalist robot manipulation, covering 42 simulation tasks, 18 real-world tasks, heterogeneous parallel simulation, RoboDojo-RealEval, XPolicyLab, and a continuously updated leaderboard. 1 arXiv:2607.04434v1 [cs.RO] 5 Jul 2026 [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Task Overview of RoboDojo. RoboDojo includes 42 simulation tasks and 18 real-world tasks for evaluating generalist robot manipulation policies. The simulation tasks are organized into five capability dimensions: Generalization, Memory, Long-Horizon, Precision, and Open, enabling efficient capability-oriented diagnosis. The real-world tasks assess policy behavior under challenging and reproducible physical deployment… view at source ↗
Figure 3
Figure 3. Skill Diversity in RoboDojo. Representative simulation tasks cover 24 manipulation skills, including grasping, placing, pushing, pulling, stacking, insertion, opening, closing, folding, alignment, tool use, and contact-sensitive operations. These tasks span diverse spatial configurations and bimanual coordination patterns, enabling skill-level diagnosis beyond repetitive pick-and-place behaviors. Open. Generalist ro… view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Real-World Task Highlights. Representative key frames of challenging real-world manipulation tasks across different robot embodiments. 3.2. Real-World Benchmark To evaluate policy performance under physical deployment conditions, we construct the RoboDojo Real-World Be…
Figure 5
Figure 5. Figure 5: Assets and parallelism in RoboDojo. RoboDojo combines physically grounded assets with heterogeneous parallel simulation for scalable benchmark construction and evaluation. 4.1.3. Heterogeneous Parallelism Efficient benchmark evaluation requires parallel simulation over…
Figure 6
Figure 6. Figure 6: Overview of the RoboDojo-RealEval system. RoboDojo-RealEval provides a standardized physical platform for reproducible real-world robot manipulation evaluation, with controlled workspace geometry, fixed robot and camera mounts, stable lighting, a touchscreen evaluation…
Figure 7
Figure 7. Figure 7: Domain randomization in RoboDojo. We visualize the effects of domain randomization in simulation, including variations in background, lighting, clutter layout, object appearance, and scene configuration. These randomized environments increase visual diversity and reduc…
Figure 8
Figure 8. Figure 8: Future extensions of RoboDojo. RoboDojo will be continuously expanded to broader manipulation scenarios and embodiments, including dexterous hand manipulation, humanoid whole-body manipulation, tactile manipulation, and mobile manipulation. RoboDojo is designed as an e…
Figure 9
Figure 9. Figure 9: shows the hardware components of the RoboDojo-RealEval platform. The platform integrates wrist cameras, a head camera, replaceable collaborative bimanual robot embodiments, a robot and camera support structure, an external frame with controlled lighting and curtains, a…

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. $N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A scaled tactile-native world-action model jointly predicts vision, touch, and action and outperforms vision-only baselines on contact-rich sim and real robot tasks.

  2. Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Pre-training a VLA model on 100k hours of auto-labeled UMI trajectories, then post-training on robot data, yields SOTA simulated manipulation and data-efficient fine-tuning.

  3. Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels

    cs.RO 2026-07 conditional novelty 5.0 of 10

    A four-layer systems framework and T0–T5 hierarchy for grading and maintaining bounded trustworthiness claims in embodied AI systems.

  4. ArmnetBench v0.1: Parallel Real-World Evaluation of Manipulation Policies on a Low-Cost Arm Farm

    cs.RO 2026-07 conditional novelty 5.0 of 10

    A managed low-cost SO-101 arm farm evaluates seven manipulation policies across twelve single-arm and bimanual tasks, releasing 3,118 graded real-world episodes under a shared training budget.

  5. Progress Reward Modeling for Robotic Learning: A Comprehensive Survey

    cs.RO 2026-07 conditional novelty 5.0 of 10

    A survey organizing progress reward modeling for robot learning into interface, method, and data/evaluation layers, with a four-family method taxonomy and a cautionary split between progress fidelity and downstream utility.

  6. Data Pyramid for Embodied Manipulation

    cs.RO 2026-07 conditional novelty 3.0 of 10

    Embodied training data form a five-layer pyramid—real-robot, UMI, ego/exo, simulation, general V–L—ordered by the trade-off between scale and robot alignment, and model capabilities track how those layers are mixed.

Reference graph

Works this paper leans on

78 extracted references · 43 linked inside Pith · cited by 6 Pith papers

  1. [1]

    Roboarena: Distributed real-world evaluation of generalist robot policies.arXiv preprint arXiv:2506.18123, 2025

    Pranav Atreya, Karl Pertsch, Tony Lee, Moo Jin Kim, Arhan Jain, Artur Kuramshin, Clemens Eppner, Cyrus Neary, Edward Hu, Fabio Ramos, et al. Roboarena: Distributed real-world evaluation of generalist robot policies.arXiv preprint arXiv:2506.18123, 2025

  2. [2]

    Motus: A unified latent action world model

    Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, Shuhe Huang, Haitian Liu, Ruowen Zhao, Yao Feng, Chendong Xiang, Yinze Rong, et al. Motus: A unified latent action world model. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 35101–35113, 2026a

  3. [3]

    H-rdt: Human manipulation enhanced bimanual robotic manipulation

    Hongzhe Bi, Lingxuan Wu, Tianwei Lin, Hengkai Tan, Zhizhong Su, Hang Su, and Jun Zhu. H-rdt: Human manipulation enhanced bimanual robotic manipulation. InProceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 18135–18143, 2026b

  4. [4]

    Gr00t n1: An open foundation model for generalist humanoid robots.arXiv preprint arXiv:2503.14734, 2025

    Johan Bjorck, Fernando Castañeda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, et al. Gr00t n1: An open foundation model for generalist humanoid robots.arXiv preprint arXiv:2503.14734, 2025

  5. [5]

    pi0: A vision-language-action flow model for general robot control.arXiv preprint arXiv:2410.24164, 2024

    Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, et al. pi0: A vision-language-action flow model for general robot control.arXiv preprint arXiv:2410.24164, 2024

  6. [6]

    Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems.arXiv preprint arXiv:2503.06669, 2025

    Qingwen Bu, Jisong Cai, Li Chen, Xiuqi Cui, Yan Ding, Siyuan Feng, Shenyuan Gao, Xindong He, Xuan Hu, Xu Huang, et al. Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems.arXiv preprint arXiv:2503.06669, 2025

  7. [7]

    Aha-wam: Asynchronous horizon-adaptive world-action modeling with observation-guided context routing.arXiv preprint arXiv:2606.09811, 2026a

    Jisong Cai, Long Ling, Shiwei Chu, Zhongshan Liu, Jiayue Kang, Zhixuan Liang, Wenjie Xu, Yinan Mao, Weinan Zhang, Xiaokang Yang, et al. Aha-wam: Asynchronous horizon-adaptive world-action modeling with observation-guided context routing.arXiv preprint arXiv:2606.09811, 2026a

  8. [8]

    Internvla-a1: Unifying understanding, generation and action for robotic manipulation.arXiv preprint arXiv:2601.02456, 2026b

    Junhao Cai, Zetao Cai, Jiafei Cao, Yilun Chen, Zeyu He, Lei Jiang, Hang Li, Hengjie Li, Yang Li, Yufei Liu, et al. Internvla-a1: Unifying understanding, generation and action for robotic manipulation.arXiv preprint arXiv:2601.02456, 2026b

Show all 78 references
  1. [9]

    Xiaomi- robotics-0: An open-sourced vision-language-action model with real-time execution.arXiv preprint arXiv:2602.12684, 2026c

    Rui Cai, Jun Guo, Xinze He, Piaopiao Jin, Jie Li, Bingxuan Lin, Futeng Liu, Wei Liu, Fei Ma, Kun Ma, et al. Xiaomi- robotics-0: An open-sourced vision-language-action model with real-time execution.arXiv preprint arXiv:2602.12684, 2026c

  2. [10]

    Pink: Python inverse kinematics based on Pinocchio, 2026

    Stéphane Caron, Yann De Mont-Marin, Rohan Budhiraja, Seung Hyeon Bang, Ivan Domrachev, Simeon Nedelchev, Peter Du, Adrien Escande, Joris Vaillant, Bruce Wingo, Santosh Patapati, Daniel San José Pro, and Nicolas Guillermo Marticorena Vidal. Pink: Python inverse kinematics based...

  3. [11]

    Univtac: A unified simulation platform for visuo-tactile manipulation data generation, learning, and benchmarking.arXiv preprint arXiv:2602.10093, 2026a

    Baijun Chen, Weijie Wan, Tianxing Chen, Xianda Guo, Congsheng Xu, Yuanyang Qi, Haojie Zhang, Longyan Wu, Tianling Xu, Zixuan Li, et al. Univtac: A unified simulation platform for visuo-tactile manipulation data generation, learning, and benchmarking.arXiv preprint arXiv:2602.1...

  4. [12]

    Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation.arXiv preprint arXiv:2506.18088, 2025a

    Tianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai, Yibin Liu, Zixuan Li, Qiwei Liang, Xianliang Lin, Yiheng Ge, Zhenyu Gu, et al. Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation.arXiv preprint ar...

  5. [13]

    G3flow: Generative 3d semantic flow for pose-aware and generalizable object manipulation

    Tianxing Chen, Yao Mu, Zhixuan Liang, Zanxin Chen, Shijia Peng, Qiangyu Chen, Mingkun Xu, Ruizhen Hu, Hongyuan Zhang, Xuelong Li, et al. G3flow: Generative 3d semantic flow for pose-aware and generalizable object manipulation. InProceedings of the Computer Vision and Pattern R...

  6. [14]

    Benchmarking generalizable bimanual manipulation: Robotwin dual-arm collaboration challenge at cvpr 2025 meis workshop.arXiv preprint arXiv:2506.23351, 2025c

    Tianxing Chen, Kaixuan Wang, Zhaohui Yang, Yuhao Zhang, Zanxin Chen, Baijun Chen, Wanxi Dong, Ziyuan Liu, Dong Chen, Tianshuo Yang, et al. Benchmarking generalizable bimanual manipulation: Robotwin dual-arm collaboration challenge at cvpr 2025 meis workshop.arXiv preprint arXi...

  7. [15]

    Rmbench: Memory-dependent robotic manipulation benchmark with insights into policy design.arXiv preprint arXiv:2603.01229, 2026b

    Tianxing Chen, Yuran Wang, Mingleyang Li, Yan Qin, Hao Shi, Zixuan Li, Yifan Hu, Yingsheng Zhang, Kaixuan Wang, Yue Chen, et al. Rmbench: Memory-dependent robotic manipulation benchmark with insights into policy design.arXiv preprint arXiv:2603.01229, 2026b. 21 RoboDojo: A Uni...

  8. [16]

    Yiting Chen, Kenneth Kimble, Edward H Adelson, Tamim Asfour, Podshara Chanrungmaneekul, Sachin Chitta, Yash Chitambar, Ziyang Chen, Ken Goldberg, Danica Kragic, et al. Manipulationnet: An infrastructure for benchmarking real-world robot manipulation with physical skill challen...

  9. [17]

    Diffusion Policy: Visuomotor Policy Learning via Action Diffusion, March 2024

    Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion, March 2024

  10. [18]

    Molmoact2: Action reasoning models for real-world deployment.arXiv preprint arXiv:2605.02881, 2026

    Haoquan Fang, Jiafei Duan, Donovan Clay, Sam Wang, Shuo Liu, Weikai Huang, Xiang Fan, Wei-Chuan Tsai, Shirui Chen, Yi Ru Wang, et al. Molmoact2: Action reasoning models for real-world deployment.arXiv preprint arXiv:2605.02881, 2026

  11. [19]

    Galaxea g0.5 technical report

    Galaxea Team. Galaxea g0.5 technical report. 2026. URLhttps://opengalaxea.github.io/G05/

  12. [20]

    Ebench: Elemental diagnosis of generalist mobile manipulation policies.arXiv preprint arXiv:2606.18239, 2026

    Ning Gao, Jinliang Zheng, Xing Gao, Haoxiang Ma, Hanqing Wang, Yukai Wang, Jiantong Chen, Zanxin Chen, Shujie Zhang, Mingda Jia, et al. Ebench: Elemental diagnosis of generalist mobile manipulation policies.arXiv preprint arXiv:2606.18239, 2026

  13. [21]

    Unified 4d world action modeling from video priors with asynchronous denoising, 2026

    Jun Guo, Qiwei Li, Peiyan Li, Zilong Chen, Nan Sun, Yifei Su, Heyun Wang, Yuan Zhang, Xinghang Li, and Huaping Liu. Unified 4d world action modeling from video priors with asynchronous denoising, 2026. URL https: //arxiv.org/abs/2604.26694

  14. [22]

    pi0.5: A vision-language-action model with open-world generalization.arXiv preprint arXiv:2504.16054, 2025

    Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, et al. pi0.5: A vision-language-action model with open-world generalization.arXiv preprint arXiv:2504.16054, 2025

  15. [24]

    Pi 0.7: a steerable generalist robotic foundation model with emergent capabilities.arXiv preprint arXiv:2604.15483, 2026b

    Physical Intelligence, Bo Ai, Ali Amin, Raichelle Aniceto, Ashwin Balakrishna, Greg Balke, Kevin Black, George Bokinsky, Shihao Cao, Thomas Charbonnier, et al. Pi 0.7: a steerable generalist robotic foundation model with emergent capabilities.arXiv preprint arXiv:2604.15483, 2026b

  16. [25]

    Stephen James, Zicong Ma, David Rovick Arrojo, and Andrew J. Davison. Rlbench: The robot learning benchmark & learning environment, 2019. URLhttps://arxiv.org/abs/1909.12271

  17. [26]

    Galaxea open-world dataset and g0 dual-system vla model.arXiv preprint arXiv:2509.00576, 2025

    Tao Jiang, Tianyuan Yuan, Yicheng Liu, Chenhao Lu, Jianning Cui, Xiao Liu, Shuiqi Cheng, Jiyang Gao, Huazhe Xu, and Hang Zhao. Galaxea open-world dataset and g0 dual-system vla model.arXiv preprint arXiv:2509.00576, 2025

  18. [27]

    Openvla: An open-source vision-language-action model.arXiv preprint arXiv:2406.09246, 2024

    Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al. Openvla: An open-source vision-language-action model.arXiv preprint arXiv:2406.09246, 2024

  19. [28]

    Fine-tuning vision-language-action models: Optimizing speed and success

    Moo Jin Kim, Chelsea Finn, and Percy Liang. Fine-tuning vision-language-action models: Optimizing speed and success. arXiv preprint arXiv:2502.19645, 2025

  20. [29]

    Autobio: A simulation and benchmark for robotic automation in digital biology laboratory.arXiv preprint arXiv:2505.14030, 2025

    Zhiqian Lan, Yuxuan Jiang, Ruiqi Wang, Xuanbing Xie, Rongkui Zhang, Yicheng Zhu, Peihang Li, Tianshuo Yang, Tianxing Chen, Haoyu Gao, et al. Autobio: A simulation and benchmark for robotic automation in digital biology laboratory.arXiv preprint arXiv:2505.14030, 2025. 22 RoboD...

  21. [30]

    Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation

    Chengshu Li, Ruohan Zhang, Josiah Wong, Cem Gokmen, Sanjana Srivastava, Roberto Martín-Martín, Chen Wang, Gabrael Levine, Michael Lingelbach, Jiankai Sun, et al. Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation. InConference on R...

  22. [31]

    Spatial forcing: Implicit spatial representation alignment for vision-language-action model.arXiv preprint arXiv:2510.12276, 2025a

    Fuhao Li, Wenxuan Song, Han Zhao, Jingbo Wang, Pengxiang Ding, Donglin Wang, Long Zeng, and Haoang Li. Spatial forcing: Implicit spatial representation alignment for vision-language-action model.arXiv preprint arXiv:2510.12276, 2025a

  23. [32]

    Simplevla-rl: Scaling vla training via reinforcement learning.arXiv preprint arXiv:2509.09674, 2025b

    Haozhan Li, Yuxin Zuo, Jiale Yu, Yuhao Zhang, Zhaohui Yang, Kaiyan Zhang, Xuekai Zhu, Yuchen Zhang, Tianxing Chen, Ganqu Cui, et al. Simplevla-rl: Scaling vla training via reinforcement learning.arXiv preprint arXiv:2509.09674, 2025b

  24. [33]

    Causal world modeling for robot control, 2026a

    Lin Li, Qihang Zhang, Yiming Luo, Shuai Yang, Ruilin Wang, Fei Han, Mingrui Yu, Zelin Gao, Nan Xue, Xing Zhu, Yujun Shen, and Yinghao Xu. Causal world modeling for robot control, 2026a. URL https://arxiv.org/abs/2601. 21998

  25. [34]

    Wall-wm: Carving world action modeling at the event joints.arXiv preprint arXiv:2606.01955, 2026b

    Shalfun Li, Victor Yao, Charles Yang, Truth Qu, Regis Cheng, Ryan Yu, Howard Lu, Newton V on, Vincent Chen, Yohann Tang, et al. Wall-wm: Carving world action modeling at the event joints.arXiv preprint arXiv:2606.01955, 2026b

  26. [36]

    Evaluating real-world robot manipulation policies in simulation.arXiv preprint arXiv:2405.05941, 2024b

    Xuanlin Li, Kyle Hsu, Jiayuan Gu, Karl Pertsch, Oier Mees, Homer Rich Walke, Chuyuan Fu, Ishikaa Lunawat, Isabel Sieh, Sean Kirmani, et al. Evaluating real-world robot manipulation policies in simulation.arXiv preprint arXiv:2405.05941, 2024b

  27. [37]

    Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023

    Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023

  28. [38]

    Rdt-1b: a diffusion foundation model for bimanual manipulation

    Songming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan, Huayu Chen, Zhengyi Wang, Ke Xu, Hang Su, and Jun Zhu. Rdt-1b: a diffusion foundation model for bimanual manipulation. InInternational Conference on Learning Representations, volume 2025, pages 29982–30009, 2025a

  29. [39]

    Avr: Active vision-driven robotic precision manipulation with viewpoint and focal length optimization

    Yushan Liu, Shilong Mu, Xintao Chao, Zizhen Li, Yao Mu, Tianxing Chen, Shoujie Li, Chuqiao Lyu, Xiao-ping Zhang, and Wenbo Ding. Avr: Active vision-driven robotic precision manipulation with viewpoint and focal length optimization. arXiv e-prints, pages arXiv–2503, 2025b

  30. [40]

    Garmentlab: A unified simulation and benchmark for garment manipulation

    Haoran Lu, Ruihai Wu, Yitong Li, Sijie Li, Ziyu Zhu, Chuanruo Ning, Yan Shen, Longzan Luo, Yuanpei Chen, and Hao Dong. Garmentlab: A unified simulation and benchmark for garment manipulation. InAdvances in Neural Information Processing Systems, 2024

  31. [41]

    Magicsim: A unified infrastructure for executable embodied interaction, 2026

    Haoran Lu, Songling Liu, Yue Chen, Guo Ye, Mutian Shen, Shuyang Yu, Yu Xiao, Jihai Zhao, Shang Wu, Jianshu Zhang, Xiangtian Gui, Chuye Hong, Yuran Wang, Maojiang Su, Jiayi Wang, Ruihai Wu, Zhaoran Wang, and Han Liu. Magicsim: A unified infrastructure for executable embodied in...

  32. [42]

    Lda-1b: Scaling latent dynamics action model via universal embodied data ingestion.arXiv preprint arXiv:2602.12215, 2026

    Jiangran Lyu, Kai Liu, Xuheng Zhang, Haoran Liao, Yusen Feng, Wenxuan Zhu, Tingrui Shen, Jiayi Chen, Jiazhao Zhang, Yifei Dong, et al. Lda-1b: Scaling latent dynamics action model via universal embodied data ingestion.arXiv preprint arXiv:2602.12215, 2026

  33. [43]

    Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters, 7(3):7327–7334, 2022

    Oier Mees, Lukas Hermann, Erick Rosete-Beas, and Wolfram Burgard. Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters, 7(3):7327–7334, 2022

  34. [44]

    Carlson, Ji Yuan Feng, Animesh Garg, Renato Gasoto, Lionel Gulich, Yijie Guo, M

    Mayank Mittal, Pascal Roth, James Tigue, Antoine Richard, Octi Zhang, Peter Du, Antonio Serrano-Muñoz, Xinjie Yao, René Zurbrügg, Nikita Rudin, Lukasz Wawrzyniak, Milad Rakhsha, Alain Denzler, Eric Heiden, Ales Borovicka, Ossama Ahmed, Iretiayo Akinola, Abrar Anwar, Mark T. Ca...

  35. [45]

    Robotwin: Dual-arm robot benchmark with generative digital twins

    Yao Mu, Tianxing Chen, Zanxin Chen, Shijia Peng, Zhiqian Lan, Zeyu Gao, Zhixuan Liang, Qiaojun Yu, Yude Zou, Mingkun Xu, et al. Robotwin: Dual-arm robot benchmark with generative digital twins. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 2764...

  36. [46]

    Robocasa365: A large-scale simulation framework for training and benchmarking generalist robots.arXiv preprint arXiv:2603.04356, 2026

    Soroush Nasiriany, Sepehr Nasiriany, Abhiram Maddukuri, and Yuke Zhu. Robocasa365: A large-scale simulation framework for training and benchmarking generalist robots.arXiv preprint arXiv:2603.04356, 2026

  37. [47]

    Isaac Sim, 2025

    NVIDIA. Isaac Sim, 2025. URLhttps://github.com/isaac-sim/IsaacSim

  38. [48]

    Memoryvla: Perceptual-cognitive memory in vision-language-action models for robotic manipulation

    Hao Shi, Bin Xie, Yingfei Liu, Lin Sun, Fengrong Liu, Tiancai Wang, Erjin Zhou, Haoqiang Fan, Xiangyu Zhang, and Gao Huang. Memoryvla: Perceptual-cognitive memory in vision-language-action models for robotic manipulation. arXiv preprint arXiv:2508.19236, 2025

  39. [49]

    Smolvla: A vision-language-action model for affordable and efficient robotics.arXiv preprint arXiv:2506.01844, 2025

    Mustafa Shukor, Dana Aubakirova, Francesco Capuano, Pepijn Kooijmans, Steven Palma, Adil Zouitine, Michel Aractingi, Caroline Pascal, Martino Russi, Andres Marafioti, et al. Smolvla: A vision-language-action model for affordable and efficient robotics.arXiv preprint arXiv:2506...

  40. [50]

    curobov2: Dynamics-aware motion generation with depth-fused distance fields for high-dof robots, 2026

    Balakumar Sundaralingam, Adithyavairavan Murali, and Stan Birchfield. curobov2: Dynamics-aware motion generation with depth-fused distance fields for high-dof robots, 2026. URLhttps://arxiv.org/abs/2603.05493

  41. [51]

    Motubrain: An advanced world action model for robot control.arXiv preprint arXiv:2604.27792, 2026

    MotuBrain Team, Chendong Xiang, Fan Bao, Haitian Liu, Hengkai Tan, Hongzhe Bi, James Li, Jiabao Liu, Jingrui Pang, Kiro Jing, et al. Motubrain: An advanced world action model for robot control.arXiv preprint arXiv:2604.27792, 2026

  42. [52]

    Rdt2: Enabling zero-shot cross-embodiment generalization by scaling up umi data, September 2025

    RDT Team. Rdt2: Enabling zero-shot cross-embodiment generalization by scaling up umi data, September 2025. URL https://github.com/thu-ml/RDT2

  43. [53]

    Spirit-v1

    Spirit AI Team et al. Spirit-v1. 5: Clean data is the enemy of great robot foundation models. spirit ai blog, 2026

  44. [54]

    Ren, Haohuan Wang, Jiaming Tang, Kyle Stachowicz, Karan Dhabalia, Michael Equi, Quan Vuong, Jost Tobias Springenberg, Sergey Levine, Chelsea Finn, and Danny Driess

    Marcel Torne, Karl Pertsch, Homer Walke, Kyle Vedder, Suraj Nair, Brian Ichter, Allen Z. Ren, Haohuan Wang, Jiaming Tang, Kyle Stachowicz, Karan Dhabalia, Michael Equi, Quan Vuong, Jost Tobias Springenberg, Sergey Levine, Chelsea Finn, and Danny Driess. Mem: Multi-scale embodi...

  45. [55]

    Dexjoco: A benchmark and toolkit for task-oriented dexterous manipulation on mujoco.arXiv preprint arXiv:2605.16257, 2026a

    Hanwen Wang, Weizhi Zhao, Xiangyu Wang, Siyuan Huang, He Lin, Boyuan Zheng, Rongtao Xu, Gang Wang, Yao Mu, He Wang, et al. Dexjoco: A benchmark and toolkit for task-oriented dexterous manipulation on mujoco.arXiv preprint arXiv:2605.16257, 2026a

  46. [56]

    Manitwin: Scaling data-generation-ready digital object dataset to 100k.arXiv preprint arXiv:2603.16866, 2026b

    Kaixuan Wang, Tianxing Chen, Jiawei Liu, Honghao Su, Shaolong Zhu, Minxuan Wang, Zixuan Li, Yue Chen, Huan- ang Gao, Yusen Qin, et al. Manitwin: Scaling data-generation-ready digital object dataset to 100k.arXiv preprint arXiv:2603.16866, 2026b

  47. [57]

    Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation.IEEE Robotics and Automation Letters, 2025

    Junjie Wen, Yichen Zhu, Jinming Li, Minjie Zhu, Zhibin Tang, Kun Wu, Zhiyuan Xu, Ning Liu, Ran Cheng, Chaomin Shen, et al. Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation.IEEE Robotics and Automation Letters, 2025

  48. [58]

    A pragmatic vla foundation model.arXiv preprint arXiv:2601.18692, 2026

    Wei Wu, Fan Lu, Yunnan Wang, Shuai Yang, Shi Liu, Fangjing Wang, Qian Zhu, He Sun, Yong Wang, Shuailei Ma, et al. A pragmatic vla foundation model.arXiv preprint arXiv:2601.18692, 2026. 24 RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Ro...

  49. [59]

    Robochallenge: Large-scale real-robot evaluation of embodied policies.arXiv preprint arXiv:2510.17950, 2025

    Adina Yakefu, Bin Xie, Chongyang Xu, Enwen Zhang, Erjin Zhou, Fan Jia, Haitao Yang, Haoqiang Fan, Haowei Zhang, Hongyang Peng, et al. Robochallenge: Large-scale real-robot evaluation of embodied policies.arXiv preprint arXiv:2510.17950, 2025

  50. [60]

    Eventvla: Event-driven visual evidence memory for long-horizon vision-language-action policies

    Ganlin Yang, Zhangzheng Tu, Yuqiang Yang, Sitong Mao, Junyi Dong, Tianxing Chen, Jiaqi Peng, Jing Xiong, Jiafei Cao, Jifeng Dai, et al. Eventvla: Event-driven visual evidence memory for long-horizon vision-language-action policies. arXiv preprint arXiv:2606.20092, 2026a

  51. [61]

    Robolab: A high-fidelity simulation benchmark for analysis of task generalist policies.arXiv preprint arXiv:2604.09860, 2026b

    Xuning Yang, Rishit Dagli, Alex Zook, Hugo Hadfield, Ankit Goyal, Stan Birchfield, Fabio Ramos, and Jonathan Tremblay. Robolab: A high-fidelity simulation benchmark for analysis of task generalist policies.arXiv preprint arXiv:2604.09860, 2026b

  52. [62]

    Abot-m0: Vla foundation model for robotic manipulation with action manifold learning

    Yandan Yang, Shuang Zeng, Tong Lin, Xinyuan Chang, Dekang Qi, Junjin Xiao, Haoyun Liu, Ronghan Chen, Yuzhi Chen, Dongjie Huo, et al. Abot-m0: Vla foundation model for robotic manipulation with action manifold learning. arXiv preprint arXiv:2602.11236, 2026c

  53. [63]

    Gigaworld-policy: An efficient action-centered world–action model.arXiv preprint arXiv:2603.17240, 2026a

    Angen Ye, Boyuan Wang, Chaojun Ni, Guan Huang, Guosheng Zhao, Hao Li, Hengtao Li, Jie Li, Jindi Lv, Jingyu Liu, et al. Gigaworld-policy: An efficient action-centered world–action model.arXiv preprint arXiv:2603.17240, 2026a

  54. [64]

    Starvla- α: Reducing complexity in vision-language-action systems.arXiv preprint arXiv:2604.11757, 2026b

    Jinhui Ye, Ning Gao, Senqiao Yang, Jinliang Zheng, Zixuan Wang, Yuxin Chen, Pengguang Chen, Yilun Chen, Shu Liu, and Jiaya Jia. Starvla- α: Reducing complexity in vision-language-action systems.arXiv preprint arXiv:2604.11757, 2026b

  55. [65]

    World action models are zero-shot policies, 2026c

    Seonghyeon Ye, Yunhao Ge, Kaiyuan Zheng, Shenyuan Gao, Sihyun Yu, George Kurian, Suneel Indupuru, You Liang Tan, Chuning Zhu, Jiannan Xiang, Ayaan Malik, Kyungmin Lee, William Liang, Nadun Ranawaka, Jiasheng Gu, Yinzhen Xu, Guanzhi Wang, Fengyuan Hu, Avnish Narayan, Johan Bjor...

  56. [66]

    Dm0: An embodied-native vision-language-action model towards physical ai.arXiv preprint arXiv:2602.14974, 2026a

    En Yu, Haoran Lv, Jianjian Sun, Kangheng Lin, Ruitao Zhang, Yukang Shi, Yuyang Chen, Ze Chen, Ziheng Zhang, Fan Jia, et al. Dm0: An embodied-native vision-language-action model towards physical ai.arXiv preprint arXiv:2602.14974, 2026a

  57. [67]

    Wall-oss-0.5 technical report.arXiv preprint arXiv:2605.30877, 2026b

    Ryan Yu, Pushi Zhang, Starrick Liu, Brae Liu, Miracle Kang, Shalfun Li, Lights Shi, Ellie Ma, Ping Yang, Chris Pan, et al. Wall-oss-0.5 technical report.arXiv preprint arXiv:2605.30877, 2026b

  58. [68]

    Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning

    Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine. Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning. InConference on robot learning, pages 1094–1100. PMLR, 2020

  59. [69]

    Qwen-robotmanip technical report: Alignment unlocks scale for robotic manipulation foundation models.arXiv preprint arXiv:2606.17846, 2026a

    Haoqi Yuan, Zhixuan Liang, Anzhe Chen, Ye Wang, Haoyang Li, Pei Lin, Yiyang Huang, Zixing Lei, Tong Zhang, Jiazhao Zhang, et al. Qwen-robotmanip technical report: Alignment unlocks scale for robotic manipulation foundation models.arXiv preprint arXiv:2606.17846, 2026a

  60. [70]

    Fast-wam: Do world action models need test-time future imagination?arXiv preprint arXiv:2603.16666, 2026b

    Tianyuan Yuan, Zibin Dong, Yicheng Liu, and Hang Zhao. Fast-wam: Do world action models need test-time future imagination?arXiv preprint arXiv:2603.16666, 2026b

  61. [71]

    3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations, 2024

    Yanjie Ze, Gu Zhang, Kangning Zhang, Chenyuan Hu, Muhan Wang, and Huazhe Xu. 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations, 2024. URLhttps://arxiv.org/abs/2403.03954

  62. [72]

    Hy-embodied-0.5-vla: From vision-language-action models to a real-world robot learning stack.arXiv preprint arXiv:2606.14409, 2026a

    He Zhang, Lingzhu Xiang, Haitao Lin, Zeyu Huang, Minghui Wang, Dingyan Zhong, Yubo Dong, Yihao Wu, Yongming Rao, Dongsheng Zhang, et al. Hy-embodied-0.5-vla: From vision-language-action models to a real-world robot learning stack.arXiv preprint arXiv:2606.14409, 2026a

  63. [73]

    A1: A fully transparent open-source, adaptive and efficient truncated vision-language-action model.arXiv preprint arXiv:2604.05672, 2026b

    Kaidong Zhang, Jian Zhang, Rongtao Xu, Yu Sun, Shuoshuo Xue, Youpeng Wen, Xiaoyu Guo, Minghao Guo, Weijia Liufu, Liu Zihou, et al. A1: A fully transparent open-source, adaptive and efficient truncated vision-language-action model.arXiv preprint arXiv:2604.05672, 2026b. 25 Robo...

  64. [74]

    Vlabench: A large-scale benchmark for language-conditioned robotics manipulation with long-horizon reasoning tasks

    Shiduo Zhang, Zhe Xu, Peiju Liu, Xiaopeng Yu, Yuan Li, Qinghui Gao, Zhaoye Fei, Zhangyue Yin, Zuxuan Wu, Yu- Gang Jiang, et al. Vlabench: A large-scale benchmark for language-conditioned robotics manipulation with long-horizon reasoning tasks. InProceedings of the IEEE/CVF Int...

  65. [76]

    Joyai-ra 0.1: A foundation model for robotic autonomy.arXiv preprint arXiv:2604.20100, 2026d

    Tianle Zhang, Zhihao Yuan, Dafeng Chi, Peidong Liu, Dongwei Li, Kejun Hu, Likui Zhang, Junnan Nie, Ziming Wei, Zengjue Chen, et al. Joyai-ra 0.1: A foundation model for robotic autonomy.arXiv preprint arXiv:2604.20100, 2026d

  66. [77]

    Dexora: Open-source vla for high-dof bimanual dexterity.arXiv preprint arXiv:2605.18722, 2026e

    Zongzheng Zhang, Jingrui Pang, Zhuo Yang, Kun Li, Minwen Liao, Saining Zhang, Guoxuan Chi, Jinbang Guo, Huan- ang Gao, Modi Shi, et al. Dexora: Open-source vla for high-dof bimanual dexterity.arXiv preprint arXiv:2605.18722, 2026e

  67. [78]

    Learning fine-grained bimanual manipulation with low-cost hardware.arXiv preprint arXiv:2304.13705, 2023

    Tony Z Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn. Learning fine-grained bimanual manipulation with low-cost hardware.arXiv preprint arXiv:2304.13705, 2023

  68. [79]

    X-vla: Soft-prompted transformer as scalable cross-embodiment vision-language-action model

    Jinliang Zheng, Jianxiong Li, Zhihao Wang, Dongxiu Liu, Xirui Kang, Yuchun Feng, Yinan Zheng, Jiayin Zou, Yilun Chen, Jia Zeng, et al. X-vla: Soft-prompted transformer as scalable cross-embodiment vision-language-action model. arXiv preprint arXiv:2510.10274, 2025

  69. [80]

    Clothesnet: An information-rich 3d garment model repository with simulated clothes environment

    Bingyang Zhou, Haoyu Zhou, Tianhai Liang, Qiaojun Yu, Siheng Zhao, Yuwei Zeng, Jun Lv, Siyuan Luo, Qiancai Wang, Xinyuan Yu, et al. Clothesnet: An information-rich 3d garment model repository with simulated clothes environment. In Proceedings of the IEEE/CVF International Conf...

  70. [81]

    Skill” is the number of operation primitives covered; “Num. Policies

    Yuke Zhu, Josiah Wong, Ajay Mandlekar, Roberto Martín-Martín, Abhishek Joshi, Kevin Lin, Abhiram Maddukuri, Soroush Nasiriany, and Yifeng Zhu. robosuite: A modular simulation framework and benchmark for robot learning. arXiv preprint arXiv:2009.12293, 2020. 26 RoboDojo: A Unif...

Pith tools

Reviewed July 11, 2026 · model on record in the stance chip above.