Pith. sign in

REVIEW 3 major objections 6 minor 286 references

Embodied AI has no internet-scale shortcut: the field’s data sources form a five-layer pyramid trading robot alignment against scalability.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 06:14 UTC pith:Q5GVA6WS

load-bearing objection Solid community map of embodied data sources; the pyramid holds up, but the recipe-to-capability story in §7 is co-occurrence, not evidence. the 3 major comments →

arxiv 2607.24744 v1 pith:Q5GVA6WS submitted 2026-07-27 cs.RO cs.CV

Data Pyramid for Embodied Manipulation

classification cs.RO cs.CV
keywords embodied data pyramidrobot learningvision-language-actionworld-action modelsegocentric datasimulationdata recipescross-embodiment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Internet-scale vision and language models can train on the open web; robots cannot, because they need observations tied to physical states and actions. This paper organizes that scarcer ecosystem as a five-layer “data pyramid”: real-robot trajectories at the apex, then UMI-style handheld-gripper demos, egocentric and exocentric human video, simulation, and finally general vision-language corpora at the base. The ordering is driven by a tension between how directly a source supports physical robot control and how cheaply it can be scaled, with quality, diversity, reusability, and physical fidelity as further dimensions. The authors then read recent embodied foundation models—embodied brains, vision-language-action models, and world-action models—through their data recipes, linking which layers are mixed in pretraining to capabilities in perception, reasoning, planning, action generation, and world prediction. They close with six open problems, from tactile and failure data to cross-embodiment action alignment and principled mixture design. A sympathetic reader cares because the taxonomy turns a scattered data landscape into a shared map for building the next generation of embodied systems.

Core claim

The embodied data ecosystem is usefully organized as a five-layer pyramid—real-robot, UMI-style, egocentric/exocentric, simulation, and general vision-language data—whose primary axes are scalability versus robot alignment, and whose secondary dimensions are quality, diversity, reusability, and physical fidelity; the pretraining recipes of embodied brain, VLA, and world-action models can be related, layer by layer, to capabilities in perception, reasoning, planning, action generation, and world prediction.

What carries the argument

The Data Pyramid: a five-layer taxonomy that ranks embodied data sources from strongest robot alignment (real-robot trajectories) to greatest scalability (general vision-language data), and uses that structure to compare collection methods, characterize trade-offs, and audit how foundation-model recipes select, align, and mix layers.

Load-bearing premise

That reported training mixtures in papers are enough to credit specific robot skills to specific data layers, even without controlled tests that hold architecture and scale fixed while changing only the mix.

What would settle it

Train matched models that differ only in pyramid-layer mixture weights (for example robot-only versus robot-plus-ego versus full five-layer mixes) under fixed architecture and compute, then check whether the paper’s claimed layer-to-capability links hold on held-out perception, planning, action, and world-prediction benchmarks—or whether robot-only recipes match the mixed ones.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Dataset builders and model teams can place new corpora on shared axes (scalability, robot alignment, fidelity) instead of treating every source as incomparable.
  • Pretraining recipes should be designed as explicit mixes across pyramid layers, not as ad hoc piles of whatever trajectories are available.
  • Scarce but high-value signals—tactile contact, failure and recovery, cross-embodiment action labels—become first-class collection targets rather than afterthoughts.
  • Egocentric human video is treated as a primary pretraining substrate between web data and robot trajectories, especially for dexterous hands once retargeting is solved.
  • World-action and VLA systems can be audited by which pyramid layers supply action-free priors versus executable action grounding.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the pyramid’s alignment–scale tension is right, pure web-scale pretraining alone will keep underdelivering on contact-rich control until higher layers grow or transfer methods harden.
  • The open “optimal recipe” problem implies the field may need mixture ablations the way language modeling needed data-mix studies—not just larger catalogs of datasets.
  • Standardizing action frames and failure labels across embodiments could matter as much as adding hours, because misaligned supervision can cancel gains from scale.
  • Tactile and recovery data, if collected at pyramid scale, would likely shift evaluation from success-rate demos toward failure awareness and contact competence.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This is a data-centric survey of the embodied-manipulation data ecosystem. The authors organize data sources into a five-layer "pyramid" — real-robot, UMI-style, egocentric/exocentric, simulation, and general vision-language data — ordered along two primary axes (scalability and robot alignment) and characterized along four secondary dimensions (quality, diversity, reusability, physical fidelity). §§2–6 review datasets, collection pipelines, embodiments, and sensing for each layer with extensive summary tables. §7 examines how embodied brain models, VLAs, and world-action models draw on the pyramid layers in their pretraining recipes, and §8 lists six open challenges. The paper also maintains an open-source curation repository.

Significance. If the taxonomy holds up, the paper provides a genuinely useful organizing device for a fragmented literature. The strengths are concrete: the per-layer tables (Tables 1–6) are unusually comprehensive and consistently annotated (embodiment counts, calibration flags, tactile/dexterous/mobile markers), the six-dimension characterization is applied uniformly across layers, and the manuscript repeatedly hedges its ordering as "an overall synthesis of the six dimensions rather than a strictly monotonic progression" (§1), which preempts the most obvious internal objection. §7's model-recipe table (Table 7) and the cross-embodiment action-representation taxonomy (§7.2.2: projection vs. zero-padding vs. semantic slots; robot-/camera-/wrist-centric frames) are original syntheses not available in prior surveys, and the open repository adds community value. The work is explicitly positioned against model-recipe-specific pyramid views (Motus, GR00T) in §1, which is appropriately disclosed.

major comments (3)
  1. [§7, abstract] §7.1–7.5 (and the abstract): the claim to "relate data composition to capabilities in perception, reasoning, planning, action generation, and world prediction" is supported by reading off co-occurrence in published pretraining recipes (Table 7), not by any controlled mixture evidence. The inference is confounded by architecture, backbone pretraining, scale, undisclosed filtering, and post-training, all of which covary across the surveyed models. The manuscript itself concedes this twice (§7.2.1: "the optimal data recipe... remains an open question"; §7 takeaway: "the contribution of each data source has not been systematically isolated"). The fix is local but necessary: (i) temper the abstract phrasing to match the hedged evidentiary status (e.g., "survey how recipes are composed and discuss hypothesized capability links"), and (ii) mark each §7.3–7.5 capability attribution explicitly as
  2. [§1, "Data Pyramid at a Glance"] The apex-to-base ordering is asserted but never scored. Since the paper concedes the order is not monotonic in any single property, the reader cannot verify the synthesis: e.g., simulation provides executable actions and privileged labels (high robot alignment on its face) yet sits below egocentric data, which provides no actions at all. A compact table scoring each layer on all six dimensions (with brief justification) would make the stipulated ordering auditable and would also expose the genuine disagreements a reader might have (e.g., UMI vs. egocentric on physical fidelity). Without it, the pyramid's central organizational claim rests on prose alone.
  3. [Figure 4, §2.4, §5.7] Figure 4 uses a heuristic keyframe-extraction procedure (following PerAct) as a proxy for trajectory-level diversity, and §2.4/§5.7/§4.4 draw comparative conclusions from it (e.g., InternData-A1's keypoints "concentrated within a small region" evidencing repetition; EgoVerse's broad spread evidencing coverage). These datasets differ in task, embodiment, camera, and workspace, so spatial spread of keyframes confounds workspace size with behavioral diversity, and single-example comparisons risk being illustrative rather than evidentiary. The figure caption should state this limitation explicitly, or the claims should be weakened to "illustrative."
minor comments (6)
  1. [Figure 2] Figure 2 plots scale evolution across layers in mutually incomparable units (demonstrations, hours, QA pairs) on what appears to be a common visual frame; the inset curves help, but the caption should warn against cross-layer quantitative comparison.
  2. [Table 7, §7.2.1] Table 7 encodes each model's data sources with layer icons but no mixture proportions; since §7.2.1 reports hour counts for a few models (Qwen-RobotManip, Xiaomi-Robotics-1), a coarse quantitative column (where disclosed) would substantially strengthen the recipe analysis.
  3. [§7.2.2] §7.2.2 states that "existing studies provide limited controlled ablations" of geometric action representations; the survey would benefit from citing the ablations that do exist (e.g., cross-embodiment transfer analyses in the Open X-Embodiment and DROID lines of work, and any mixture-weight ablations in cited models such as π0.5/GR00T) so readers can locate the nearest available evidence.
  4. [§5.5 vs §7.5] §5.5 groups world-model-based data engines under simulation, while §7.5 treats WAMs as consumers of pyramid layers; a sentence reconciling this dual role (simulator as data source vs. model trained on data) would prevent reader confusion about where "world models as simulators" sit in the taxonomy.
  5. [Tables 1–5] Several tables (e.g., Table 1 "Arm" column, Table 4 "Embod.") use S/D/H abbreviations defined in captions; consider repeating the legend in each table footer for standalone readability.
  6. [§6.7] §6.7 repeats the citation "[228]" twice in one sentence ("Failure-oriented data [228] such as RoboFail [228]").

Circularity Check

0 steps flagged

No circular derivation: the pyramid is an explicit taxonomy and §7 recipe–capability links are interpretive survey claims, not predictions forced by fitted inputs or self-citation chains.

full rationale

This manuscript is a data-centric survey and organizational taxonomy, not a first-principles derivation that claims to predict observables from independent premises. The five-layer pyramid is introduced as a stipulated synthesis ordered by scalability versus robot alignment (plus quality, diversity, reusability, physical fidelity), and the text itself states that the ordering is an overall synthesis rather than a strictly monotonic law along every axis. Category membership is therefore definitional framework design, not a claimed reduction of Y from X where X is secretly defined as Y. Section 7 relates published pretraining recipes (Table 7) to model-family capabilities by co-occurrence and qualitative complementarity; those inferences may be confounded by architecture and scale, but confounding is an evidence-strength issue, not circularity: nothing is fitted to a subset and then re-labeled a prediction, and no uniqueness theorem or load-bearing ansatz is imported from overlapping-author prior work to forbid alternatives. Prior pyramid-like views (e.g., Motus, GR00T) are cited as motivation and then criticized for limited category-level analysis, not used as the sole warrant that forces the present taxonomy. The paper repeatedly hedges that optimal data recipes remain open and that robot-only models can still perform strongly. There is no equation chain, no fitted-parameter-as-prediction step, and no self-citation uniqueness loop. Score 0 with no circular steps.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 1 invented entities

Load-bearing structure is definitional and domain-level: the field is assumed to be usefully partitioned into five data source types; scalability and robot alignment are assumed to be the right primary tension; secondary axes (quality, diversity, reusability, physical fidelity) are assumed comparable across heterogeneous literatures; and public training-recipe descriptions are assumed informative enough for capability discussion. No fitted constants. The main invented construct is the ordered Data Pyramid itself.

axioms (4)
  • domain assumption Embodied foundation-model pretraining is usefully analyzed primarily through heterogeneous data source composition rather than architecture alone.
    Stated in the introduction’s two framing questions and §7 overview; underpins the entire data-centric analysis.
  • ad hoc to paper Scalability and robot alignment are the dominant organizing tension and justify a single apex-to-base order of five layers.
    §1 explicitly chooses this ordering as an “overall synthesis” not strictly monotonic on every property; alternative primary axes (e.g., action-label density alone) could reorder layers.
  • domain assumption Quality, diversity, reusability, and physical fidelity are meaningful, comparable category-level dimensions across real, human, sim, and web data.
    Introduced in §1 and applied in each layer’s advantages/limitations; dimensions are qualitative, not operationalized with shared metrics.
  • domain assumption Published descriptions of model pretraining mixes are adequate evidence for relating data layers to perception, reasoning, planning, action, and world-prediction capabilities.
    §7 and Table 7; paper later admits optimal recipes and causal isolation remain open (§7.2.1, §8.6).
invented entities (1)
  • Embodied Data Pyramid (five-layer taxonomy ordered by scalability vs. robot alignment) no independent evidence
    purpose: Provide a category-level map of embodied data sources and a lens for reading foundation-model data recipes and open collection problems.
    Named organizing contribution of the paper; builds on but generalizes prior model-specific pyramid-like views the authors cite.

pith-pipeline@v1.2.0-grok45-kimik3 · 55100 in / 3272 out tokens · 79414 ms · 2026-07-31T06:14:59.964745+00:00 · methodology

0 comments
read the original abstract

Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a "pyramid" spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data. We organize the pyramid around the tension between scalability and robot alignment, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity. We then analyze recent embodied foundation models through the lens of their data recipes, examining how different sources are selected, aligned, and mixed during pretraining. For embodied brain models, vision-language-action models, and world-action models alike, we relate data composition to capabilities in perception, reasoning, planning, action generation, and world prediction. We close by discussing six open challenges: building large-scale tactile datasets, collecting failure and recovery data, developing scalable data-collection pipelines, aligning actions across embodiments, leveraging egocentric data for dexterous manipulation, and designing principled data recipes for robot learning. We hope this work paves the foundation for the design of next-generation embodied systems.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

286 extracted references · 3 canonical work pages

  1. [1]

    https://huggingface.co/datasets/builddotai/ Egocentric-100K, 2025

    Builddotai/egocentric-100k·datasets at hugging face. https://huggingface.co/datasets/builddotai/ Egocentric-100K, 2025. URLhttps://huggingface.co/datasets/builddotai/egocentric-100k

  2. [2]

    TacSL: A library for visuotactile sensor simulation and learning.IEEE Trans

    Iretiayo Akinola, Jie Xu, Jan Carius, Dieter Fox, and Yashraj Narang. TacSL: A library for visuotactile sensor simulation and learning.IEEE Trans. Robot., 2025. URLhttps://arxiv.org/abs/2408.06506

  3. [3]

    World simulation with video foundation models for physical AI.arXiv preprint arXiv:2511.00062, 2025

    Arslan Ali, Junjie Bai, Maciej Bala, Yogesh Balaji, Aaron Blakeman, Tiffany Cai, Jiaxin Cao, Tianshi Cao, Elizabeth Cha, Yu-Wei Chao, et al. World simulation with video foundation models for physical AI.arXiv preprint arXiv:2511.00062, 2025. URLhttps://arxiv.org/abs/2511.00062

  4. [4]

    HapTile: a haptic-informed Vision-Tactile-Language-Action dataset for contact-rich imitation learning.arXiv preprint arXiv:2606.04825, 2026

    Amirhosein Alian, Yongqiang Zhao, Shiyi Gu, Xuyang Zhang, Zhuo Chen, Christopher E Mower, Haitham Bou-Ammar, and Shan Luo. HapTile: a haptic-informed Vision-Tactile-Language-Action dataset for contact-rich imitation learning.arXiv preprint arXiv:2606.04825, 2026. URLhttps://arxiv.org/abs/2606.04825

  5. [5]

    Scalable behavior cloning with open data, training, and evaluation, 2026

    Arthur Allshire, Himanshu Gaurav Singh, Ritvik Singh, Adam Rashid, Hongsuk Choi, David McAllister, Justin Yu, Yiyuan Chen, Huang Huang, Pieter Abbeel, Xi Chen, Rocky Duan, Phillip Isola, Jitendra Malik, Fred Shentu, Guanya Shi, Philipp Wu, and Angjoo Kanazawa. Scalable behavior cloning with open data, training, and evaluation, 2026. URLhttps://arxiv.org/a...

  6. [6]

    Lawrence Zitnick, and Devi Parikh

    Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. VQA: Visual question answering. InIEEE/CVF Int. Conf. Comput. Vis., 2015. URL https: //arxiv.org/abs/1505.00468

  7. [8]

    Muckley, Ammar Rizvi, et al

    Mahmoud Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Mojtaba Komeili, Matthew J. Muckley, Ammar Rizvi, et al. V-JEPA 2: self-supervised video models enable understanding, prediction and planning.arXiv preprint arXiv:2506.09985, 2025. URLhttps://arxiv.org/abs/2506.09985

  8. [9]

    ScanQA: 3D question answering for spatial scene understanding

    Daichi Azuma, Taiki Miyanishi, Shuhei Kurita, and Motoaki Kawanabe. ScanQA: 3D question answering for spatial scene understanding. InIEEE/CVF Conf. Comput. Vis. Pattern Recog., 2022. URLhttps://arxiv.org/ abs/2112.10482

  9. [10]

    Crandall, and Chen Yu

    Sven Bambach, Stefan Lee, David J. Crandall, and Chen Yu. Lending a hand: detecting hands and recognizing activities in complex egocentric interactions. InIEEE/CVF Int. Conf. Comput. Vis., 2015. URL https: //doi.org/10.1109/iccv.2015.226

  10. [11]

    Introducing hot3d: an egocentric dataset for 3D hand and object tracking.arXiv preprint arXiv:2406.09598, 2024

    Prithviraj Banerjee, Sindi Shkodrani, Pierre Moulon, Shreyas Hampali, Fan Zhang, Jade Fountain, Edward Miller, Selen Basol, Richard Newcombe, Robert Wang, Jakob Julian Engel, and Tomas Hodan. Introducing hot3d: an egocentric dataset for 3D hand and object tracking.arXiv preprint arXiv:2406.09598, 2024. URL https://arxiv.org/abs/2406.09598. 46

  11. [12]

    DexArt: Benchmarking generalizable dexterous manipulation with articulated objects

    Chen Bao, Helin Xu, Yuzhe Qin, and Xiaolong Wang. DexArt: Benchmarking generalizable dexterous manipulation with articulated objects. InIEEE/CVF Conf. Comput. Vis. Pattern Recog., 2023. URL https://arxiv.org/abs/2305.05706

  12. [13]

    Arkitscenes: a diverse real-world dataset for 3D indoor scene understanding using mobile RGB-D data

    Gilad Baruch, Zhuoyuan Chen, Afshin Dehghan, Tal Dimry, Yuri Feigin, Peter Fu, Thomas Gebauer, Brandon Joffe, Daniel Kurz, Arik Schwartz, and Elad Shulman. Arkitscenes: a diverse real-world dataset for 3D indoor scene understanding using mobile RGB-D data. InAdv. Neural Inf. Process. Syst., 2021. URLhttps: //arxiv.org/abs/2111.08897

  13. [14]

    Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, et al

    Charles Beattie, Joel Z. Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, et al. DeepMind lab.arXiv preprint arXiv:1612.03801, 2016. URLhttps://arxiv.org/ abs/1612.03801

  14. [15]

    Track2Act: Predicting point tracks from internet videos enables generalizable robot manipulation

    Homanga Bharadhwaj, Roozbeh Mottaghi, Abhinav Gupta, and Shubham Tulsiani. Track2Act: Predicting point tracks from internet videos enables generalizable robot manipulation. InEur. Conf. Comput. Vis., 2024. URL https://arxiv.org/abs/2405.01527

  15. [16]

    RoboAgent: generalization and efficiency in robot manipulation via semantic augmentations and action chunking

    Homanga Bharadhwaj, Jay Vakil, Mohit Sharma, Abhinav Gupta, Shubham Tulsiani, and Vikash Kumar. RoboAgent: generalization and efficiency in robot manipulation via semantic augmentations and action chunking. InIEEE Int. Conf. Robot. Autom., 2024. URLhttps://arxiv.org/abs/2309.01918

  16. [17]

    MapleGrasp: mask-guided feature pooling for language-driven efficient robotic grasping

    Vineet Bhat, Naman Patel, Prashanth Krishnamurthy, Ramesh Karri, and Farshad Khorrami. MapleGrasp: mask-guided feature pooling for language-driven efficient robotic grasping. InIEEE/CVF Winter Conf. Appl. Comput. Vis., 2026. URLhttps://arxiv.org/abs/2506.06535

  17. [18]

    Motus: a unified latent action world model

    Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, Shuhe Huang, Haitian Liu, Ruowen Zhao, Yao Feng, et al. Motus: a unified latent action world model. InIEEE/CVF Conf. Comput. Vis. Pattern Recog., 2026. URL https://arxiv.org/abs/2512.13030

  18. [19]

    H-RDT: human manipulation enhanced bimanual robotic manipulation

    Hongzhe Bi, Lingxuan Wu, Tianwei Lin, Hengkai Tan, Zhizhong Su, Hang Su, and Jun Zhu. H-RDT: human manipulation enhanced bimanual robotic manipulation. InAAAI Conf. Artif. Intell., 2026. URL https: //doi.org/10.1609/aaai.v40i22.38875

  19. [20]

    GR00T n1: an open foundation model for generalist humanoid robots

    Johan Bjorck, Fernando Castañeda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, et al. GR00T n1: an open foundation model for generalist humanoid robots. arXiv preprint arXiv:2503.14734, 2025. URLhttps://arxiv.org/abs/2503.14734

  20. [21]

    π0: a vision-language-action flow model for general robot control

    Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, et al. π0: a vision-language-action flow model for general robot control. InRobot. Sci. Syst., 2025. URL https://arxiv.org/abs/2410.24164

  21. [22]

    WEAR: An outdoor sports dataset for wearable and egocentric activity recognition.Proc

    Marius Bock, Hilde Kuehne, Kristof Van Laerhoven, and Michael Moeller. WEAR: An outdoor sports dataset for wearable and egocentric activity recognition.Proc. ACM Interact. Mob. Wearable Ubiquitous Technol., 2024. URLhttps://arxiv.org/abs/2304.05088

  22. [23]

    RT-2: vision-language-action models transfer web knowledge to robotic control

    Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, et al. RT-2: vision-language-action models transfer web knowledge to robotic control. In Conf. Robot Learn., 2023. URLhttps://arxiv.org/abs/2307.15818

  23. [24]

    RT-1: robotics transformer for real-world control at scale

    Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, et al. RT-1: robotics transformer for real-world control at scale. InRobot. Sci. Syst., 2023. URLhttps://arxiv.org/abs/2212.06817

  24. [25]

    Genie: generative interactive environments

    Jake Bruce, Michael Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, et al. Genie: generative interactive environments. InInt. Conf. Mach. Learn., 2024. URLhttps://arxiv.org/abs/2402.15391

  25. [26]

    AgiBot world colosseo: a large-scale manipulation platform for scalable and intelligent embodied systems

    Qingwen Bu, Jisong Cai, Li Chen, Xiuqi Cui, Yan Ding, Siyuan Feng, Shenyuan Gao, Xindong He, Xuan Hu, Xu Huang, et al. AgiBot world colosseo: a large-scale manipulation platform for scalable and intelligent embodied systems. InIEEE/RSJ Int. Conf. Intell. Robots Syst., 2025. URLhttps://arxiv.org/abs/2503.06669

  26. [27]

    UniVLA: learning to act anywhere with task-centric latent actions

    Qingwen Bu, Yanting Yang, Jisong Cai, Shenyuan Gao, Guanghui Ren, Maoqing Yao, Ping Luo, and Hongyang Li. UniVLA: learning to act anywhere with task-centric latent actions. InRobot. Sci. Syst., 2025. URL https://doi.org/10.15607/rss.2025.xxi.014. 47

  27. [28]

    Internvla-a1: unifying understanding, generation and action for robotic manipulation.arXiv preprint arXiv:2601.02456, 2026

    Junhao Cai, Zetao Cai, Jiafei Cao, Yilun Chen, Zeyu He, Lei Jiang, Hang Li, Hengjie Li, Yang Li, Yufei Liu, et al. Internvla-a1: unifying understanding, generation and action for robotic manipulation.arXiv preprint arXiv:2601.02456, 2026. URLhttps://arxiv.org/abs/2601.02456

  28. [29]

    Xiaomi-robotics-0: An open-sourced vision-language-action model with real-time execution.arXiv preprint arXiv:2602.12684, 2026

    Rui Cai, Jun Guo, Xinze He, Piaopiao Jin, Jie Li, Bingxuan Lin, Futeng Liu, Wei Liu, et al. Xiaomi-robotics-0: An open-sourced vision-language-action model with real-time execution.arXiv preprint arXiv:2602.12684, 2026. URLhttps://arxiv.org/abs/2602.12684

  29. [31]

    Scaling spatial intelligence with multimodal foundation models

    Zhongang Cai, Ruisi Wang, Chenyang Gu, Fanyi Pu, Junxiang Xu, Yubo Wang, Wanqi Yin, Zhitao Yang, et al. Scaling spatial intelligence with multimodal foundation models. InIEEE/CVF Conf. Comput. Vis. Pattern Recog., 2026. URLhttps://arxiv.org/abs/2511.13719

  30. [32]

    SuctionNet-1Billion: a large-scale benchmark for suction grasping.IEEE Robot

    Hanwen Cao, Hao-Shu Fang, Wenhai Liu, and Cewu Lu. SuctionNet-1Billion: a large-scale benchmark for suction grasping.IEEE Robot. Autom. Lett., 2021. URLhttps://arxiv.org/abs/2103.12311

  31. [33]

    Physx-3D: Physical-grounded 3D asset generation

    Ziang Cao, Zhaoxi Chen, Liang Pan, and Ziwei Liu. Physx-3D: Physical-grounded 3D asset generation. InAdv. Neural Inf. Process. Syst., 2025. URLhttps://arxiv.org/abs/2507.12465

  32. [34]

    Physx-anything: Simulation-ready physical 3D assets from single image

    Ziang Cao, Fangzhou Hong, Zhaoxi Chen, Liang Pan, and Ziwei Liu. Physx-anything: Simulation-ready physical 3D assets from single image. InIEEE/CVF Conf. Comput. Vis. Pattern Recog., 2026. URLhttps: //arxiv.org/abs/2511.13648

  33. [35]

    PhysX-Omni: Unified simulation-ready physical 3D generation for rigid, deformable, and articulated objects

    Ziang Cao, Yinghao Liu, Haitian Li, Runmao Yao, Fangzhou Hong, Zhaoxi Chen, Liang Pan, and Ziwei Liu. PhysX-Omni: Unified simulation-ready physical 3D generation for rigid, deformable, and articulated objects. arXiv preprint arXiv:2605.21572, 2026. URLhttps://arxiv.org/abs/2605.21572

  34. [36]

    WorldVLA: towards autoregressive action world model.arXiv preprint arXiv:2506.21539, 2025

    Jun Cen, Chaohui Yu, Hangjie Yuan, Yuming Jiang, Siteng Huang, Jiayan Guo, Xin Li, Yibing Song, Hao Luo, Fan Wang, Deli Zhao, and Hao Chen. WorldVLA: towards autoregressive action world model.arXiv preprint arXiv:2506.21539, 2025. URLhttps://arxiv.org/abs/2506.21539

  35. [37]

    Chandra, Iman Nematollahi, Chenguang Huang, Tim Welschehold, Wolfram Burgard, and Abhinav Valada

    Akshay L. Chandra, Iman Nematollahi, Chenguang Huang, Tim Welschehold, Wolfram Burgard, and Abhinav Valada. DiWA: diffusion policy adaptation with world models. InConf. Robot Learn., 2025. URL https: //arxiv.org/abs/2508.03645

  36. [38]

    Matterport3D: learning from RGB-D data in indoor environments

    Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Nießner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3D: learning from RGB-D data in indoor environments. InInt. Conf. 3D Vis., 2017. URLhttps://doi.org/10.1109/3dv.2017.00081

  37. [40]

    IndEgo: A dataset of industrial scenarios and collaborative work for egocentric assistants

    Vivek Chavan, Yasmina Imgrund, Tung Dao, Sanwantri Bai, Bosong Wang, Ze Lu, Oliver Heimann, and Jörg Krüger. IndEgo: A dataset of industrial scenarios and collaborative work for egocentric assistants. InAdv. Neural Inf. Process. Syst., 2025. URLhttps://arxiv.org/abs/2511.19684

  38. [41]

    GR-2: a generative video-language-action model with web-scale knowledge for robot manipulation.arXiv preprint arXiv:2410.06158, 2024

    Chi-Lam Cheang, Guangzeng Chen, Ya Jing, Tao Kong, Hang Li, Yifeng Li, Yuxiao Liu, Hongtao Wu, Jiafeng Xu, Yichu Yang, Hanbo Zhang, and Minzhao Zhu. GR-2: a generative video-language-action model with web-scale knowledge for robot manipulation.arXiv preprint arXiv:2410.06158, 2024. URLhttps://arxiv.org/abs/2410. 06158

  39. [42]

    Gr-3 technical report.arXiv preprint arXiv:2507.15493, 2025

    Chilam Cheang, Sijin Chen, Zhongren Cui, Yingdong Hu, Liqun Huang, Tao Kong, Hang Li, Yifeng Li, et al. Gr-3 technical report.arXiv preprint arXiv:2507.15493, 2025. URLhttps://arxiv.org/abs/2507.15493

  40. [43]

    Univtac: A unified simulation platform for visuo-tactile manipulation data generation, learning, and benchmarking.arXiv preprint arXiv:2602.10093, 2026

    Baijun Chen, Weijie Wan, Tianxing Chen, Xianda Guo, Congsheng Xu, Yuanyang Qi, Haojie Zhang, Longyan Wu, Tianling Xu, Zixuan Li, et al. Univtac: A unified simulation platform for visuo-tactile manipulation data generation, learning, and benchmarking.arXiv preprint arXiv:2602.10093, 2026

  41. [44]

    Visa-flow: accelerating robot skill learning via large-scale video semantic action flow

    Changhe Chen, Quantao Yang, Xiaohao Xu, Nima Fazeli, and Olov Andersson. Visa-flow: accelerating robot skill learning via large-scale video semantic action flow. InIEEE Int. Conf. Robot. Autom., 2026. URL https://arxiv.org/abs/2505.01288. 48

  42. [46]

    ShareGPT4V: improving large multi-modal models with better captions

    LinChen, JinsongLi, XiaoyiDong, PanZhang, ConghuiHe, JiaqiWang, FengZhao, andDahuaLin. ShareGPT4V: improving large multi-modal models with better captions. InEur. Conf. Comput. Vis., 2024. URL https: //arxiv.org/abs/2311.12793

  43. [48]

    Daxbench: benchmarking deformable object manipulation with differentiable physics

    Siwei Chen, Yiqing Xu, Cunjun Yu, Linfeng Li, Xiao Ma, Zhongwen Xu, and David Hsu. Daxbench: benchmarking deformable object manipulation with differentiable physics. InInt. Conf. Learn. Represent., 2023. URL https://arxiv.org/abs/2210.13066

  44. [49]

    Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation.arXiv preprint arXiv:2506.18088, 2025

    Tianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai, Yibin Liu, Zixuan Li, Qiwei Liang, Xianliang Lin, Yiheng Ge, Zhenyu Gu, et al. Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation.arXiv preprint arXiv:2506.18088, 2025. URLhttps: //arxiv.org/abs/2506.18088

  45. [50]

    Robodojo: A unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies.arXiv preprint arXiv:2607.04434, 2026

    Tianxing Chen, Yue Chen, Zixuan Li, Junyuan Tang, Kailun Su, Weijie Wan, Baijun Chen, Haoran Lu, Haowen Yan, Honghao Su, et al. Robodojo: A unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies.arXiv preprint arXiv:2607.04434, 2026

  46. [51]

    URLhttps://arxiv.org/abs/2607.00678

  47. [52]

    Villa-x: enhancing latent action modeling in vision-language-action models

    Xiaoyu Chen, Hangxing Wei, Pushi Zhang, Chuheng Zhang, Kaixin Wang, Yanjiang Guo, Rushuai Yang, Yucen Wang, Xinquan Xiao, Li Zhao, et al. Villa-x: enhancing latent action modeling in vision-language-action models. InInt. Conf. Learn. Represent., 2026. URLhttps://arxiv.org/abs/2507.23682

  48. [54]

    Moto: Latent motion token as the bridging language for learning robot manipulation from videos

    Yi Chen, Yuying Ge, Weiliang Tang, Yizhuo Li, Yixiao Ge, Mingyu Ding, Ying Shan, and Xihui Liu. Moto: Latent motion token as the bridging language for learning robot manipulation from videos. InIEEE/CVF Int. Conf. Comput. Vis., 2025. URLhttps://arxiv.org/abs/2412.04445

  49. [55]

    Rmbench: Memory-dependent robotic manipulation benchmark with insights into policy design.arXiv preprint arXiv:2603.01229, 2026

    Tianxing Chen, Yuran Wang, Mingleyang Li, Yan Qin, Hao Shi, Zixuan Li, Yifan Hu, Yingsheng Zhang, Kaixuan Wang, Yue Chen, et al. Rmbench: Memory-dependent robotic manipulation benchmark with insights into policy design.arXiv preprint arXiv:2603.01229, 2026

  50. [56]

    Towards human-level bimanual dexterous manipulation with reinforcement learning

    Yuanpei Chen, Tianhao Wu, Shengjie Wang, Xidong Feng, Jiechuang Jiang, Stephen Marcus McAleer, Yiran Geng, Hao Dong, Zongqing Lu, Song-Chun Zhu, and Yaodong Yang. Towards human-level bimanual dexterous manipulation with reinforcement learning. InAdv. Neural Inf. Process. Syst., 2022. URLhttps://arxiv.org/abs/ 2206.08686

  51. [57]

    RH20T-P: a primitive-level robotic manipulation dataset towards composable generalization agents in real-world scenarios

    Zeren Chen, Zhelun Shi, Xiaoya Lu, Lehan He, Sucheng Qian, Enshen Zhou, Zhenfei Yin, Wanli Ouyang, Jing Shao, Yu Qiao, Cewu Lu, and Lu Sheng. RH20T-P: a primitive-level robotic manipulation dataset towards composable generalization agents in real-world scenarios. InIEEE/RSJ Int. Conf. Intell. Robots Syst., 2025. URLhttps://arxiv.org/abs/2403.19622

  52. [58]

    URLhttps://arxiv.org/abs/2510.13778

  53. [59]

    Open-television: teleoperation with immersive active visual feedback

    Xuxin Cheng, Jialong Li, Shiqi Yang, Ge Yang, and Xiaolong Wang. Open-television: teleoperation with immersive active visual feedback. InConf. Robot Learn., 2024. URLhttps://arxiv.org/abs/2407.01512

  54. [60]

    EgoPlan-Bench: benchmarking multimodal large language models for human-level planning.Int

    Yi Chen, Yuying Ge, Yixiao Ge, Mingyu Ding, Bohao Li, Rui Wang, Ruifeng Xu, Ying Shan, and Xihui Liu. EgoPlan-Bench: benchmarking multimodal large language models for human-level planning.Int. J. Comput. Vis., 2026. URLhttps://arxiv.org/abs/2312.06722

  55. [61]

    BiGym: A demo-driven mobile bi-manual manipulation benchmark

    Nikita Chernyadev, Nicholas Backshall, Xiao Ma, Yunfan Lu, Younggyo Seo, and Stephen James. BiGym: A demo-driven mobile bi-manual manipulation benchmark. InConf. Robot Learn., 2024. URLhttps://arxiv.org/ abs/2407.07788

  56. [62]

    Universal manipulation interface: in-the-wild robot teaching without in-the-wild robots

    Cheng Chi, Zhenjia Xu, Chuer Pan, Eric Cousineau, Benjamin Burchfiel, Siyuan Feng, Russ Tedrake, and Shuran Song. Universal manipulation interface: in-the-wild robot teaching without in-the-wild robots. InRobot. Sci. Syst., 2024. URLhttps://arxiv.org/abs/2402.10329

  57. [63]

    Tacchi: A pluggable and low computational cost elastomer deformation simulator for optical tactile sensors.IEEE Robot

    Zixi Chen, Shixin Zhang, Shan Luo, Fuchun Sun, and Bin Fang. Tacchi: A pluggable and low computational cost elastomer deformation simulator for optical tactile sensors.IEEE Robot. Autom. Lett., 2023. URL https://arxiv.org/abs/2301.08343

  58. [64]

    Open X-Embodiment: Robotic learning datasets and RT-X models

    OX-Embodiment Collaboration, Abby O’Neill, Abdul Rehman, Abhinav Gupta, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, et al. Open X-Embodiment: Robotic learning datasets and RT-X models. InIEEE Int. Conf. Robot. Autom., 2024. URLhttps://arxiv.org/abs/2310. 08864

  59. [65]

    Kovalev, and Aleksandr I

    Egor Cherepanov, Nikita Kachaev, Alexey K. Kovalev, and Aleksandr I. Panov. Memory, benchmark & robots: a benchmark for solving complex tasks with reinforcement learning. InInt. Conf. Learn. Represent., 2026. URL https://doi.org/10.48550/arXiv.2502.10550. 49

  60. [66]

    PyBullet, a python module for physics simulation for games, robotics and machine learning.http://pybullet.org, 2016–2021

    Erwin Coumans and Yunfei Bai. PyBullet, a python module for physics simulation for games, robotics and machine learning.http://pybullet.org, 2016–2021. URLhttps://pybullet.org/wordpress/

  61. [67]

    A survey of sim-to-real methods in RL: progress, prospects and challenges with foundation models

    Longchao Da, Justin Turnau, Thirulogasankar Pranav Kutralingam, Alvaro Velasquez, Paulo Shakarian, and Hua Wei. A survey of sim-to-real methods in RL: progress, prospects and challenges with foundation models. arXiv preprint arXiv:2502.13187, 2025. URLhttps://arxiv.org/abs/2502.13187

  62. [69]

    URLhttps://arxiv.org/abs/2604.22748

  63. [70]

    RACER: Rich language-guided failure recovery policies for imitation learning

    Yinpei Dai, Jayjun Lee, Nima Fazeli, and Joyce Chai. RACER: Rich language-guided failure recovery policies for imitation learning. InIEEE Int. Conf. Robot. Autom., 2025. URLhttps://arxiv.org/abs/2409.14674

  64. [71]

    Yago Vicente, Thomas Dideriksen, Himanshu Arora, Matthieu Guillaumin, and Jitendra Malik

    Jasmine Collins, Shubham Goel, Kenan Deng, Achleshwar Luthra, Leon Xu, Erhan Gundogdu, Xi Zhang, Tomas F. Yago Vicente, Thomas Dideriksen, Himanshu Arora, Matthieu Guillaumin, and Jitendra Malik. ABO: Dataset and benchmarks for real-world 3D object understanding. InIEEE/CVF Conf. Comput. Vis. Pattern Recog.,

  65. [72]

    Scaling egocentric vision: the epic-kitchens dataset

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. Scaling egocentric vision: the epic-kitchens dataset. InEur. Conf. Comput. Vis., 2018. URLhttps://arxiv.org/abs/1804.02748

  66. [73]

    Rescaling egocentric vision.Int

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. Rescaling egocentric vision.Int. J. Comput. Vis., 2022. URLhttps://arxiv.org/abs/2006.13256

  67. [74]

    RynnBrain: open embodied foundation models.arXiv preprint arXiv:2602.14979, 2026

    Ronghao Dang, Jiayan Guo, Bohan Hou, Sicong Leng, Kehan Li, Xin Li, Jiangpin Liu, Yunxuan Mao, et al. RynnBrain: open embodied foundation models.arXiv preprint arXiv:2602.14979, 2026. URLhttps://arxiv.org/ abs/2602.14979

  68. [75]

    Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner

    Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. ScanNet: richly-annotated 3D reconstructions of indoor scenes. InIEEE/CVF Conf. Comput. Vis. Pattern Recog., 2017. URLhttps://arxiv.org/abs/1702.04405

  69. [77]

    URLhttps://arxiv.org/abs/2602.00557

  70. [78]

    Objaverse: a universe of annotated 3D objects

    Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: a universe of annotated 3D objects. InIEEE/CVF Conf. Comput. Vis. Pattern Recog., 2023. URLhttps://doi.org/10.1109/cvpr52729.2023.01263

  71. [79]

    RoboMME: benchmarking and understanding memory for robotic generalist policies

    Yinpei Dai, Hongze Fu, Jayjun Lee, Yuejiang Liu, Haoran Zhang, Jianing Yang, Chelsea Finn, Nima Fazeli, and Joyce Chai. RoboMME: benchmarking and understanding memory for robotic generalist policies. InInt. Conf. Mach. Learn., 2026. URLhttps://doi.org/10.48550/arXiv.2603.04639

  72. [80]

    GraspVLA: a grasping foundation model pre-trained on billion-scale synthetic action data

    Shengliang Deng, Mi Yan, Songlin Wei, Haixin Ma, Yuxin Yang, Jiayi Chen, Zhiqi Zhang, Taoyu Yang, et al. GraspVLA: a grasping foundation model pre-trained on billion-scale synthetic action data. InConf. Robot Learn.,

  73. [81]

    HumanNet: Scaling human-centric video learning to one million hours.arXiv preprint arXiv:2605.06747, 2026

    Yufan Deng and Daquan Zhou. HumanNet: Scaling human-centric video learning to one million hours.arXiv preprint arXiv:2605.06747, 2026. URLhttps://arxiv.org/abs/2605.06747

  74. [82]

    Rethinking video generation model for the embodied world

    Yufan Deng, Zilin Pan, Hongyu Zhang, Xiaojie Li, Ruoqing Hu, Yufei Ding, Yiming Zou, Yan Zeng, and Daquan Zhou. Rethinking video generation model for the embodied world. InInt. Conf. Mach. Learn., 2026. URL https://arxiv.org/abs/2601.15282

  75. [83]

    RoboNet: large-scale multi-robot learning

    Sudeep Dasari, Frederik Ebert, Stephen Tian, Suraj Nair, Bernadette Bucher, Karl Schmeckpeper, Siddharth Singh, Sergey Levine, and Chelsea Finn. RoboNet: large-scale multi-robot learning. InConf. Robot Learn., 2020. URLhttps://arxiv.org/abs/1910.11215

  76. [84]

    ProcTHOR: large-scale embodied AI using procedural generation

    Matt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs, Jordi Salvador, Kiana Ehsani, Winson Han, Eric Kolve, Ali Farhadi, Aniruddha Kembhavi, and Roozbeh Mottaghi. ProcTHOR: large-scale embodied AI using procedural generation. InAdv. Neural Inf. Process. Syst., 2022. URLhttps://doi.org/10.52202/068431-0433

  77. [85]

    Objaverse-xl: a universe of 10m+ 3D objects

    Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, et al. Objaverse-xl: a universe of 10m+ 3D objects. InAdv. Neural Inf. Process. Syst., 2023. URL https://doi.org/10.52202/075280-1554. 50

  78. [86]

    Bunny-VisionPro: Real-time bimanual dexterous teleoperation for imitation learning

    Runyu Ding, Yuzhe Qin, Jiyue Zhu, Chengzhe Jia, Shiqi Yang, Ruihan Yang, Xiaojuan Qi, and Xiaolong Wang. Bunny-VisionPro: Real-time bimanual dexterous teleoperation for imitation learning. InIEEE/RSJ Int. Conf. Intell. Robots Syst., 2025. URLhttps://arxiv.org/abs/2407.03162

  79. [87]

    Molmo and PixMo: Open weights and open data for state-of-the-art vision-language models

    Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi, Yue Yang, Jae Sung Park, Mohammadreza Salehi, Niklas Muennighoff, et al. Molmo and PixMo: Open weights and open data for state-of-the-art vision-language models. InIEEE/CVF Conf. Comput. Vis. Pattern Recog., 2025. URLhttps://arxiv.org/abs/2409.17146

  80. [88]

    Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, et al. PaLM-E: an embodied multimodal language model. InInt. Conf. Mach. Learn., 2023. URLhttps://arxiv.org/abs/2303.03378

Showing first 80 references.