{"work":{"id":"13253de2-3d89-415c-8c2f-3adb25d4c337","openalex_id":"https://openalex.org/W4393027744","doi":"10.48550/arxiv.2403.12945","arxiv_id":"2403.12945","raw_key":null,"title":"DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset","authors":null,"authors_text":"Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, Siddharth Karamcheti","year":2024,"venue":"cs.RO","abstract":"The creation of large, diverse, high-quality robot manipulation datasets is an important stepping stone on the path toward more capable and robust robotic manipulation policies. However, creating such datasets is challenging: collecting robot manipulation data in diverse environments poses logistical and safety challenges and requires substantial investments in hardware and human labour. As a result, even the most general robot manipulation policies today are mostly trained on data collected in a small number of environments with limited scene and task diversity. In this work, we introduce DROID (Distributed Robot Interaction Dataset), a diverse robot manipulation dataset with 76k demonstration trajectories or 350 hours of interaction data, collected across 564 scenes and 84 tasks by 50 data collectors in North America, Asia, and Europe over the course of 12 months. We demonstrate that training with DROID leads to policies with higher performance and improved generalization ability. We open source the full dataset, policy learning code, and a detailed guide for reproducing our robot hardware setup.","external_url":"https://arxiv.org/abs/2403.12945","cited_by_count":3,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2403.12945","created_at":"2026-05-09T05:10:16.367226+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset","render_title":"DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset"},"hub":{"state":{"work_id":"13253de2-3d89-415c-8c2f-3adb25d4c337","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":213,"external_cited_by_count":3,"distinct_field_count":6,"first_pith_cited_at":"2024-05-09T17:30:16+00:00","last_pith_cited_at":"2026-07-09T17:30:49+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T21:49:23.562872+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"dataset","n":23},{"context_role":"background","n":15},{"context_role":"method","n":2},{"context_role":"baseline","n":1},{"context_role":"other","n":1}],"polarity_counts":[{"context_polarity":"use_dataset","n":22},{"context_polarity":"background","n":15},{"context_polarity":"unclear","n":3},{"context_polarity":"baseline","n":1},{"context_polarity":"use_method","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset","claims":[{"claim_text":"The creation of large, diverse, high-quality robot manipulation datasets is an important stepping stone on the path toward more capable and robust robotic manipulation policies. However, creating such datasets is challenging: collecting robot manipulation data in diverse environments poses logistical and safety challenges and requires substantial investments in hardware and human labour. As a result, even the most general robot manipulation policies today are mostly trained on data collected in a small number of environments with limited scene and task diversity. In this work, we introduce DRO","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Batch size is set to 64, we set the query length of each modality 9 and diffusion steps in DiT to 10. We weight the dynamic region, depth and segmentation prediction losses as λdyn=0.1, λdepth=0.001, λsem=0.1, and the action loss as λDiT=1, respectively. We first pre-train DreamVLA on the language-free split of the CALVIN [117] and on the full DROID dataset [82]. For the LIBERO benchmark, we first pretrain DreamVLA on LIBERO-90 and then finetune on each track. The model predicts entire frames in","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"[29] Ahmed Khaled, Konstantin Mishchenko, and Peter Richtarik. Tighter theory for local SGD on identical and heterogeneous data. InProceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics, volume 108 ofProceedings of Machine Learning Research, pages 4519-4529. PMLR, 2020. URLhttps://proceedings.mlr.press/v108/bayoumi20a.html. [30] Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, Siddharth Karamcheti, Soroush Nasiriany, Mo","claim_type":"other","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"15 PhysX-Mobility [97] CVPR '26 CAD Models 2K+ objects / 47 cat. Geometry, Kinematics, Physics (per-part)✓ ✓ ✓ /external-link-alt 16 DTC [205] CVPR '25 Real Scans 2K objects / 40 cat. Geometry, 4K PBR, Evaluation Seqs. ✗ ✗ ✓ /external-link-alt 17 ManiTwin [206] arXiv '26 Mixed 100K+ objects Geometry, Manipulation Anno., Sim-ready▲ ▲ ✓ /external-link-alt DROID [214], RH20T [215], BridgeData V2 [216]-provide high-fidelity trajectories but are costly to scale. Simulation- based benchmarks leverage ","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"then SpatialVLA is trained with a standard auto-regressive next-token prediction objective in eq. (1). Importantly, the embeddings of text tokens Etext are frozen to maintain the general world knowledge in pre-trained VLM, and the experimental results show this frozen operation is beneficial for the instruction following ability. Moreover, as discussed in OpenVLA [30], DROID dataset [29] are removed from the data mixture for the final third of pre-training to improve the quality of the pre-train","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"3, 32, 36 [37] Tao Jiang, Peng Lu, Li Zhang, Ningsheng Ma, Rui Han, Chengqi Lyu, Yining Li, and Kai Chen. Rtmpose: Real-time multi-person pose estimation based on mmpose.arXiv preprint arXiv:2303.07399, 2023. 7 [38] Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models.NeurIPS, 2022. 8 [39] AlexanderKhazatsky, KarlPertsch, SurajNair, AshwinBalakrishna, SudeepDasari, SiddharthKaramcheti, Soroush Nasiriany, Mohan Kumar Srirama, L","claim_type":"background","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"This refined primitive distribu- tion provides evidence thatR&B-EnCoRecan effectively filter out patently irrelevant information, generating high-quality relevant reasoning strategies tailored to the specific needs of various embodiments of legged locomotion navigation. D. Autonomous Vehicles We extendR&B-EnCoReto the autonomous vehicle (A V) nuScenes dataset [32] and study how reasoning traces for driv- ing VLAs changes under our method. We leverage traces from pioneering LLM agent-based planne","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset because it crossed a citation-hub threshold. Current citing contexts most often use it as dataset evidence (22 contexts).","role_counts":[{"n":22,"context_role":"dataset"},{"n":15,"context_role":"background"},{"n":2,"context_role":"method"},{"n":1,"context_role":"baseline"},{"n":1,"context_role":"other"}]},"error":null,"updated_at":"2026-05-23T01:34:06.861008+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"2ffa9e44-7b49-4ef0-b5fc-720f113bdc87","orcid":null,"display_name":"Alexander Khazatsky"},{"id":"9821de6a-6143-4939-8865-7d9c02424af9","orcid":null,"display_name":"Karl Pertsch"},{"id":"863a3a51-2bb9-4ec3-9cf1-7535b3809d8b","orcid":null,"display_name":"Suraj Nair"},{"id":"c277791e-8df4-4262-99f1-6b73bf8a525e","orcid":null,"display_name":"Ashwin Balakrishna"},{"id":"1be5ea08-317a-40b7-aeb9-b5cce80f6662","orcid":null,"display_name":"Sudeep Dasari"},{"id":"9a5d4ee6-e5d5-4afd-89cf-8f80084b69ba","orcid":null,"display_name":"Siddharth Karamcheti"}]},"error":null,"updated_at":"2026-05-23T01:34:06.854568+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T12:10:05.749408+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"$\\pi_0$: A Vision-Language-Action Flow Model for General Robot Control","work_id":"f790abdc-a796-482f-a40d-f8ee035ecfc2","shared_citers":36},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model","work_id":"3e7e65c5-5aed-4fe9-8414-2092bcb31cc7","shared_citers":27},{"title":"$\\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization","work_id":"d1ad7304-d09a-49bc-809e-846439f6aff9","shared_citers":26},{"title":"RT-1: Robotics Transformer for Real-World Control at Scale","work_id":"e11bda85-8531-46bc-a07f-d0ade3643ab1","shared_citers":24},{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots","work_id":"e2db69c7-ee8a-4cb7-a761-7b8de1dfcf97","shared_citers":23},{"title":"RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control","work_id":"ff438a8a-8003-4fae-9131-acd418b3597b","shared_citers":20},{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware","work_id":"6fe159e0-fa73-481a-88d4-4719c15140be","shared_citers":15},{"title":"Octo: An Open-Source Generalist Robot Policy","work_id":"f9ca0722-8855-48c3-a27a-0eefb7e19253","shared_citers":15},{"title":"Open X-Embodiment: Robotic Learning Datasets and RT-X Models","work_id":"62f0fb6c-e6ae-4dc4-95a4-d9dd64b240e8","shared_citers":14},{"title":"FAST: Efficient Action Tokenization for Vision-Language-Action Models","work_id":"83a8f966-6cfa-4f21-81f3-87440aae238f","shared_citers":13},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":13},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":12},{"title":"Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success","work_id":"04f46bb3-4346-47e8-bf09-c75d91f96e87","shared_citers":11},{"title":"RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation","work_id":"9b985126-4a2f-4bdf-b014-2a7524ec634e","shared_citers":11},{"title":"RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation","work_id":"12319725-bc7d-4c32-a229-ad270a7460bc","shared_citers":10},{"title":"CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation","work_id":"4b158d3e-3dff-4412-85cd-baa879465a5e","shared_citers":9},{"title":"Do As I Can, Not As I Say: Grounding Language in Robotic Affordances","work_id":"037320f1-b0a9-4cbe-a639-bfb25409ce71","shared_citers":9},{"title":"PaliGemma: A versatile 3B VLM for transfer","work_id":"df6f48b3-5792-47c7-9614-cb856ea31ad9","shared_citers":9},{"title":"SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model","work_id":"592041b3-3ca2-4836-8dd4-f8095d8a692b","shared_citers":9},{"title":"Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning","work_id":"3d63039f-41b0-4a31-af31-6fc10f5c1b1b","shared_citers":8},{"title":"LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning","work_id":"662203ad-084f-42c4-8e60-977b3173755b","shared_citers":8},{"title":"PaLM-E: An Embodied Multimodal Language Model","work_id":"5b99811a-1d93-47e2-9d59-f4045a0b74a2","shared_citers":8},{"title":"World Action Models are Zero-shot Policies","work_id":"9a85fc69-74df-450e-94cd-69d186e9e830","shared_citers":8},{"title":"3D-VLA: A 3D Vision-Language-Action Generative World Model","work_id":"aebf924c-e761-437e-9cee-f1ccc2e427bd","shared_citers":7}],"time_series":[{"n":3,"year":2024},{"n":6,"year":2025},{"n":48,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T12:10:05.805038+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T12:09:54.438119+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset","claims":[{"claim_text":"The creation of large, diverse, high-quality robot manipulation datasets is an important stepping stone on the path toward more capable and robust robotic manipulation policies. However, creating such datasets is challenging: collecting robot manipulation data in diverse environments poses logistical and safety challenges and requires substantial investments in hardware and human labour. As a result, even the most general robot manipulation policies today are mostly trained on data collected in a small number of environments with limited scene and task diversity. In this work, we introduce DRO","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Batch size is set to 64, we set the query length of each modality 9 and diffusion steps in DiT to 10. We weight the dynamic region, depth and segmentation prediction losses as λdyn=0.1, λdepth=0.001, λsem=0.1, and the action loss as λDiT=1, respectively. We first pre-train DreamVLA on the language-free split of the CALVIN [117] and on the full DROID dataset [82]. For the LIBERO benchmark, we first pretrain DreamVLA on LIBERO-90 and then finetune on each track. The model predicts entire frames in","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"[29] Ahmed Khaled, Konstantin Mishchenko, and Peter Richtarik. Tighter theory for local SGD on identical and heterogeneous data. InProceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics, volume 108 ofProceedings of Machine Learning Research, pages 4519-4529. PMLR, 2020. URLhttps://proceedings.mlr.press/v108/bayoumi20a.html. [30] Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, Siddharth Karamcheti, Soroush Nasiriany, Mo","claim_type":"other","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"15 PhysX-Mobility [97] CVPR '26 CAD Models 2K+ objects / 47 cat. Geometry, Kinematics, Physics (per-part)✓ ✓ ✓ /external-link-alt 16 DTC [205] CVPR '25 Real Scans 2K objects / 40 cat. Geometry, 4K PBR, Evaluation Seqs. ✗ ✗ ✓ /external-link-alt 17 ManiTwin [206] arXiv '26 Mixed 100K+ objects Geometry, Manipulation Anno., Sim-ready▲ ▲ ✓ /external-link-alt DROID [214], RH20T [215], BridgeData V2 [216]-provide high-fidelity trajectories but are costly to scale. Simulation- based benchmarks leverage ","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"then SpatialVLA is trained with a standard auto-regressive next-token prediction objective in eq. (1). Importantly, the embeddings of text tokens Etext are frozen to maintain the general world knowledge in pre-trained VLM, and the experimental results show this frozen operation is beneficial for the instruction following ability. Moreover, as discussed in OpenVLA [30], DROID dataset [29] are removed from the data mixture for the final third of pre-training to improve the quality of the pre-train","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"3, 32, 36 [37] Tao Jiang, Peng Lu, Li Zhang, Ningsheng Ma, Rui Han, Chengqi Lyu, Yining Li, and Kai Chen. Rtmpose: Real-time multi-person pose estimation based on mmpose.arXiv preprint arXiv:2303.07399, 2023. 7 [38] Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models.NeurIPS, 2022. 8 [39] AlexanderKhazatsky, KarlPertsch, SurajNair, AshwinBalakrishna, SudeepDasari, SiddharthKaramcheti, Soroush Nasiriany, Mohan Kumar Srirama, L","claim_type":"background","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"This refined primitive distribu- tion provides evidence thatR&B-EnCoRecan effectively filter out patently irrelevant information, generating high-quality relevant reasoning strategies tailored to the specific needs of various embodiments of legged locomotion navigation. D. Autonomous Vehicles We extendR&B-EnCoReto the autonomous vehicle (A V) nuScenes dataset [32] and study how reasoning traces for driv- ing VLAs changes under our method. We leverage traces from pioneering LLM agent-based planne","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset because it crossed a citation-hub threshold. Current citing contexts most often use it as dataset evidence (22 contexts).","role_counts":[{"n":22,"context_role":"dataset"},{"n":15,"context_role":"background"},{"n":2,"context_role":"method"},{"n":1,"context_role":"baseline"},{"n":1,"context_role":"other"}]},"error":null,"updated_at":"2026-05-23T01:34:06.865098+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset","claims":[{"claim_text":"The creation of large, diverse, high-quality robot manipulation datasets is an important stepping stone on the path toward more capable and robust robotic manipulation policies. However, creating such datasets is challenging: collecting robot manipulation data in diverse environments poses logistical and safety challenges and requires substantial investments in hardware and human labour. As a result, even the most general robot manipulation policies today are mostly trained on data collected in a small number of environments with limited scene and task diversity. In this work, we introduce DRO","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T12:10:01.221632+00:00"}},"summary":{"title":"DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset","claims":[{"claim_text":"The creation of large, diverse, high-quality robot manipulation datasets is an important stepping stone on the path toward more capable and robust robotic manipulation policies. However, creating such datasets is challenging: collecting robot manipulation data in diverse environments poses logistical and safety challenges and requires substantial investments in hardware and human labour. As a result, even the most general robot manipulation policies today are mostly trained on data collected in a small number of environments with limited scene and task diversity. In this work, we introduce DRO","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"$\\pi_0$: A Vision-Language-Action Flow Model for General Robot Control","work_id":"f790abdc-a796-482f-a40d-f8ee035ecfc2","shared_citers":36},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model","work_id":"3e7e65c5-5aed-4fe9-8414-2092bcb31cc7","shared_citers":27},{"title":"$\\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization","work_id":"d1ad7304-d09a-49bc-809e-846439f6aff9","shared_citers":26},{"title":"RT-1: Robotics Transformer for Real-World Control at Scale","work_id":"e11bda85-8531-46bc-a07f-d0ade3643ab1","shared_citers":24},{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots","work_id":"e2db69c7-ee8a-4cb7-a761-7b8de1dfcf97","shared_citers":23},{"title":"RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control","work_id":"ff438a8a-8003-4fae-9131-acd418b3597b","shared_citers":20},{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware","work_id":"6fe159e0-fa73-481a-88d4-4719c15140be","shared_citers":15},{"title":"Octo: An Open-Source Generalist Robot Policy","work_id":"f9ca0722-8855-48c3-a27a-0eefb7e19253","shared_citers":15},{"title":"Open X-Embodiment: Robotic Learning Datasets and RT-X Models","work_id":"62f0fb6c-e6ae-4dc4-95a4-d9dd64b240e8","shared_citers":14},{"title":"FAST: Efficient Action Tokenization for Vision-Language-Action Models","work_id":"83a8f966-6cfa-4f21-81f3-87440aae238f","shared_citers":13},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":13},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":12},{"title":"Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success","work_id":"04f46bb3-4346-47e8-bf09-c75d91f96e87","shared_citers":11},{"title":"RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation","work_id":"9b985126-4a2f-4bdf-b014-2a7524ec634e","shared_citers":11},{"title":"RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation","work_id":"12319725-bc7d-4c32-a229-ad270a7460bc","shared_citers":10},{"title":"CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation","work_id":"4b158d3e-3dff-4412-85cd-baa879465a5e","shared_citers":9},{"title":"Do As I Can, Not As I Say: Grounding Language in Robotic Affordances","work_id":"037320f1-b0a9-4cbe-a639-bfb25409ce71","shared_citers":9},{"title":"PaliGemma: A versatile 3B VLM for transfer","work_id":"df6f48b3-5792-47c7-9614-cb856ea31ad9","shared_citers":9},{"title":"SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model","work_id":"592041b3-3ca2-4836-8dd4-f8095d8a692b","shared_citers":9},{"title":"Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning","work_id":"3d63039f-41b0-4a31-af31-6fc10f5c1b1b","shared_citers":8},{"title":"LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning","work_id":"662203ad-084f-42c4-8e60-977b3173755b","shared_citers":8},{"title":"PaLM-E: An Embodied Multimodal Language Model","work_id":"5b99811a-1d93-47e2-9d59-f4045a0b74a2","shared_citers":8},{"title":"World Action Models are Zero-shot Policies","work_id":"9a85fc69-74df-450e-94cd-69d186e9e830","shared_citers":8},{"title":"3D-VLA: A 3D Vision-Language-Action Generative World Model","work_id":"aebf924c-e761-437e-9cee-f1ccc2e427bd","shared_citers":7}],"time_series":[{"n":3,"year":2024},{"n":6,"year":2025},{"n":48,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"2ffa9e44-7b49-4ef0-b5fc-720f113bdc87","orcid":null,"display_name":"Alexander Khazatsky","source":"manual","import_confidence":0.72},{"id":"c277791e-8df4-4262-99f1-6b73bf8a525e","orcid":null,"display_name":"Ashwin Balakrishna","source":"manual","import_confidence":0.72},{"id":"9821de6a-6143-4939-8865-7d9c02424af9","orcid":null,"display_name":"Karl Pertsch","source":"manual","import_confidence":0.72},{"id":"9a5d4ee6-e5d5-4afd-89cf-8f80084b69ba","orcid":null,"display_name":"Siddharth Karamcheti","source":"manual","import_confidence":0.72},{"id":"1be5ea08-317a-40b7-aeb9-b5cce80f6662","orcid":null,"display_name":"Sudeep Dasari","source":"manual","import_confidence":0.72},{"id":"863a3a51-2bb9-4ec3-9cf1-7535b3809d8b","orcid":null,"display_name":"Suraj Nair","source":"manual","import_confidence":0.72}]}}