{"work":{"id":"4b158d3e-3dff-4412-85cd-baa879465a5e","openalex_id":"https://openalex.org/W4405031329","doi":"10.48550/arxiv.2411.19650","arxiv_id":"2411.19650","raw_key":null,"title":"CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation","authors":null,"authors_text":"Qixiu Li, Yaobo Liang, Zeyu Wang, Lin Luo, Xi Chen, Mozheng Liao","year":2024,"venue":"cs.RO","abstract":"The advancement of large Vision-Language-Action (VLA) models has significantly improved robotic manipulation in terms of language-guided task execution and generalization to unseen scenarios. While existing VLAs adapted from pretrained large Vision-Language-Models (VLM) have demonstrated promising generalizability, their task performance is still unsatisfactory as indicated by the low tasks success rates in different environments. In this paper, we present a new advanced VLA architecture derived from VLM. Unlike previous works that directly repurpose VLM for action prediction by simple action quantization, we propose a omponentized VLA architecture that has a specialized action module conditioned on VLM output. We systematically study the design of the action module and demonstrates the strong performance enhancement with diffusion action transformers for action sequence modeling, as well as their favorable scaling behaviors. We also conduct comprehensive experiments and ablation studies to evaluate the efficacy of our models with varied designs. The evaluation on 5 robot embodiments in simulation and real work shows that our model not only significantly surpasses existing VLAs in task performance and but also exhibits remarkable adaptation to new robots and generalization to unseen objects and backgrounds. It exceeds the average success rates of OpenVLA which has similar model size (7B) with ours by over 35% in simulated evaluation and 55% in real robot experiments. It also outperforms the large RT-2-X model (55B) by 18% absolute success rates in simulation. Code and models can be found on our project page (https://cogact.github.io/).","external_url":"https://arxiv.org/abs/2411.19650","cited_by_count":5,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2411.19650","created_at":"2026-05-09T06:05:34.791328+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation","render_title":"CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation"},"hub":{"state":{"work_id":"4b158d3e-3dff-4412-85cd-baa879465a5e","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":127,"external_cited_by_count":5,"distinct_field_count":7,"first_pith_cited_at":"2024-05-23T01:43:54+00:00","last_pith_cited_at":"2026-07-09T16:15:43+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-17T03:19:21.252379+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":26},{"context_role":"baseline","n":5},{"context_role":"method","n":2},{"context_role":"other","n":1}],"polarity_counts":[{"context_polarity":"background","n":25},{"context_polarity":"baseline","n":6},{"context_polarity":"unclear","n":2},{"context_polarity":"use_method","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation","claims":[{"claim_text":"The advancement of large Vision-Language-Action (VLA) models has significantly improved robotic manipulation in terms of language-guided task execution and generalization to unseen scenarios. While existing VLAs adapted from pretrained large Vision-Language-Models (VLM) have demonstrated promising generalizability, their task performance is still unsatisfactory as indicated by the low tasks success rates in different environments. In this paper, we present a new advanced VLA architecture derived from VLM. Unlike previous works that directly repurpose VLM for action prediction by simple action ","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Global and local prompts cooperation via optimal transport for federated learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12151-12161, 2024. [38] Qinbin Li, Bingsheng He, and Dawn Song. Model-contrastive federated learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10713-10722, 2021. [39] Qixiu Li, Yaobo Liang, Zeyu Wang, Lin Luo, Xi Chen, Mozheng Liao, Fangyun Wei, Yu Deng, Sicheng Xu, Yizhong Z","claim_type":"other","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"Concat DDPM OXE, [SC]Sim: MetaWorld;Real(Franka, UR5): place, stack, flip mug, close drawer, open box CogACT [122]♢DINOv2, SigLIP LLaMA 2 DiT Concat DDIM OXE, [SC]Sim: SimplerEnv;Real(Realman, Franka): pick, place, stack, open/close oven DexVLA [123]♢ViT, ResNet-50 Qwen2-VL 2B, DistilBERT, ScaleDP Concat DDPM [SC]Real(Franka, UR5e, AgileX): pick, fold shirt, bus table, pour HybridVLA [124]♢DINOv2, SigLIP, CLIP LLaMA 2, Phi-2 MLP Concat DDIM OXE, DROIDSim: RLBench;Real(Franka, AgileX): pick- plac","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"and1×10 −5/32 for stages 1, 2, and 3, respectively. 7 Table 1: Quantitative results of VLAs for fine-tuned robotic manipulation tasks. Representation Category Method LIBERO [40] CSOT-Bench (Ours) Spatial Object Goal Long Average Scene Object Task Average Implicit OpenVLA [2] 84.7 88.4 79.2 53.7 76.5 74.6 82.4 72.0 76.3 Octo [20] 78.9 85.7 84.6 51.1 75.1 66.3 78.0 71.5 72.0 CogACT [42] 87.5 90.2 78.4 53.2 76.5 76.1 84.5 73.8 78.1 DiffusionPolicy [8] 78.3 92.5 68.3 50.5 72.4 68.7 80.2 65.3 71.4 Sp","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"strong zero-shot transfer: RT-1 [6] unifies vision, language, and action in a single transformer for real-time kitchen manipulation; RT-2 [5] jointly finetunes large vision-language models on web and robot data to support semantic planning and object reasoning; diffusion-based RDT-1B [28] andπ[3] learn diverse bimanual dynamics from over a million episodes. Vision-language-action systems such as OpenVLA [20] and CogACT [24], together with adaptations like Octo [35], LAPA [55], and OpenVLA-OFT [1","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"[23] Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. Offline reinforcement learning: Tutorial, review, and perspectives on open problems.arXiv preprint arXiv:2005.01643, 2020. 3 [24] Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li, Zejun Ma, and Chunyuan Li. Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models.arXiv preprint arXiv:2407.07895, 2024. 2 [25] Qixiu Li, Yaobo Liang, Zeyu Wang, Lin Luo, Xi Chen, Mozheng Liao, Fangyun Wei, Yu D","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Ren, Homer Walke, Quan Vuong, Lucy Xiaoyang Shi, and Sergey Levine. Knowledge insulating vision-language-action models: Train fast, run fast, generalize better, 2025. [26] NVIDIA GEAR Team, Allison Azzolini, Johan Bjorck, Valts Blukis, et al. Gr00t n1.6: An im- proved open foundation model for generalist humanoid robots. https://research.nvidia. com/labs/gear/gr00t-n1_6/, December 2025. [27] Qixiu Li, Yaobo Liang, Zeyu Wang, Lin Luo, Xi Chen, Mozheng Liao, Fangyun Wei, Yu Deng, Sicheng Xu, Yizho","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (26 contexts).","role_counts":[{"n":26,"context_role":"background"},{"n":5,"context_role":"baseline"},{"n":2,"context_role":"method"},{"n":1,"context_role":"other"}]},"error":null,"updated_at":"2026-07-02T22:03:00.506770+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"3fe1716c-d186-49d7-bdd9-011717d279f8","orcid":null,"display_name":"Qixiu Li"},{"id":"8d9ec7c8-98c4-4cb1-baca-04ecf7cb963b","orcid":null,"display_name":"Yaobo Liang"},{"id":"dc369dc7-07dc-487b-bc4f-d0c626b571bb","orcid":null,"display_name":"Zeyu Wang"},{"id":"10c07dc1-5748-4aa2-ba19-0c8e329bcf14","orcid":null,"display_name":"Lin Luo"},{"id":"3c910835-0332-416e-87ba-3e49dd8c1b91","orcid":null,"display_name":"Xi Chen"},{"id":"31d3c180-a31a-4349-9aaa-cabadcceecf6","orcid":null,"display_name":"Mozheng Liao"}]},"error":null,"updated_at":"2026-07-02T22:03:01.058136+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T18:20:09.498448+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"$\\pi_0$: A Vision-Language-Action Flow Model for General Robot Control","work_id":"f790abdc-a796-482f-a40d-f8ee035ecfc2","shared_citers":26},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model","work_id":"3e7e65c5-5aed-4fe9-8414-2092bcb31cc7","shared_citers":23},{"title":"Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success","work_id":"04f46bb3-4346-47e8-bf09-c75d91f96e87","shared_citers":20},{"title":"RT-1: Robotics Transformer for Real-World Control at Scale","work_id":"e11bda85-8531-46bc-a07f-d0ade3643ab1","shared_citers":19},{"title":"$\\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization","work_id":"d1ad7304-d09a-49bc-809e-846439f6aff9","shared_citers":18},{"title":"Octo: An Open-Source Generalist Robot Policy","work_id":"f9ca0722-8855-48c3-a27a-0eefb7e19253","shared_citers":18},{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots","work_id":"e2db69c7-ee8a-4cb7-a761-7b8de1dfcf97","shared_citers":16},{"title":"FAST: Efficient Action Tokenization for Vision-Language-Action Models","work_id":"83a8f966-6cfa-4f21-81f3-87440aae238f","shared_citers":13},{"title":"RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation","work_id":"12319725-bc7d-4c32-a229-ad270a7460bc","shared_citers":13},{"title":"SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model","work_id":"592041b3-3ca2-4836-8dd4-f8095d8a692b","shared_citers":12},{"title":"GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation","work_id":"843ab5eb-2815-4db8-b3bc-890b23fa5ffa","shared_citers":10},{"title":"RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control","work_id":"ff438a8a-8003-4fae-9131-acd418b3597b","shared_citers":10},{"title":"WorldVLA: Towards Autoregressive Action World Model","work_id":"d8c0c873-b2fc-44a5-a0c8-0d4a698783fb","shared_citers":10},{"title":"DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset","work_id":"13253de2-3d89-415c-8c2f-3adb25d4c337","shared_citers":9},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":9},{"title":"SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics","work_id":"0c5e9314-5fa7-4613-ad12-605a71d561d2","shared_citers":9},{"title":"X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model","work_id":"13faca8d-e96d-4e6c-a441-9f2683d11934","shared_citers":9},{"title":"Dexvla: Vision-language model with plug-in diffusion expert for general robot control","work_id":"3564a757-5726-4b2a-a28e-114a4a467dfb","shared_citers":8},{"title":"LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning","work_id":"662203ad-084f-42c4-8e60-977b3173755b","shared_citers":8},{"title":"LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models","work_id":"e35c8c6d-977d-4af1-963a-766ba98703ce","shared_citers":8},{"title":"RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation","work_id":"9b985126-4a2f-4bdf-b014-2a7524ec634e","shared_citers":8},{"title":"3D-VLA: A 3D Vision-Language-Action Generative World Model","work_id":"aebf924c-e761-437e-9cee-f1ccc2e427bd","shared_citers":7},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":7},{"title":"Gemini Robotics: Bringing AI into the Physical World","work_id":"f7c5ce10-8364-4fbe-964f-2802b81c3a98","shared_citers":7}],"time_series":[{"n":2,"year":2025},{"n":33,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T18:19:58.236690+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T18:19:38.161293+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation","claims":[{"claim_text":"The advancement of large Vision-Language-Action (VLA) models has significantly improved robotic manipulation in terms of language-guided task execution and generalization to unseen scenarios. While existing VLAs adapted from pretrained large Vision-Language-Models (VLM) have demonstrated promising generalizability, their task performance is still unsatisfactory as indicated by the low tasks success rates in different environments. In this paper, we present a new advanced VLA architecture derived from VLM. Unlike previous works that directly repurpose VLM for action prediction by simple action ","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Global and local prompts cooperation via optimal transport for federated learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12151-12161, 2024. [38] Qinbin Li, Bingsheng He, and Dawn Song. Model-contrastive federated learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10713-10722, 2021. [39] Qixiu Li, Yaobo Liang, Zeyu Wang, Lin Luo, Xi Chen, Mozheng Liao, Fangyun Wei, Yu Deng, Sicheng Xu, Yizhong Z","claim_type":"other","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"Concat DDPM OXE, [SC]Sim: MetaWorld;Real(Franka, UR5): place, stack, flip mug, close drawer, open box CogACT [122]♢DINOv2, SigLIP LLaMA 2 DiT Concat DDIM OXE, [SC]Sim: SimplerEnv;Real(Realman, Franka): pick, place, stack, open/close oven DexVLA [123]♢ViT, ResNet-50 Qwen2-VL 2B, DistilBERT, ScaleDP Concat DDPM [SC]Real(Franka, UR5e, AgileX): pick, fold shirt, bus table, pour HybridVLA [124]♢DINOv2, SigLIP, CLIP LLaMA 2, Phi-2 MLP Concat DDIM OXE, DROIDSim: RLBench;Real(Franka, AgileX): pick- plac","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"and1×10 −5/32 for stages 1, 2, and 3, respectively. 7 Table 1: Quantitative results of VLAs for fine-tuned robotic manipulation tasks. Representation Category Method LIBERO [40] CSOT-Bench (Ours) Spatial Object Goal Long Average Scene Object Task Average Implicit OpenVLA [2] 84.7 88.4 79.2 53.7 76.5 74.6 82.4 72.0 76.3 Octo [20] 78.9 85.7 84.6 51.1 75.1 66.3 78.0 71.5 72.0 CogACT [42] 87.5 90.2 78.4 53.2 76.5 76.1 84.5 73.8 78.1 DiffusionPolicy [8] 78.3 92.5 68.3 50.5 72.4 68.7 80.2 65.3 71.4 Sp","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"strong zero-shot transfer: RT-1 [6] unifies vision, language, and action in a single transformer for real-time kitchen manipulation; RT-2 [5] jointly finetunes large vision-language models on web and robot data to support semantic planning and object reasoning; diffusion-based RDT-1B [28] andπ[3] learn diverse bimanual dynamics from over a million episodes. Vision-language-action systems such as OpenVLA [20] and CogACT [24], together with adaptations like Octo [35], LAPA [55], and OpenVLA-OFT [1","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"[23] Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. Offline reinforcement learning: Tutorial, review, and perspectives on open problems.arXiv preprint arXiv:2005.01643, 2020. 3 [24] Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li, Zejun Ma, and Chunyuan Li. Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models.arXiv preprint arXiv:2407.07895, 2024. 2 [25] Qixiu Li, Yaobo Liang, Zeyu Wang, Lin Luo, Xi Chen, Mozheng Liao, Fangyun Wei, Yu D","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Ren, Homer Walke, Quan Vuong, Lucy Xiaoyang Shi, and Sergey Levine. Knowledge insulating vision-language-action models: Train fast, run fast, generalize better, 2025. [26] NVIDIA GEAR Team, Allison Azzolini, Johan Bjorck, Valts Blukis, et al. Gr00t n1.6: An im- proved open foundation model for generalist humanoid robots. https://research.nvidia. com/labs/gear/gr00t-n1_6/, December 2025. [27] Qixiu Li, Yaobo Liang, Zeyu Wang, Lin Luo, Xi Chen, Mozheng Liao, Fangyun Wei, Yu Deng, Sicheng Xu, Yizho","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (26 contexts).","role_counts":[{"n":26,"context_role":"background"},{"n":5,"context_role":"baseline"},{"n":2,"context_role":"method"},{"n":1,"context_role":"other"}]},"error":null,"updated_at":"2026-07-02T22:03:01.060587+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation","claims":[{"claim_text":"The advancement of large Vision-Language-Action (VLA) models has significantly improved robotic manipulation in terms of language-guided task execution and generalization to unseen scenarios. While existing VLAs adapted from pretrained large Vision-Language-Models (VLM) have demonstrated promising generalizability, their task performance is still unsatisfactory as indicated by the low tasks success rates in different environments. In this paper, we present a new advanced VLA architecture derived from VLM. Unlike previous works that directly repurpose VLM for action prediction by simple action ","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T18:19:58.121274+00:00"}},"summary":{"title":"CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation","claims":[{"claim_text":"The advancement of large Vision-Language-Action (VLA) models has significantly improved robotic manipulation in terms of language-guided task execution and generalization to unseen scenarios. While existing VLAs adapted from pretrained large Vision-Language-Models (VLM) have demonstrated promising generalizability, their task performance is still unsatisfactory as indicated by the low tasks success rates in different environments. In this paper, we present a new advanced VLA architecture derived from VLM. Unlike previous works that directly repurpose VLM for action prediction by simple action ","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"$\\pi_0$: A Vision-Language-Action Flow Model for General Robot Control","work_id":"f790abdc-a796-482f-a40d-f8ee035ecfc2","shared_citers":26},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model","work_id":"3e7e65c5-5aed-4fe9-8414-2092bcb31cc7","shared_citers":23},{"title":"Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success","work_id":"04f46bb3-4346-47e8-bf09-c75d91f96e87","shared_citers":20},{"title":"RT-1: Robotics Transformer for Real-World Control at Scale","work_id":"e11bda85-8531-46bc-a07f-d0ade3643ab1","shared_citers":19},{"title":"$\\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization","work_id":"d1ad7304-d09a-49bc-809e-846439f6aff9","shared_citers":18},{"title":"Octo: An Open-Source Generalist Robot Policy","work_id":"f9ca0722-8855-48c3-a27a-0eefb7e19253","shared_citers":18},{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots","work_id":"e2db69c7-ee8a-4cb7-a761-7b8de1dfcf97","shared_citers":16},{"title":"FAST: Efficient Action Tokenization for Vision-Language-Action Models","work_id":"83a8f966-6cfa-4f21-81f3-87440aae238f","shared_citers":13},{"title":"RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation","work_id":"12319725-bc7d-4c32-a229-ad270a7460bc","shared_citers":13},{"title":"SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model","work_id":"592041b3-3ca2-4836-8dd4-f8095d8a692b","shared_citers":12},{"title":"GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation","work_id":"843ab5eb-2815-4db8-b3bc-890b23fa5ffa","shared_citers":10},{"title":"RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control","work_id":"ff438a8a-8003-4fae-9131-acd418b3597b","shared_citers":10},{"title":"WorldVLA: Towards Autoregressive Action World Model","work_id":"d8c0c873-b2fc-44a5-a0c8-0d4a698783fb","shared_citers":10},{"title":"DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset","work_id":"13253de2-3d89-415c-8c2f-3adb25d4c337","shared_citers":9},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":9},{"title":"SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics","work_id":"0c5e9314-5fa7-4613-ad12-605a71d561d2","shared_citers":9},{"title":"X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model","work_id":"13faca8d-e96d-4e6c-a441-9f2683d11934","shared_citers":9},{"title":"Dexvla: Vision-language model with plug-in diffusion expert for general robot control","work_id":"3564a757-5726-4b2a-a28e-114a4a467dfb","shared_citers":8},{"title":"LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning","work_id":"662203ad-084f-42c4-8e60-977b3173755b","shared_citers":8},{"title":"LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models","work_id":"e35c8c6d-977d-4af1-963a-766ba98703ce","shared_citers":8},{"title":"RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation","work_id":"9b985126-4a2f-4bdf-b014-2a7524ec634e","shared_citers":8},{"title":"3D-VLA: A 3D Vision-Language-Action Generative World Model","work_id":"aebf924c-e761-437e-9cee-f1ccc2e427bd","shared_citers":7},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":7},{"title":"Gemini Robotics: Bringing AI into the Physical World","work_id":"f7c5ce10-8364-4fbe-964f-2802b81c3a98","shared_citers":7}],"time_series":[{"n":2,"year":2025},{"n":33,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"10c07dc1-5748-4aa2-ba19-0c8e329bcf14","orcid":null,"display_name":"Lin Luo","source":"manual","import_confidence":0.72},{"id":"31d3c180-a31a-4349-9aaa-cabadcceecf6","orcid":null,"display_name":"Mozheng Liao","source":"manual","import_confidence":0.72},{"id":"3fe1716c-d186-49d7-bdd9-011717d279f8","orcid":null,"display_name":"Qixiu Li","source":"manual","import_confidence":0.72},{"id":"3c910835-0332-416e-87ba-3e49dd8c1b91","orcid":null,"display_name":"Xi Chen","source":"manual","import_confidence":0.72},{"id":"8d9ec7c8-98c4-4cb1-baca-04ecf7cb963b","orcid":null,"display_name":"Yaobo Liang","source":"manual","import_confidence":0.72},{"id":"dc369dc7-07dc-487b-bc4f-d0c626b571bb","orcid":null,"display_name":"Zeyu Wang","source":"manual","import_confidence":0.72}]}}