{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:FK7GK5EV3IMN7ICWKQM3RNM2WV","short_pith_number":"pith:FK7GK5EV","schema_version":"1.0","canonical_sha256":"2abe657495da18dfa0565419b8b59ab56bd59c6fe90f5222e0c68461a2196e99","source":{"kind":"arxiv","id":"2403.09631","version":1},"attestation_state":"computed","paper":{"title":"3D-VLA: A 3D Vision-Language-Action Generative World Model","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"3D-VLA connects 3D perception to robot actions by embedding a generative world model inside a language model.","cross_cats":["cs.AI","cs.CL","cs.RO"],"primary_cat":"cs.CV","authors_text":"Chuang Gan, Haoyu Zhen, Jincheng Yang, Peihao Chen, Xiaowen Qiu, Xin Yan, Yilun Du, Yining Hong","submitted_at":"2024-03-14T17:58:41Z","abstract_excerpt":"Recent vision-language-action (VLA) models rely on 2D inputs, lacking integration with the broader realm of the 3D physical world. Furthermore, they perform action prediction by learning a direct mapping from perception to action, neglecting the vast dynamics of the world and the relations between actions and dynamics. In contrast, human beings are endowed with world models that depict imagination about future scenarios to plan actions accordingly. To this end, we propose 3D-VLA by introducing a new family of embodied foundation models that seamlessly link 3D perception, reasoning, and action "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":true,"formal_links_present":true},"canonical_record":{"source":{"id":"2403.09631","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CV","submitted_at":"2024-03-14T17:58:41Z","cross_cats_sorted":["cs.AI","cs.CL","cs.RO"],"title_canon_sha256":"7ee1dd59c180164676e4c9dca9a0adf39fd0d0976b6cfd38611f584c0bcf439a","abstract_canon_sha256":"b982b6157d5680c892d356911975d8aed838b02e6808a65f0c5469d9ca30a48a"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T07:56:12.328490Z","signature_b64":"lgTNmRaJ1En9RAukfMuH9Rtnqt8lEFJoYwx7Uh7koHTehZd7Q8ogC6fLkoZjrJb6iOgJVQTfCDdK9t0LJcsgBg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"2abe657495da18dfa0565419b8b59ab56bd59c6fe90f5222e0c68461a2196e99","last_reissued_at":"2026-07-05T07:56:12.327971Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T07:56:12.327971Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"3D-VLA: A 3D Vision-Language-Action Generative World Model","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"3D-VLA connects 3D perception to robot actions by embedding a generative world model inside a language model.","cross_cats":["cs.AI","cs.CL","cs.RO"],"primary_cat":"cs.CV","authors_text":"Chuang Gan, Haoyu Zhen, Jincheng Yang, Peihao Chen, Xiaowen Qiu, Xin Yan, Yilun Du, Yining Hong","submitted_at":"2024-03-14T17:58:41Z","abstract_excerpt":"Recent vision-language-action (VLA) models rely on 2D inputs, lacking integration with the broader realm of the 3D physical world. Furthermore, they perform action prediction by learning a direct mapping from perception to action, neglecting the vast dynamics of the world and the relations between actions and dynamics. In contrast, human beings are endowed with world models that depict imagination about future scenarios to plan actions accordingly. To this end, we propose 3D-VLA by introducing a new family of embodied foundation models that seamlessly link 3D perception, reasoning, and action "},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Our experiments on held-in datasets demonstrate that 3D-VLA significantly improves the reasoning, multimodal generation, and planning capabilities in embodied environments, showcasing its potential in real-world applications.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That a dataset curated by extracting 3D information from existing robotics datasets is diverse and representative enough to train a general-purpose 3D-VLA model that generalizes beyond the training distributions.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"3D-VLA is a new embodied foundation model that uses a 3D LLM plus aligned diffusion models to generate future images and point clouds for improved reasoning and action planning in 3D environments.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"3D-VLA connects 3D perception to robot actions by embedding a generative world model inside a language model.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"d117d967fefbf9976db9ca25bfc0619775dfab4e8fbeec3697fdae17e8acd89d"},"source":{"id":"2403.09631","kind":"arxiv","version":1},"verdict":{"id":"d853d7cb-65cd-4fef-86dc-52ded44ca0ef","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-13T18:13:35.728197Z","strongest_claim":"Our experiments on held-in datasets demonstrate that 3D-VLA significantly improves the reasoning, multimodal generation, and planning capabilities in embodied environments, showcasing its potential in real-world applications.","one_line_summary":"3D-VLA is a new embodied foundation model that uses a 3D LLM plus aligned diffusion models to generate future images and point clouds for improved reasoning and action planning in 3D environments.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That a dataset curated by extracting 3D information from existing robotics datasets is diverse and representative enough to train a general-purpose 3D-VLA model that generalizes beyond the training distributions.","pith_extraction_headline":"3D-VLA connects 3D perception to robot actions by embedding a generative world model inside a language model."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2403.09631/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":62,"sample":[{"doi":"","year":2022,"title":"Flamingo: a visual language model for few-shot learning","work_id":"8b5c09cb-7aa2-46b3-b62f-8f9b1a63df1a","ref_index":1,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2023,"title":"ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth","work_id":"c902e427-c54a-4fc4-8aef-d0243f90ea39","ref_index":2,"cited_arxiv_id":"2302.12288","is_internal_anchor":true},{"doi":"","year":2023,"title":"Zero-shot robotic manipulation with pretrained image-editing diffusion models","work_id":"7955fbb1-e5f4-4902-903c-a9032cfa1993","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2022,"title":"RT-1: Robotics Transformer for Real-World Control at Scale","work_id":"e11bda85-8531-46bc-a07f-d0ade3643ab1","ref_index":4,"cited_arxiv_id":"2212.06817","is_internal_anchor":true},{"doi":"","year":2023,"title":"RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control","work_id":"ff438a8a-8003-4fae-9131-acd418b3597b","ref_index":5,"cited_arxiv_id":"2307.15818","is_internal_anchor":true}],"resolved_work":62,"snapshot_sha256":"8f7657298a337098ffeda428273167b94ef357c012a000bda4381e781f4d20ae","internal_anchors":14},"formal_canon":{"evidence_count":3,"snapshot_sha256":"565a3c63fb5c933a3c9a717a1970ceb7b6d3f4341b8f7c9c283b9d0b9b6c3928"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2403.09631","created_at":"2026-07-05T07:56:12.328035+00:00"},{"alias_kind":"arxiv_version","alias_value":"2403.09631v1","created_at":"2026-07-05T07:56:12.328035+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2403.09631","created_at":"2026-07-05T07:56:12.328035+00:00"},{"alias_kind":"pith_short_12","alias_value":"FK7GK5EV3IMN","created_at":"2026-07-05T07:56:12.328035+00:00"},{"alias_kind":"pith_short_16","alias_value":"FK7GK5EV3IMN7ICW","created_at":"2026-07-05T07:56:12.328035+00:00"},{"alias_kind":"pith_short_8","alias_value":"FK7GK5EV","created_at":"2026-07-05T07:56:12.328035+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":74,"internal_anchor_count":74,"sample":[{"citing_arxiv_id":"2607.08448","citing_title":"Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents","ref_index":35,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26423","citing_title":"CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation","ref_index":47,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26800","citing_title":"SSI-Policy: Learning Structured Scene Interfaces for Vision-Language Robotic Manipulation","ref_index":29,"is_internal_anchor":true},{"citing_arxiv_id":"2606.20999","citing_title":"Inductive Generalization for Robotic Manipulation","ref_index":5,"is_internal_anchor":true},{"citing_arxiv_id":"2606.17598","citing_title":"MuseVLA: An Adaptive Multimodal Sensing Vision-Language-Action Model for Robotic Manipulation","ref_index":29,"is_internal_anchor":true},{"citing_arxiv_id":"2606.13053","citing_title":"EA-WM: Event-Aware World Models with Task-Specification Grounding for Long-Horizon Manipulation","ref_index":14,"is_internal_anchor":true},{"citing_arxiv_id":"2606.12403","citing_title":"World Pilot: Steering Vision-Language-Action Models with World-Action Priors","ref_index":2,"is_internal_anchor":true},{"citing_arxiv_id":"2606.10568","citing_title":"VeriSpace: Spatially Grounded Action Verification for Vision-Language-Action Models","ref_index":50,"is_internal_anchor":true},{"citing_arxiv_id":"2606.07100","citing_title":"LARA: Latent Action Representation Alignment for Vision-Language-Action Models","ref_index":53,"is_internal_anchor":true},{"citing_arxiv_id":"2606.06155","citing_title":"AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding","ref_index":84,"is_internal_anchor":true},{"citing_arxiv_id":"2607.00351","citing_title":"Unleashing More Actions via Action Compositional Training for VLA Models","ref_index":17,"is_internal_anchor":true},{"citing_arxiv_id":"2606.04436","citing_title":"3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training","ref_index":11,"is_internal_anchor":true},{"citing_arxiv_id":"2606.03943","citing_title":"PointAction: 3D Points as Universal Action Representations for Robot Control","ref_index":68,"is_internal_anchor":true},{"citing_arxiv_id":"2606.04264","citing_title":"UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation","ref_index":68,"is_internal_anchor":true},{"citing_arxiv_id":"2605.30877","citing_title":"Wall-OSS-0.5 Technical Report","ref_index":80,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26423","citing_title":"CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation","ref_index":47,"is_internal_anchor":true},{"citing_arxiv_id":"2605.14950","citing_title":"Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model","ref_index":57,"is_internal_anchor":true},{"citing_arxiv_id":"2606.07100","citing_title":"LARA: Latent Action Representation Alignment for Vision-Language-Action Models","ref_index":53,"is_internal_anchor":true},{"citing_arxiv_id":"2605.24642","citing_title":"Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models","ref_index":52,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26800","citing_title":"SSI-Policy: Learning Structured Scene Interfaces for Vision-Language Robotic Manipulation","ref_index":29,"is_internal_anchor":true},{"citing_arxiv_id":"2605.26282","citing_title":"Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization","ref_index":37,"is_internal_anchor":true},{"citing_arxiv_id":"2605.25829","citing_title":"OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation","ref_index":64,"is_internal_anchor":true},{"citing_arxiv_id":"2605.28548","citing_title":"GEM: Generative Supervision Helps Embodied Intelligence","ref_index":100,"is_internal_anchor":true},{"citing_arxiv_id":"2606.24472","citing_title":"G$^3$VLA: Geometric inductive bias for Vision-Language-Action Models","ref_index":9,"is_internal_anchor":true},{"citing_arxiv_id":"2506.14135","citing_title":"GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation","ref_index":66,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":3,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/FK7GK5EV3IMN7ICWKQM3RNM2WV","json":"https://pith.science/pith/FK7GK5EV3IMN7ICWKQM3RNM2WV.json","graph_json":"https://pith.science/api/pith-number/FK7GK5EV3IMN7ICWKQM3RNM2WV/graph.json","events_json":"https://pith.science/api/pith-number/FK7GK5EV3IMN7ICWKQM3RNM2WV/events.json","paper":"https://pith.science/paper/FK7GK5EV"},"agent_actions":{"view_html":"https://pith.science/pith/FK7GK5EV3IMN7ICWKQM3RNM2WV","download_json":"https://pith.science/pith/FK7GK5EV3IMN7ICWKQM3RNM2WV.json","view_paper":"https://pith.science/paper/FK7GK5EV","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2403.09631&json=true","fetch_graph":"https://pith.science/api/pith-number/FK7GK5EV3IMN7ICWKQM3RNM2WV/graph.json","fetch_events":"https://pith.science/api/pith-number/FK7GK5EV3IMN7ICWKQM3RNM2WV/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/FK7GK5EV3IMN7ICWKQM3RNM2WV/action/timestamp_anchor","attest_storage":"https://pith.science/pith/FK7GK5EV3IMN7ICWKQM3RNM2WV/action/storage_attestation","attest_author":"https://pith.science/pith/FK7GK5EV3IMN7ICWKQM3RNM2WV/action/author_attestation","sign_citation":"https://pith.science/pith/FK7GK5EV3IMN7ICWKQM3RNM2WV/action/citation_signature","submit_replication":"https://pith.science/pith/FK7GK5EV3IMN7ICWKQM3RNM2WV/action/replication_record"}},"created_at":"2026-07-05T07:56:12.328035+00:00","updated_at":"2026-07-05T07:56:12.328035+00:00"}