{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:H6WEYW6UZJ7E5WKC4T52ZW2HL2","short_pith_number":"pith:H6WEYW6U","schema_version":"1.0","canonical_sha256":"3fac4c5bd4ca7e4ed942e4fbacdb475e9a89c0a262156d1b5a84b030096cb4e1","source":{"kind":"arxiv","id":"2505.19381","version":4},"attestation_state":"computed","paper":{"title":"DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CV","cs.RO"],"primary_cat":"cs.AI","authors_text":"Anqing Jiang, Hao Jiang, Hao Sun, Hao Zhao, Jijun Wang, Jinghao Chai, Qian Cao, Xianda Guo, Yiru Wang, Yu Gao, Yunda Dong, Yuweng Heng, Zhigang Sun, Zongzheng Zhang","submitted_at":"2025-05-26T00:49:35Z","abstract_excerpt":"Research interest in end-to-end autonomous driving has surged owing to its fully differentiable design integrating modular tasks, i.e. perception, prediction and planing, which enables optimization in pursuit of the ultimate goal. Despite the great potential of the end-to-end paradigm, existing methods suffer from several aspects including expensive BEV (bird's eye view) computation, action diversity, and sub-optimal decision in complex real-world scenarios. To address these challenges, we propose a novel hybrid sparse-dense diffusion policy, empowered by a Vision-Language Model (VLM), called "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2505.19381","kind":"arxiv","version":4},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.AI","submitted_at":"2025-05-26T00:49:35Z","cross_cats_sorted":["cs.CV","cs.RO"],"title_canon_sha256":"17103aa54308ba55fa53ff4ebc401ac9ff4d483a71b26da392a743e508f83e2d","abstract_canon_sha256":"5c2f91f34aeea4610d5e9184d3019ba35738ca922096f6ef6c91fb55cf1e9587"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:14:29.490873Z","signature_b64":"zxSYc0eyQVg2V9Z7kFAkVElaIgWNPpqZl2H2YuN5uWWzjnDZyYdmV7tzjJNCFKrpWeXopbwFP0pJNiRUgze9Dw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"3fac4c5bd4ca7e4ed942e4fbacdb475e9a89c0a262156d1b5a84b030096cb4e1","last_reissued_at":"2026-07-05T11:14:29.490394Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:14:29.490394Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CV","cs.RO"],"primary_cat":"cs.AI","authors_text":"Anqing Jiang, Hao Jiang, Hao Sun, Hao Zhao, Jijun Wang, Jinghao Chai, Qian Cao, Xianda Guo, Yiru Wang, Yu Gao, Yunda Dong, Yuweng Heng, Zhigang Sun, Zongzheng Zhang","submitted_at":"2025-05-26T00:49:35Z","abstract_excerpt":"Research interest in end-to-end autonomous driving has surged owing to its fully differentiable design integrating modular tasks, i.e. perception, prediction and planing, which enables optimization in pursuit of the ultimate goal. Despite the great potential of the end-to-end paradigm, existing methods suffer from several aspects including expensive BEV (bird's eye view) computation, action diversity, and sub-optimal decision in complex real-world scenarios. To address these challenges, we propose a novel hybrid sparse-dense diffusion policy, empowered by a Vision-Language Model (VLM), called "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2505.19381","kind":"arxiv","version":4},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2505.19381/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2505.19381","created_at":"2026-07-05T11:14:29.490455+00:00"},{"alias_kind":"arxiv_version","alias_value":"2505.19381v4","created_at":"2026-07-05T11:14:29.490455+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2505.19381","created_at":"2026-07-05T11:14:29.490455+00:00"},{"alias_kind":"pith_short_12","alias_value":"H6WEYW6UZJ7E","created_at":"2026-07-05T11:14:29.490455+00:00"},{"alias_kind":"pith_short_16","alias_value":"H6WEYW6UZJ7E5WKC","created_at":"2026-07-05T11:14:29.490455+00:00"},{"alias_kind":"pith_short_8","alias_value":"H6WEYW6U","created_at":"2026-07-05T11:14:29.490455+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":19,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2607.08375","citing_title":"WCog-VLA: A Dual-Level World-Cognitive Vision-Language-Action Model for End-to-End Autonomous Driving","ref_index":23,"is_internal_anchor":true},{"citing_arxiv_id":"2606.12396","citing_title":"VLGA: Vision-Language-Geometry-Action Models for Autonomous Driving","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2606.07170","citing_title":"Test-Time Trajectory Optimization for Autonomous Driving","ref_index":61,"is_internal_anchor":false},{"citing_arxiv_id":"2605.31476","citing_title":"IDOL: Inverse-Dynamics-Guided Future Prediction for End-to-End Autonomous Driving","ref_index":25,"is_internal_anchor":false},{"citing_arxiv_id":"2606.31830","citing_title":"PriorEye: Geospatial Visual Priors for End-to-End Autonomous Driving","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2606.29879","citing_title":"LWDrive: Layer-Wise World-Model-Guided Vision-Language Model Planning for Autonomous Driving","ref_index":50,"is_internal_anchor":false},{"citing_arxiv_id":"2606.29879","citing_title":"LWDrive: Layer-Wise World-Model-Guided Vision-Language Model Planning for Autonomous Driving","ref_index":50,"is_internal_anchor":false},{"citing_arxiv_id":"2605.23270","citing_title":"ChainFlow-VLA: Causal Flow Planning with Vision-Language Models","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22089","citing_title":"LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2507.17596","citing_title":"PRIX: Learning to Plan from Raw Pixels for End-to-End Autonomous Driving","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2510.12796","citing_title":"DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving","ref_index":21,"is_internal_anchor":false},{"citing_arxiv_id":"2603.07686","citing_title":"UniUncer: Unified Dynamic Static Uncertainty for End to End Driving","ref_index":35,"is_internal_anchor":false},{"citing_arxiv_id":"2603.13842","citing_title":"Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10426","citing_title":"CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2604.00813","citing_title":"DVGT-2: Vision-Geometry-Action Model for Autonomous Driving at Scale","ref_index":24,"is_internal_anchor":false},{"citing_arxiv_id":"2604.02714","citing_title":"ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving","ref_index":21,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08975","citing_title":"Latency Analysis and Optimization of Alpamayo 1 via Efficient Trajectory Generation","ref_index":23,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10426","citing_title":"CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2605.09701","citing_title":"DriveFuture: Future-Aware Latent World Models for Autonomous Driving","ref_index":60,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/H6WEYW6UZJ7E5WKC4T52ZW2HL2","json":"https://pith.science/pith/H6WEYW6UZJ7E5WKC4T52ZW2HL2.json","graph_json":"https://pith.science/api/pith-number/H6WEYW6UZJ7E5WKC4T52ZW2HL2/graph.json","events_json":"https://pith.science/api/pith-number/H6WEYW6UZJ7E5WKC4T52ZW2HL2/events.json","paper":"https://pith.science/paper/H6WEYW6U"},"agent_actions":{"view_html":"https://pith.science/pith/H6WEYW6UZJ7E5WKC4T52ZW2HL2","download_json":"https://pith.science/pith/H6WEYW6UZJ7E5WKC4T52ZW2HL2.json","view_paper":"https://pith.science/paper/H6WEYW6U","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2505.19381&json=true","fetch_graph":"https://pith.science/api/pith-number/H6WEYW6UZJ7E5WKC4T52ZW2HL2/graph.json","fetch_events":"https://pith.science/api/pith-number/H6WEYW6UZJ7E5WKC4T52ZW2HL2/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/H6WEYW6UZJ7E5WKC4T52ZW2HL2/action/timestamp_anchor","attest_storage":"https://pith.science/pith/H6WEYW6UZJ7E5WKC4T52ZW2HL2/action/storage_attestation","attest_author":"https://pith.science/pith/H6WEYW6UZJ7E5WKC4T52ZW2HL2/action/author_attestation","sign_citation":"https://pith.science/pith/H6WEYW6UZJ7E5WKC4T52ZW2HL2/action/citation_signature","submit_replication":"https://pith.science/pith/H6WEYW6UZJ7E5WKC4T52ZW2HL2/action/replication_record"}},"created_at":"2026-07-05T11:14:29.490455+00:00","updated_at":"2026-07-05T11:14:29.490455+00:00"}