{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:KVR52WS45ROSKGTD33OBKUT3HY","short_pith_number":"pith:KVR52WS4","schema_version":"1.0","canonical_sha256":"5563dd5a5cec5d251a63dedc15527b3e25266f3acb2ac37f955c40cece128427","source":{"kind":"arxiv","id":"2506.00411","version":1},"attestation_state":"computed","paper":{"title":"LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.RO","authors_text":"Jiaxuan Sun, Siqi Kou, Yihan Wang, Yi Yang, Zhijie Deng","submitted_at":"2025-05-31T06:01:03Z","abstract_excerpt":"Real-world embodied agents face long-horizon tasks, characterized by high-level goals demanding multi-step solutions beyond single actions. Successfully navigating these requires both high-level task planning (i.e., decomposing goals into sub-tasks) and low-level motion control (i.e., generating precise robot actions). While existing vision language action (VLA) models and hierarchical architectures offer potential in embodied tasks, the former often falter in planning, and the latter can suffer from coordination issues, both hampering performance. We introduce a new unified VLA framework for "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2506.00411","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.RO","submitted_at":"2025-05-31T06:01:03Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"2b6b47cd1e887e4f06cbb2ed9976ed86ace1ba8d3d0d643aedad1b0fe6577e0b","abstract_canon_sha256":"41d41b1020e063976fd3ce7137c5ce9178f6967c4c603ad057814cce57cb1b22"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:13:34.335048Z","signature_b64":"/jkmB7s9op6m8qqJJo8Et/0SD4msFrWPD1tK5maGiNg8KRsMCSeT3G5GPNr+jlQuptZPbzANMsS+NPdXT/GmBw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"5563dd5a5cec5d251a63dedc15527b3e25266f3acb2ac37f955c40cece128427","last_reissued_at":"2026-07-05T11:13:34.334507Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:13:34.334507Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.RO","authors_text":"Jiaxuan Sun, Siqi Kou, Yihan Wang, Yi Yang, Zhijie Deng","submitted_at":"2025-05-31T06:01:03Z","abstract_excerpt":"Real-world embodied agents face long-horizon tasks, characterized by high-level goals demanding multi-step solutions beyond single actions. Successfully navigating these requires both high-level task planning (i.e., decomposing goals into sub-tasks) and low-level motion control (i.e., generating precise robot actions). While existing vision language action (VLA) models and hierarchical architectures offer potential in embodied tasks, the former often falter in planning, and the latter can suffer from coordination issues, both hampering performance. We introduce a new unified VLA framework for "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2506.00411","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2506.00411/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2506.00411","created_at":"2026-07-05T11:13:34.334578+00:00"},{"alias_kind":"arxiv_version","alias_value":"2506.00411v1","created_at":"2026-07-05T11:13:34.334578+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2506.00411","created_at":"2026-07-05T11:13:34.334578+00:00"},{"alias_kind":"pith_short_12","alias_value":"KVR52WS45ROS","created_at":"2026-07-05T11:13:34.334578+00:00"},{"alias_kind":"pith_short_16","alias_value":"KVR52WS45ROSKGTD","created_at":"2026-07-05T11:13:34.334578+00:00"},{"alias_kind":"pith_short_8","alias_value":"KVR52WS4","created_at":"2026-07-05T11:13:34.334578+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":9,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.12402","citing_title":"DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2606.12497","citing_title":"$\\mu$VLA: On Recurrent Memory for Partially Observable Manipulation in VLA Models","ref_index":72,"is_internal_anchor":false},{"citing_arxiv_id":"2606.05660","citing_title":"Safe Embodied AI for Long-horizon Tasks: A Cross-layer Analysis of Robotic Manipulation","ref_index":280,"is_internal_anchor":false},{"citing_arxiv_id":"2508.13073","citing_title":"Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey","ref_index":106,"is_internal_anchor":false},{"citing_arxiv_id":"2602.13193","citing_title":"Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control","ref_index":45,"is_internal_anchor":false},{"citing_arxiv_id":"2605.13119","citing_title":"Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models","ref_index":32,"is_internal_anchor":false},{"citing_arxiv_id":"2604.02965","citing_title":"Open-Loop Planning, Closed-Loop Verification: Speculative Verification for VLA","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2604.08340","citing_title":"Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time","ref_index":86,"is_internal_anchor":false},{"citing_arxiv_id":"2604.17880","citing_title":"ST-$\\pi$: Structured SpatioTemporal VLA for Robotic Manipulation","ref_index":39,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/KVR52WS45ROSKGTD33OBKUT3HY","json":"https://pith.science/pith/KVR52WS45ROSKGTD33OBKUT3HY.json","graph_json":"https://pith.science/api/pith-number/KVR52WS45ROSKGTD33OBKUT3HY/graph.json","events_json":"https://pith.science/api/pith-number/KVR52WS45ROSKGTD33OBKUT3HY/events.json","paper":"https://pith.science/paper/KVR52WS4"},"agent_actions":{"view_html":"https://pith.science/pith/KVR52WS45ROSKGTD33OBKUT3HY","download_json":"https://pith.science/pith/KVR52WS45ROSKGTD33OBKUT3HY.json","view_paper":"https://pith.science/paper/KVR52WS4","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2506.00411&json=true","fetch_graph":"https://pith.science/api/pith-number/KVR52WS45ROSKGTD33OBKUT3HY/graph.json","fetch_events":"https://pith.science/api/pith-number/KVR52WS45ROSKGTD33OBKUT3HY/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/KVR52WS45ROSKGTD33OBKUT3HY/action/timestamp_anchor","attest_storage":"https://pith.science/pith/KVR52WS45ROSKGTD33OBKUT3HY/action/storage_attestation","attest_author":"https://pith.science/pith/KVR52WS45ROSKGTD33OBKUT3HY/action/author_attestation","sign_citation":"https://pith.science/pith/KVR52WS45ROSKGTD33OBKUT3HY/action/citation_signature","submit_replication":"https://pith.science/pith/KVR52WS45ROSKGTD33OBKUT3HY/action/replication_record"}},"created_at":"2026-07-05T11:13:34.334578+00:00","updated_at":"2026-07-05T11:13:34.334578+00:00"}