{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:NCCZAI3Q5ZRPA5KVPJ4PREWCWY","short_pith_number":"pith:NCCZAI3Q","schema_version":"1.0","canonical_sha256":"6885902370ee62f075557a78f892c2b61dac8dbc3ff7915e6cbbd63049e476ce","source":{"kind":"arxiv","id":"2406.11815","version":1},"attestation_state":"computed","paper":{"title":"LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CV","cs.LG"],"primary_cat":"cs.RO","authors_text":"Baifeng Shi, Dantong Niu, Giscard Biamby, Jerome Quenum, Roei Herzig, Trevor Darrell, Yutong Bai, Yuvan Sharma","submitted_at":"2024-06-17T17:55:29Z","abstract_excerpt":"In recent years, instruction-tuned Large Multimodal Models (LMMs) have been successful at several tasks, including image captioning and visual question answering; yet leveraging these models remains an open question for robotics. Prior LMMs for robotics applications have been extensively trained on language and action data, but their ability to generalize in different settings has often been less than desired. To address this, we introduce LLARVA, a model trained with a novel instruction tuning method that leverages structured prompts to unify a range of robotic learning tasks, scenarios, and "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2406.11815","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.RO","submitted_at":"2024-06-17T17:55:29Z","cross_cats_sorted":["cs.CV","cs.LG"],"title_canon_sha256":"8913ee406cbe0eda60b0f180bf5d2f440ef039745128b400bbe14019005e649e","abstract_canon_sha256":"52f6dfd82684d92535a492e6e7a6916c065b58f4b9cb62ac441d2e9b928f8c07"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:32:58.183608Z","signature_b64":"ok/4IWEK5oHvE6JeABcyyXwgRa88L38EK9wkv6lidIJdZndhwIJd6ru2Zev2CqhX0eJu3FaeHmcVcfvLodwmDg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"6885902370ee62f075557a78f892c2b61dac8dbc3ff7915e6cbbd63049e476ce","last_reissued_at":"2026-07-05T08:32:58.183180Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:32:58.183180Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CV","cs.LG"],"primary_cat":"cs.RO","authors_text":"Baifeng Shi, Dantong Niu, Giscard Biamby, Jerome Quenum, Roei Herzig, Trevor Darrell, Yutong Bai, Yuvan Sharma","submitted_at":"2024-06-17T17:55:29Z","abstract_excerpt":"In recent years, instruction-tuned Large Multimodal Models (LMMs) have been successful at several tasks, including image captioning and visual question answering; yet leveraging these models remains an open question for robotics. Prior LMMs for robotics applications have been extensively trained on language and action data, but their ability to generalize in different settings has often been less than desired. To address this, we introduce LLARVA, a model trained with a novel instruction tuning method that leverages structured prompts to unify a range of robotic learning tasks, scenarios, and "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2406.11815","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2406.11815/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2406.11815","created_at":"2026-07-05T08:32:58.183237+00:00"},{"alias_kind":"arxiv_version","alias_value":"2406.11815v1","created_at":"2026-07-05T08:32:58.183237+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2406.11815","created_at":"2026-07-05T08:32:58.183237+00:00"},{"alias_kind":"pith_short_12","alias_value":"NCCZAI3Q5ZRP","created_at":"2026-07-05T08:32:58.183237+00:00"},{"alias_kind":"pith_short_16","alias_value":"NCCZAI3Q5ZRPA5KV","created_at":"2026-07-05T08:32:58.183237+00:00"},{"alias_kind":"pith_short_8","alias_value":"NCCZAI3Q","created_at":"2026-07-05T08:32:58.183237+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":8,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.12978","citing_title":"Trajectory-Level Redirection Attacks on Vision-Language-Action Models","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2606.03385","citing_title":"Grasp-Then-Plan with Failure Attribution: A Closed Two-Stage Framework for Precise and Generalizable Robotic Manipulation","ref_index":75,"is_internal_anchor":false},{"citing_arxiv_id":"2605.30231","citing_title":"Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning","ref_index":36,"is_internal_anchor":false},{"citing_arxiv_id":"2504.16054","citing_title":"$\\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization","ref_index":61,"is_internal_anchor":false},{"citing_arxiv_id":"2507.16815","citing_title":"ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning","ref_index":30,"is_internal_anchor":false},{"citing_arxiv_id":"2601.07060","citing_title":"PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation","ref_index":93,"is_internal_anchor":false},{"citing_arxiv_id":"2412.10345","citing_title":"TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies","ref_index":89,"is_internal_anchor":false},{"citing_arxiv_id":"2605.01448","citing_title":"Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic Manipulation","ref_index":16,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/NCCZAI3Q5ZRPA5KVPJ4PREWCWY","json":"https://pith.science/pith/NCCZAI3Q5ZRPA5KVPJ4PREWCWY.json","graph_json":"https://pith.science/api/pith-number/NCCZAI3Q5ZRPA5KVPJ4PREWCWY/graph.json","events_json":"https://pith.science/api/pith-number/NCCZAI3Q5ZRPA5KVPJ4PREWCWY/events.json","paper":"https://pith.science/paper/NCCZAI3Q"},"agent_actions":{"view_html":"https://pith.science/pith/NCCZAI3Q5ZRPA5KVPJ4PREWCWY","download_json":"https://pith.science/pith/NCCZAI3Q5ZRPA5KVPJ4PREWCWY.json","view_paper":"https://pith.science/paper/NCCZAI3Q","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2406.11815&json=true","fetch_graph":"https://pith.science/api/pith-number/NCCZAI3Q5ZRPA5KVPJ4PREWCWY/graph.json","fetch_events":"https://pith.science/api/pith-number/NCCZAI3Q5ZRPA5KVPJ4PREWCWY/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/NCCZAI3Q5ZRPA5KVPJ4PREWCWY/action/timestamp_anchor","attest_storage":"https://pith.science/pith/NCCZAI3Q5ZRPA5KVPJ4PREWCWY/action/storage_attestation","attest_author":"https://pith.science/pith/NCCZAI3Q5ZRPA5KVPJ4PREWCWY/action/author_attestation","sign_citation":"https://pith.science/pith/NCCZAI3Q5ZRPA5KVPJ4PREWCWY/action/citation_signature","submit_replication":"https://pith.science/pith/NCCZAI3Q5ZRPA5KVPJ4PREWCWY/action/replication_record"}},"created_at":"2026-07-05T08:32:58.183237+00:00","updated_at":"2026-07-05T08:32:58.183237+00:00"}