{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:EJZD36WTQ5WLHPN5Y4USJ5MYAY","short_pith_number":"pith:EJZD36WT","schema_version":"1.0","canonical_sha256":"22723dfad3876cb3bdbdc72924f598063fe53b214f35564778b32c5067d7360d","source":{"kind":"arxiv","id":"2312.07533","version":4},"attestation_state":"computed","paper":{"title":"VILA: On Pre-training for Visual Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":[],"primary_cat":"cs.CV","authors_text":"Andrew Tao, Hongxu Yin, Huizi Mao, Jan Kautz, Ji Lin, Mohammad Shoeybi, Pavlo Molchanov, Song Han, Wei Ping, Yao Lu","submitted_at":"2023-12-12T18:58:18Z","abstract_excerpt":"Visual language models (VLMs) rapidly progressed with the recent success of large language models. There have been growing efforts on visual instruction tuning to extend the LLM with visual inputs, but lacks an in-depth study of the visual language pre-training process, where the model learns to perform joint modeling on both modalities. In this work, we examine the design options for VLM pre-training by augmenting LLM towards VLM through step-by-step controllable comparisons. We introduce three main findings: (1) freezing LLMs during pre-training can achieve decent zero-shot performance, but "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2312.07533","kind":"arxiv","version":4},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CV","submitted_at":"2023-12-12T18:58:18Z","cross_cats_sorted":[],"title_canon_sha256":"033e0857004416d965a0594f66e41e0e26b364d6b26afdb5410f100374e0e995","abstract_canon_sha256":"238c6b16de68d76b8e3654c311aca2dbd42b3449ff6933e4cfa34eee4cd511fb"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:19:57.883936Z","signature_b64":"Upjs9u33ZqkJieGTSZcGv5lUmgK71q57O2xaHr3t4YctfUMmMYdwhKTybA+i6+49O2p8d9ZK4sZSIJQ283aDDw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"22723dfad3876cb3bdbdc72924f598063fe53b214f35564778b32c5067d7360d","last_reissued_at":"2026-07-05T08:19:57.883458Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:19:57.883458Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"VILA: On Pre-training for Visual Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":[],"primary_cat":"cs.CV","authors_text":"Andrew Tao, Hongxu Yin, Huizi Mao, Jan Kautz, Ji Lin, Mohammad Shoeybi, Pavlo Molchanov, Song Han, Wei Ping, Yao Lu","submitted_at":"2023-12-12T18:58:18Z","abstract_excerpt":"Visual language models (VLMs) rapidly progressed with the recent success of large language models. There have been growing efforts on visual instruction tuning to extend the LLM with visual inputs, but lacks an in-depth study of the visual language pre-training process, where the model learns to perform joint modeling on both modalities. In this work, we examine the design options for VLM pre-training by augmenting LLM towards VLM through step-by-step controllable comparisons. We introduce three main findings: (1) freezing LLMs during pre-training can achieve decent zero-shot performance, but "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2312.07533","kind":"arxiv","version":4},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2312.07533/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2312.07533","created_at":"2026-07-05T08:19:57.883516+00:00"},{"alias_kind":"arxiv_version","alias_value":"2312.07533v4","created_at":"2026-07-05T08:19:57.883516+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2312.07533","created_at":"2026-07-05T08:19:57.883516+00:00"},{"alias_kind":"pith_short_12","alias_value":"EJZD36WTQ5WL","created_at":"2026-07-05T08:19:57.883516+00:00"},{"alias_kind":"pith_short_16","alias_value":"EJZD36WTQ5WLHPN5","created_at":"2026-07-05T08:19:57.883516+00:00"},{"alias_kind":"pith_short_8","alias_value":"EJZD36WT","created_at":"2026-07-05T08:19:57.883516+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":19,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.23581","citing_title":"Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse","ref_index":38,"is_internal_anchor":false},{"citing_arxiv_id":"2606.23557","citing_title":"Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views","ref_index":28,"is_internal_anchor":false},{"citing_arxiv_id":"2606.21734","citing_title":"HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning","ref_index":290,"is_internal_anchor":false},{"citing_arxiv_id":"2607.02089","citing_title":"ESC: Emotional Self-Correction for Reliable Vision-Language Models","ref_index":48,"is_internal_anchor":false},{"citing_arxiv_id":"2606.17539","citing_title":"Reinforcing Dual-Path Reasoning in Spatial Vision Language Models","ref_index":104,"is_internal_anchor":false},{"citing_arxiv_id":"2605.27318","citing_title":"Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2408.04840","citing_title":"mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models","ref_index":227,"is_internal_anchor":false},{"citing_arxiv_id":"2401.16158","citing_title":"Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2601.14724","citing_title":"HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2403.09611","citing_title":"MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training","ref_index":71,"is_internal_anchor":false},{"citing_arxiv_id":"2404.16994","citing_title":"PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2311.16502","citing_title":"MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI","ref_index":39,"is_internal_anchor":false},{"citing_arxiv_id":"2605.13328","citing_title":"What Limits Vision-and-Language Navigation ?","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2501.13826","citing_title":"Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2604.23941","citing_title":"GoClick: Lightweight Element Grounding Model for Autonomous GUI Interaction","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2506.01844","citing_title":"SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2604.10985","citing_title":"Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2604.09749","citing_title":"See Fair, Speak Truth: Equitable Attention Improves Grounding and Reduces Hallucination in Vision-Language Alignment","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2406.09246","citing_title":"OpenVLA: An Open-Source Vision-Language-Action Model","ref_index":88,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/EJZD36WTQ5WLHPN5Y4USJ5MYAY","json":"https://pith.science/pith/EJZD36WTQ5WLHPN5Y4USJ5MYAY.json","graph_json":"https://pith.science/api/pith-number/EJZD36WTQ5WLHPN5Y4USJ5MYAY/graph.json","events_json":"https://pith.science/api/pith-number/EJZD36WTQ5WLHPN5Y4USJ5MYAY/events.json","paper":"https://pith.science/paper/EJZD36WT"},"agent_actions":{"view_html":"https://pith.science/pith/EJZD36WTQ5WLHPN5Y4USJ5MYAY","download_json":"https://pith.science/pith/EJZD36WTQ5WLHPN5Y4USJ5MYAY.json","view_paper":"https://pith.science/paper/EJZD36WT","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2312.07533&json=true","fetch_graph":"https://pith.science/api/pith-number/EJZD36WTQ5WLHPN5Y4USJ5MYAY/graph.json","fetch_events":"https://pith.science/api/pith-number/EJZD36WTQ5WLHPN5Y4USJ5MYAY/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/EJZD36WTQ5WLHPN5Y4USJ5MYAY/action/timestamp_anchor","attest_storage":"https://pith.science/pith/EJZD36WTQ5WLHPN5Y4USJ5MYAY/action/storage_attestation","attest_author":"https://pith.science/pith/EJZD36WTQ5WLHPN5Y4USJ5MYAY/action/author_attestation","sign_citation":"https://pith.science/pith/EJZD36WTQ5WLHPN5Y4USJ5MYAY/action/citation_signature","submit_replication":"https://pith.science/pith/EJZD36WTQ5WLHPN5Y4USJ5MYAY/action/replication_record"}},"created_at":"2026-07-05T08:19:57.883516+00:00","updated_at":"2026-07-05T08:19:57.883516+00:00"}