{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:UCNOS4BHI6NHTQ64SGZ7VWJWF2","short_pith_number":"pith:UCNOS4BH","schema_version":"1.0","canonical_sha256":"a09ae97027479a79c3dc91b3fad9362e95767759f4daa502c7e5fd12ba1b451c","source":{"kind":"arxiv","id":"2310.00653","version":1},"attestation_state":"computed","paper":{"title":"Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CV","authors_text":"Chongyi Wang, Dahai Li, Hai-Tao Zheng, Haoye Zhang, Jiao Xue, Jinyi Hu, Maosong Sun, Shan Wang, Tianyu Yu, Yinxv Pan, Yuan Yao, Yue Zhao, Zhiyuan Liu","submitted_at":"2023-10-01T12:35:18Z","abstract_excerpt":"Recent Multimodal Large Language Models (MLLMs) exhibit impressive abilities to perceive images and follow open-ended instructions. The capabilities of MLLMs depend on two crucial factors: the model architecture to facilitate the feature alignment of visual modules and large language models; the multimodal instruction tuning datasets for human instruction following. (i) For the model architecture, most existing models introduce an external bridge module to connect vision encoders with language models, which needs an additional feature-alignment pre-training. In this work, we discover that comp"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2310.00653","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CV","submitted_at":"2023-10-01T12:35:18Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"dd0aa8c1b66cc3f95c9569260f9ac42823bfeb7561ee020f303254348fe8dc88","abstract_canon_sha256":"6dc01e1873e56d223a220a4a228d8e141844f7d86a90003d68f763d47cb89577"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T06:56:08.743850Z","signature_b64":"H2hvOtCWeVd2OGGGUP+knZLBDyyDI7WY5DBJCe9tERxey+GZPnhPSorUIRTEc4WpWPqXf6uQgRTGGEhGGZAyBg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"a09ae97027479a79c3dc91b3fad9362e95767759f4daa502c7e5fd12ba1b451c","last_reissued_at":"2026-07-05T06:56:08.743386Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T06:56:08.743386Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CV","authors_text":"Chongyi Wang, Dahai Li, Hai-Tao Zheng, Haoye Zhang, Jiao Xue, Jinyi Hu, Maosong Sun, Shan Wang, Tianyu Yu, Yinxv Pan, Yuan Yao, Yue Zhao, Zhiyuan Liu","submitted_at":"2023-10-01T12:35:18Z","abstract_excerpt":"Recent Multimodal Large Language Models (MLLMs) exhibit impressive abilities to perceive images and follow open-ended instructions. The capabilities of MLLMs depend on two crucial factors: the model architecture to facilitate the feature alignment of visual modules and large language models; the multimodal instruction tuning datasets for human instruction following. (i) For the model architecture, most existing models introduce an external bridge module to connect vision encoders with language models, which needs an additional feature-alignment pre-training. In this work, we discover that comp"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2310.00653","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2310.00653/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2310.00653","created_at":"2026-07-05T06:56:08.743467+00:00"},{"alias_kind":"arxiv_version","alias_value":"2310.00653v1","created_at":"2026-07-05T06:56:08.743467+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2310.00653","created_at":"2026-07-05T06:56:08.743467+00:00"},{"alias_kind":"pith_short_12","alias_value":"UCNOS4BHI6NH","created_at":"2026-07-05T06:56:08.743467+00:00"},{"alias_kind":"pith_short_16","alias_value":"UCNOS4BHI6NHTQ64","created_at":"2026-07-05T06:56:08.743467+00:00"},{"alias_kind":"pith_short_8","alias_value":"UCNOS4BH","created_at":"2026-07-05T06:56:08.743467+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":6,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.11853","citing_title":"Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning","ref_index":168,"is_internal_anchor":false},{"citing_arxiv_id":"2412.04468","citing_title":"NVILA: Efficient Frontier Visual Language Models","ref_index":130,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15300","citing_title":"Deep Pre-Alignment for VLMs","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2404.18930","citing_title":"Hallucination of Multimodal Large Language Models: A Survey","ref_index":197,"is_internal_anchor":false},{"citing_arxiv_id":"2408.01800","citing_title":"MiniCPM-V: A GPT-4V Level MLLM on Your Phone","ref_index":110,"is_internal_anchor":false},{"citing_arxiv_id":"2306.13394","citing_title":"MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models","ref_index":56,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/UCNOS4BHI6NHTQ64SGZ7VWJWF2","json":"https://pith.science/pith/UCNOS4BHI6NHTQ64SGZ7VWJWF2.json","graph_json":"https://pith.science/api/pith-number/UCNOS4BHI6NHTQ64SGZ7VWJWF2/graph.json","events_json":"https://pith.science/api/pith-number/UCNOS4BHI6NHTQ64SGZ7VWJWF2/events.json","paper":"https://pith.science/paper/UCNOS4BH"},"agent_actions":{"view_html":"https://pith.science/pith/UCNOS4BHI6NHTQ64SGZ7VWJWF2","download_json":"https://pith.science/pith/UCNOS4BHI6NHTQ64SGZ7VWJWF2.json","view_paper":"https://pith.science/paper/UCNOS4BH","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2310.00653&json=true","fetch_graph":"https://pith.science/api/pith-number/UCNOS4BHI6NHTQ64SGZ7VWJWF2/graph.json","fetch_events":"https://pith.science/api/pith-number/UCNOS4BHI6NHTQ64SGZ7VWJWF2/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/UCNOS4BHI6NHTQ64SGZ7VWJWF2/action/timestamp_anchor","attest_storage":"https://pith.science/pith/UCNOS4BHI6NHTQ64SGZ7VWJWF2/action/storage_attestation","attest_author":"https://pith.science/pith/UCNOS4BHI6NHTQ64SGZ7VWJWF2/action/author_attestation","sign_citation":"https://pith.science/pith/UCNOS4BHI6NHTQ64SGZ7VWJWF2/action/citation_signature","submit_replication":"https://pith.science/pith/UCNOS4BHI6NHTQ64SGZ7VWJWF2/action/replication_record"}},"created_at":"2026-07-05T06:56:08.743467+00:00","updated_at":"2026-07-05T06:56:08.743467+00:00"}