{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:MCS4PBZV7EKQ2XTRVWUD3BUO7Y","short_pith_number":"pith:MCS4PBZV","schema_version":"1.0","canonical_sha256":"60a5c78735f9150d5e71ada83d868efe3effa1627f37b632889e127301a82a76","source":{"kind":"arxiv","id":"2406.08418","version":3},"attestation_state":"computed","paper":{"title":"OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CV","authors_text":"Bin Wang, Botian Shi, Bo Zhang, Changyao Tian, Chao Xu, Conghui He, Dahua Lin, Erfei Cui, Guanzhou Chen, Hao Tian, Jiasheng Zhou, Jiashuo Yu, Jifeng Dai, Junjun He, Lewei Lu, Licheng Wen, Limin Wang, Min Dou, Pei Chu, Pinlong Cai, Qingyun Li, Shenglong Ye, Tong Lu, Wei Li, Weiyun Wang, Wenhai Wang, Wenjian Zhang, Xiangchao Yan, Xingjian Wei, Xizhou Zhu, Yali Wang, Yinan He, Yi Wang, Yu Qiao, Yushi Chen, Zhangwei Gao, Zhe Chen, Zhenjiang Jin, Zhenxiang Li, Zhongying Tu","submitted_at":"2024-06-12T17:01:04Z","abstract_excerpt":"Image-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resembles human reading habits. Recent studies have shown that such data aids multimodal in-context learning and maintains the capabilities of large language models during multimodal fine-tuning. However, the limited scale and diversity of current image-text interleaved data restrict the development of multimodal large language models. In this paper, we introduce OmniCorpus, a 10 billion-scale image-text interleaved datas"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2406.08418","kind":"arxiv","version":3},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CV","submitted_at":"2024-06-12T17:01:04Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"41ab20bd5c39420366cd08e368bcbc916f4652800f7bc52d338af40485856079","abstract_canon_sha256":"18c95ab3aecfaba62dfa2ec61afa4d9118b8f405c54321ef68258f72e7017bb8"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:43:10.961064Z","signature_b64":"fWmRzS9OPc3HSDipDd2Xjozs6WXwXxLeboBdI8SkwtNRLbKSDKohDUe+sBvYQRHvqPQI9jeBtzhxQw0tOO+bBA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"60a5c78735f9150d5e71ada83d868efe3effa1627f37b632889e127301a82a76","last_reissued_at":"2026-07-05T08:43:10.960590Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:43:10.960590Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CV","authors_text":"Bin Wang, Botian Shi, Bo Zhang, Changyao Tian, Chao Xu, Conghui He, Dahua Lin, Erfei Cui, Guanzhou Chen, Hao Tian, Jiasheng Zhou, Jiashuo Yu, Jifeng Dai, Junjun He, Lewei Lu, Licheng Wen, Limin Wang, Min Dou, Pei Chu, Pinlong Cai, Qingyun Li, Shenglong Ye, Tong Lu, Wei Li, Weiyun Wang, Wenhai Wang, Wenjian Zhang, Xiangchao Yan, Xingjian Wei, Xizhou Zhu, Yali Wang, Yinan He, Yi Wang, Yu Qiao, Yushi Chen, Zhangwei Gao, Zhe Chen, Zhenjiang Jin, Zhenxiang Li, Zhongying Tu","submitted_at":"2024-06-12T17:01:04Z","abstract_excerpt":"Image-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resembles human reading habits. Recent studies have shown that such data aids multimodal in-context learning and maintains the capabilities of large language models during multimodal fine-tuning. However, the limited scale and diversity of current image-text interleaved data restrict the development of multimodal large language models. In this paper, we introduce OmniCorpus, a 10 billion-scale image-text interleaved datas"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2406.08418","kind":"arxiv","version":3},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2406.08418/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2406.08418","created_at":"2026-07-05T08:43:10.960648+00:00"},{"alias_kind":"arxiv_version","alias_value":"2406.08418v3","created_at":"2026-07-05T08:43:10.960648+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2406.08418","created_at":"2026-07-05T08:43:10.960648+00:00"},{"alias_kind":"pith_short_12","alias_value":"MCS4PBZV7EKQ","created_at":"2026-07-05T08:43:10.960648+00:00"},{"alias_kind":"pith_short_16","alias_value":"MCS4PBZV7EKQ2XTR","created_at":"2026-07-05T08:43:10.960648+00:00"},{"alias_kind":"pith_short_8","alias_value":"MCS4PBZV","created_at":"2026-07-05T08:43:10.960648+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":13,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.17030","citing_title":"Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation","ref_index":31,"is_internal_anchor":false},{"citing_arxiv_id":"2606.11188","citing_title":"ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations","ref_index":43,"is_internal_anchor":false},{"citing_arxiv_id":"2606.28551","citing_title":"DataComp-VLM: Improved Open Datasets for Vision-Language Models","ref_index":161,"is_internal_anchor":false},{"citing_arxiv_id":"2606.28551","citing_title":"DataComp-VLM: Improved Open Datasets for Vision-Language Models","ref_index":161,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22344","citing_title":"Bernini: Latent Semantic Planning for Video Diffusion","ref_index":38,"is_internal_anchor":false},{"citing_arxiv_id":"2501.12386","citing_title":"InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2411.10442","citing_title":"Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization","ref_index":49,"is_internal_anchor":false},{"citing_arxiv_id":"2509.18154","citing_title":"MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12305","citing_title":"Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2604.08644","citing_title":"EXAONE 4.5 Technical Report","ref_index":28,"is_internal_anchor":false},{"citing_arxiv_id":"2507.01006","citing_title":"GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2505.14683","citing_title":"Emerging Properties in Unified Multimodal Pretraining","ref_index":39,"is_internal_anchor":false},{"citing_arxiv_id":"2412.05271","citing_title":"Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling","ref_index":133,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/MCS4PBZV7EKQ2XTRVWUD3BUO7Y","json":"https://pith.science/pith/MCS4PBZV7EKQ2XTRVWUD3BUO7Y.json","graph_json":"https://pith.science/api/pith-number/MCS4PBZV7EKQ2XTRVWUD3BUO7Y/graph.json","events_json":"https://pith.science/api/pith-number/MCS4PBZV7EKQ2XTRVWUD3BUO7Y/events.json","paper":"https://pith.science/paper/MCS4PBZV"},"agent_actions":{"view_html":"https://pith.science/pith/MCS4PBZV7EKQ2XTRVWUD3BUO7Y","download_json":"https://pith.science/pith/MCS4PBZV7EKQ2XTRVWUD3BUO7Y.json","view_paper":"https://pith.science/paper/MCS4PBZV","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2406.08418&json=true","fetch_graph":"https://pith.science/api/pith-number/MCS4PBZV7EKQ2XTRVWUD3BUO7Y/graph.json","fetch_events":"https://pith.science/api/pith-number/MCS4PBZV7EKQ2XTRVWUD3BUO7Y/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/MCS4PBZV7EKQ2XTRVWUD3BUO7Y/action/timestamp_anchor","attest_storage":"https://pith.science/pith/MCS4PBZV7EKQ2XTRVWUD3BUO7Y/action/storage_attestation","attest_author":"https://pith.science/pith/MCS4PBZV7EKQ2XTRVWUD3BUO7Y/action/author_attestation","sign_citation":"https://pith.science/pith/MCS4PBZV7EKQ2XTRVWUD3BUO7Y/action/citation_signature","submit_replication":"https://pith.science/pith/MCS4PBZV7EKQ2XTRVWUD3BUO7Y/action/replication_record"}},"created_at":"2026-07-05T08:43:10.960648+00:00","updated_at":"2026-07-05T08:43:10.960648+00:00"}