Pith. sign in

Paper Citation Record · LEDGER

OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2406.08418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.08418 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:41:33.517914Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:18:43.867707Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 824f55b6-504a-4d89-aa60-d90613f33225 · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:16:17.260929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:852f13f28d239d92df35fa2b94ab81f0aa5df0467a12671707ae5a5452e8370a

Observation 6f2357ce-2b2b-4ee5-b88d-a80f4916dbc8 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 133

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.160664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:319e1ef4914863dffe0a6db13ff2c3ebcbe099880a8e09bb2d0cf25b02c745a2

Observation ac0ba78b-5f72-4496-89f4-a62cafe8728f · inbound

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining cites this paper.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.517914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.517914Z digest=sha256:0311646f610bc395db35d516be3c85c88c530d504a486008184a84f5d943fdbf

Observation f1b88b23-ddb9-408e-bb7e-bfcb08664a1d · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.735829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:68402f47c040ca386e57541f4cd1a45504483b5900ed4f2a5947e133bb65839a

Observation 85ae94bf-1568-48f6-8ec1-a7df8729faeb · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 221

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:34.989102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:34.989102Z digest=sha256:8e4a389eccf869bd2c5756f754eb05d606af0f48d43a1548db6e88357fdb3266

Observation e957f681-90ad-44fd-b6e4-8a60d50f561c · inbound

Emerging Properties in Unified Multimodal Pretraining cites this paper.

Emerging Properties in Unified Multimodal Pretraining OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:23:42.093546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T16:23:41.854132Z digest=sha256:d6b5d205b19360cbeeecb63f445e92fd042b0577fcd038762dc0eaeb6da0de4f

Observation 43d12450-2c33-4b38-841a-93432052401f · inbound

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning cites this paper.

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:48:26.861808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T04:48:26.355351Z digest=sha256:831b2dd1c87e16adbbdde0b555272f83e9784486b1ddc0f8cdb0d7c7c68682af

Observation c852b742-35d6-41e3-8e07-17ec85210c5f · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:02.050118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:02.050118Z digest=sha256:5d9b064817d7d4720efedab0325e22514de194fb0f15d3644a4d2128091eca89

Observation 34bf0cf9-76bf-47c6-a541-abe3413a2711 · inbound

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning cites this paper.

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T00:59:52.075206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:59:52.075206Z digest=sha256:6baa57ce59536b7bfde82e298eafb8a17bc84fc3168ac6835eaa4dc9ff4ff852

Observation ab0f426d-e632-49d3-9d1c-f60c320fcc9e · inbound

MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe cites this paper.

MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:07:27.436302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T17:07:27.277040Z digest=sha256:473b23b90a46fd9e896b6021c7cba9d92125e26aa88c1dcb463101e504967788

Observation 17a19361-0ea5-4e8b-9ff8-16d408649dd5 · inbound

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models cites this paper.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:25.305850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:25.305850Z digest=sha256:3e02e8a9c36ed6ad073f09228714b46b7286cf06f9a6b2c82fdc1838b2ccdde2

Observation fe5a3143-1304-4a6d-8fd8-71012da7726d · inbound

EXAONE 4.5 Technical Report cites this paper.

EXAONE 4.5 Technical Report OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:10:57.656341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T17:47:34.692414Z digest=sha256:e8aca6272fe3bc324e2bc1cb2ed91c4f627b9964a150e618b66ddb4ea9d20819

Observation ed4986ff-549e-4aee-8c4a-8a21cccce593 · inbound

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation cites this paper.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.782843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:52384fe44ab57291fae4b33f08baa992848cf4f26f7f23f6e71ebf8ab0e80bfd

Observation 83bdc400-7f10-406d-b2a6-a5c5268b0e7f · inbound

Bernini: Latent Semantic Planning for Video Diffusion cites this paper.

Bernini: Latent Semantic Planning for Video Diffusion OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:41:10.547052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T06:39:47.124605Z digest=sha256:92bc496ac8e936fbfc1970aae5406a4b0187511fb71e22d2cef9845e53060641

Observation 7e73a6c3-6dba-435c-8b6c-48fc1e285a9e · inbound

ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations cites this paper.

ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:38.515381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T13:29:11.526106Z digest=sha256:cc49aaede6327041016f6fead58ef28cfbccd4f7b2ff14e84aa4e5dcbd6579fb

Observation 5f6ca5a9-977d-4d38-a4c9-93d2453d2ecf · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:43.868974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:fef104905bf47de514a24678738f8b1143c4e8076b169388077500b7a7b3ee11

Observation 9204d090-fde8-4020-8c14-a098b4082c56 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 161

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:47.667104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:ea2c1e7387c544df1009508d417609cf9a0486a68c1e79736ac463848115239a

Observation be8464c3-3919-4ec1-a51e-1d77cc30a5a5 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 161

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:23.935699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:c009887e7bf15939bfd68c6c0fb8a7675831600af59ed58bb0edd8feaa86982c

Observation c0cf02e3-bc76-4a1f-8f75-b1509b5be0f5 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 300

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:a69a264b7876b0091962c3b0ccb2c291f323378d0643c3109bacb8e39f300fcb

Observation 650ce7e9-078c-454c-9954-ca660d138e27 · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 233

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:363cbe008277369ff0eecfd2074e3c8723d382366f8a4f135bb9e2fbff276f5e