Pith. sign in

Paper Citation Record · LEDGER

CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2406.10462.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.10462 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:43:54.222341Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T20:55:04.590393Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation df1f0d85-c592-480c-b6ec-472cdc934e69 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.638404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:153211b3338bbf2535a47f387116b9c14af65087a94965ff2ae4a170067060b4

Observation 5ae716e0-f550-4984-a7e7-c66ce3d745bf · inbound

CI-VID: A Coherent Interleaved Text-Video Dataset cites this paper.

CI-VID: A Coherent Interleaved Text-Video Dataset CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:54.222341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:43:54.222341Z digest=sha256:4ed07bcac3da26b5474d241b02010bdcf0460649536c7f0de43727ea60737b0b

Observation 10ef18b9-ce55-42a5-9e42-06e8c46ad0dc · inbound

COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts cites this paper.

COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:41:26.760116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T09:47:27.891288Z digest=sha256:7f0ca1e0272d988bf2235235ec6f484e70544d5cdc3bcfa61a4e5e8b7d04d84a

Observation 916736f9-cbb1-4c1c-bd0e-abf0a03e2611 · inbound

COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts cites this paper.

COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:17:59.275253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:15:30.952247Z digest=sha256:81730928ddc13171a1741dc28bdc69aa66d51541ddd3c77a6ef9bcf01637e6d4

Observation c47cd7de-b9ad-4a15-9df7-61f39115f6b1 · inbound

Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA cites this paper.

Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:55:04.591960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T20:46:54.667418Z digest=sha256:42563818fd94eb8f598aa43286aeebbe4c682d6fd67907edd14b22fd5d963e7d