Pith. sign in

Paper Citation Record · LEDGER

From Pixels to Prose: A Large Dataset of Dense Image Captions

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2406.10328.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.10328 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:09:56.862819Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:23:28.438005Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d15b7e7e-f44d-4dd6-9422-b149b8c24e0a · inbound

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation cites this paper.

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:09:16.246786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T22:09:16.001309Z digest=sha256:b1e01e42d21e4c4831c82f0a7328c2e23e390c7d4301b30ed61b51824849c31c

Observation fdf7a16c-0e59-4fc1-91d5-130f6ccbcf13 · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:09:25.282416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:f24f2ea1566ec4bda885d024ae6258d6c2ebcdc17f8f8f7125da5891a8b748b6

Observation 3e95b839-3773-4b95-99c6-bf0a01e15116 · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:52:01.832523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:f50c2bddb9ce67c79544e0387b819caf58093604229c9570d6236d4ac57f25b6

Observation 8f4a2872-484e-4da9-85af-e0d56c56e405 · inbound

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions cites this paper.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:56.862819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:56.862819Z digest=sha256:ddb6ca40f566e7780b063ce19e77a03010d8c0163c981695edddae1e16ab6abb

Observation 2cc224ce-97fb-4208-a204-b9c5ec103c97 · inbound

Document-Level Text Generation with Minimum Bayes Risk Decoding using Optimal Transport cites this paper.

Document-Level Text Generation with Minimum Bayes Risk Decoding using Optimal Transport From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.258130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.258130Z digest=sha256:dfc7b220364005ef63445ea343f436e54c52f027b700747f777504a977446ff4

Observation 4aaaa547-1faf-4365-b5ba-e27ee147e874 · inbound

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings cites this paper.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.893267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.893267Z digest=sha256:ae2fe6b2ce316ceef36d0c1097db6ee5ea6eccaf34933819b2b0764bc69f96fc

Observation becd85ce-93b1-4790-9592-8da2c45b33fb · inbound

MS-DPPs: Multi-Source Determinantal Point Processes for Contextual Diversity Refinement of Composite Attributes in Text to Image Retrieval cites this paper.

MS-DPPs: Multi-Source Determinantal Point Processes for Contextual Diversity Refinement of Composite Attributes in Text to Image Retrieval From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:06:03.980611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:06:03.980611Z digest=sha256:fd6d481060d8995731f77f5b251b6b52e8cc5e38b4afec823eca11360b1502bd

Observation ce87387f-6774-4174-adf5-76bfa3e64914 · inbound

ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation cites this paper.

ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T05:59:15.603444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:59:15.603444Z digest=sha256:7dcf34c823412295fefc3aaadeb4c24f56156d724f19ced004c0da0bb7c0e9c1

Observation 93220cf9-883b-4f79-8d90-b1cfa87f3d93 · inbound

Transition Models: Rethinking the Generative Learning Objective cites this paper.

Transition Models: Rethinking the Generative Learning Objective From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T10:19:54.462627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:19:54.462627Z digest=sha256:955a3827ea191f4b1cc637f96fc2f7f0236c6110400c463ecb1dfc6da9d9c8a5

Observation a40919a6-6724-4fa7-b093-3f5f9d14a557 · inbound

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark cites this paper.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.803853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.803853Z digest=sha256:4aa5b0635eca543f5b09970f65eee32928e64dcf6d4f054b278fe09efbfcd830

Observation 1e3a1767-f545-4f45-b653-e5e7622b1d96 · inbound

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens cites this paper.

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:08:15.863682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T12:04:19.761430Z digest=sha256:391028db37eae1ffa320dac1266f63cf7593a1365968ca9ae5d549fb042d0a7d

Observation 6fb820e6-d82a-455e-ac0b-fcf50011016b · inbound

BEiTScore: Reference-free Image Captioning Evaluation with an Efficient Cross-Encoder Model cites this paper.

BEiTScore: Reference-free Image Captioning Evaluation with an Efficient Cross-Encoder Model From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:04:45.988241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T09:01:24.453821Z digest=sha256:d05f3dc202bf15f971a878458d9d4c9c1ccac38a9ae730f1d1850b162c932276

Observation 9416758b-e8e1-4ecd-9f1b-7af77abf22f4 · inbound

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning cites this paper.

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.439444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T13:13:57.599970Z digest=sha256:e184ef4ebdb15cc29d6097a88a430599d521b071a2f4c8e93613daa43bd3b327

Observation f2f6443d-f7f4-4dff-bcb1-acd6bcf3d080 · inbound

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget cites this paper.

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-02T06:14:02.725335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:14:02.725335Z digest=sha256:567646a5bc2d229ab94e5f6f8c641ca3f461b623c82c8519fde979f99bc6b9ad

Observation 99b7a9c9-e0ce-4fd2-9a40-4f126a3beb0d · inbound

MIDAL: A Dataset of Math Image Descriptions for Accessible Learning cites this paper.

MIDAL: A Dataset of Math Image Descriptions for Accessible Learning From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T00:49:46.472687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:49:46.472687Z digest=sha256:fbf47597fa243e9ec6935fa229462ea91b249bba4154459d0b66d924e324fc0e