Pith. sign in

Paper Citation Record · LEDGER

DOCCI: Descriptions of Connected and Contrasting Images

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2404.19753.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.19753 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:22:43.261937Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T18:23:50.307986Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 599b19be-1f61-4920-8a80-20f3806464dd · inbound

SketchFlex: Facilitating Spatial-Semantic Coherence in Text-to-Image Generation with Region-Based Sketches cites this paper.

SketchFlex: Facilitating Spatial-Semantic Coherence in Text-to-Image Generation with Region-Based Sketches DOCCI: Descriptions of Connected and Contrasting Images

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T12:22:43.261937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:22:43.261937Z digest=sha256:9f9ccccc2aeec758bfcf63fe19e9f4d9dea8ccb4c511ba6fe3a85b1fd9fbf100

Observation 2a229a72-0d73-45f5-876a-07f8928da016 · inbound

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions cites this paper.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions DOCCI: Descriptions of Connected and Contrasting Images

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:55.647399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:55.647399Z digest=sha256:f334737d0b96504f73229fc0f2cd8ae79cff3d0647b0576e7859e7caf94159ce

Observation c076fba8-d99e-4b99-a298-f412ffb53052 · inbound

OmniGen2: Towards Instruction-Aligned Multimodal Generation cites this paper.

OmniGen2: Towards Instruction-Aligned Multimodal Generation DOCCI: Descriptions of Connected and Contrasting Images

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.787801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:928f397acb1046013bed7ed93fd01e4e7e844852e75036d87109089db13e46a0

Observation c008511e-efbf-4603-8c3e-76703c2a874e · inbound

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World cites this paper.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World DOCCI: Descriptions of Connected and Contrasting Images

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.623070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.623070Z digest=sha256:8e7906124bac56fb0f196ae8805148822d4fcde70b5373ed33eaf79bd0b5ce16

Observation c1b18570-5a77-4e14-9520-86630e70cc47 · inbound

ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation cites this paper.

ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation DOCCI: Descriptions of Connected and Contrasting Images

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T05:59:15.559791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:59:15.559791Z digest=sha256:3205f87dd1e69490262ea268f4ed8a50da2d6db62f6c3e824ae04c3cbf751e16

Observation 26230f63-8435-458b-aceb-cb054c5dfbb1 · inbound

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning cites this paper.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning DOCCI: Descriptions of Connected and Contrasting Images

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:37.077028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:37.077028Z digest=sha256:4ab42a06cdcb2d29b90f36e24c2d3bf5312cb79d2edfbd220eaeb024920e6924

Observation bd7421d5-a9c5-4f20-a85d-f586c9bf2b99 · inbound

DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions cites this paper.

DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions DOCCI: Descriptions of Connected and Contrasting Images

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:10:52.914892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:39:29.592129Z digest=sha256:0e34f781444d0069351d01410484903cf92416c8e6918d653e7be7d281776ec4

Observation 1a017bf0-f698-41c8-adbd-dedb9824c434 · inbound

ReflectCAP: Detailed Image Captioning with Reflective Memory cites this paper.

ReflectCAP: Detailed Image Captioning with Reflective Memory DOCCI: Descriptions of Connected and Contrasting Images

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:11:02.350789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:11:10.598413Z digest=sha256:ed6dcf604ea238f39a9d10b8adbfd0fcaf2a883881cd323a5368b8cf03267018

Observation 3269f5d8-5b06-4512-80fc-667a87290235 · inbound

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini cites this paper.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini DOCCI: Descriptions of Connected and Contrasting Images

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.309525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:19648755d97009a1e238af149e432a84839ae4683a44891f10fc51ab0d886c8b