Pith. sign in

Paper Citation Record · LEDGER

DOCCI: Descriptions of Connected and Contrasting Images

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2404.19753.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.19753 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:23:44.221338Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T18:23:50.307986Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 983ceef5-8c25-47a7-a7bc-aaef2c9317c1 · inbound

FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity cites this paper.

FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity DOCCI: Descriptions of Connected and Contrasting Images

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:23:44.221338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:23:44.221338Z digest=sha256:f292c2d9f07af94211e52110537b4cec82c5d312e9dee38c0ae1ca1ed26b0c40

Observation 9656d041-4de3-403c-8078-a99f12a9ea3b · inbound

FLAIR: VLM with Fine-grained Language-informed Image Representations cites this paper.

FLAIR: VLM with Fine-grained Language-informed Image Representations DOCCI: Descriptions of Connected and Contrasting Images

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:06.780939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:06.780939Z digest=sha256:cac0687820cdaf839c1daafb42be38970c627ef1cce089cd514c7e907a7c1103

Observation 18b39670-52ec-483d-a12f-b4458503fee8 · inbound

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent cites this paper.

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent DOCCI: Descriptions of Connected and Contrasting Images

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:24:45.480790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:24:45.480790Z digest=sha256:d7f98157466f14f7685e0ce11901e2f91a3f6d6543a22fa27673ed160c00dd07

Observation 52ba97fe-2ef0-4242-ad23-3068fef9803c · inbound

Learning Visual Composition through Improved Semantic Guidance cites this paper.

Learning Visual Composition through Improved Semantic Guidance DOCCI: Descriptions of Connected and Contrasting Images

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:31:57.518309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:31:57.518309Z digest=sha256:310ff868c162e073112bf24259ec7385360734c31952fb99919e929d4f52b7fc

Observation 796b09f0-3f6d-4d33-ad5e-6fdca5defe66 · inbound

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage cites this paper.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage DOCCI: Descriptions of Connected and Contrasting Images

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.718006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.718006Z digest=sha256:25558c91adfcc7af804336aac8d43fdddd9e50cf24bc509c804cd97d340aa967

Observation 599b19be-1f61-4920-8a80-20f3806464dd · inbound

SketchFlex: Facilitating Spatial-Semantic Coherence in Text-to-Image Generation with Region-Based Sketches cites this paper.

SketchFlex: Facilitating Spatial-Semantic Coherence in Text-to-Image Generation with Region-Based Sketches DOCCI: Descriptions of Connected and Contrasting Images

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T12:22:43.261937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:22:43.261937Z digest=sha256:4c5b0b46ab0f43b0b69fbb8ccb4da31feaefdd23d1370e88f415d3006ed10c60

Observation 2a229a72-0d73-45f5-876a-07f8928da016 · inbound

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions cites this paper.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions DOCCI: Descriptions of Connected and Contrasting Images

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:55.647399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:55.647399Z digest=sha256:b6ad13890923070906e9576475bfb9889e3e85809ca3c7fe38f66c5465c22a30

Observation c076fba8-d99e-4b99-a298-f412ffb53052 · inbound

OmniGen2: Towards Instruction-Aligned Multimodal Generation cites this paper.

OmniGen2: Towards Instruction-Aligned Multimodal Generation DOCCI: Descriptions of Connected and Contrasting Images

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.787801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:95a0611e2f7faabcce94be83e4d08b55a6fd9440240cd77752ba26efdb3c9e73

Observation c008511e-efbf-4603-8c3e-76703c2a874e · inbound

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World cites this paper.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World DOCCI: Descriptions of Connected and Contrasting Images

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.623070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.623070Z digest=sha256:8de37f8bf4d2d26344b1ba68858e4a45c321c451c0ffde5a868e21373a7fb97a

Observation c1b18570-5a77-4e14-9520-86630e70cc47 · inbound

ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation cites this paper.

ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation DOCCI: Descriptions of Connected and Contrasting Images

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T05:59:15.559791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:59:15.559791Z digest=sha256:c951e29ed55287b2555fba2f80667b430f8e4451d1fecac41f1c5368ecbc4034

Observation 26230f63-8435-458b-aceb-cb054c5dfbb1 · inbound

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning cites this paper.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning DOCCI: Descriptions of Connected and Contrasting Images

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:37.077028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:37.077028Z digest=sha256:0215a88ec4eaadb7012c76bb611bc83275dc3aba068e3cd36dcc69be09412e7d

Observation bd7421d5-a9c5-4f20-a85d-f586c9bf2b99 · inbound

DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions cites this paper.

DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions DOCCI: Descriptions of Connected and Contrasting Images

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:10:52.914892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:39:29.592129Z digest=sha256:161a426cf9a9cab31dc078006ed2b2c13bdeb07ad28fba081606506d9028ea3f

Observation 1a017bf0-f698-41c8-adbd-dedb9824c434 · inbound

ReflectCAP: Detailed Image Captioning with Reflective Memory cites this paper.

ReflectCAP: Detailed Image Captioning with Reflective Memory DOCCI: Descriptions of Connected and Contrasting Images

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:11:02.350789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T16:11:10.598413Z digest=sha256:42cd69e0137c70672b92179f988d76e3a34a0ddcc2af7a2ae80b67697ae178f7

Observation 3269f5d8-5b06-4512-80fc-667a87290235 · inbound

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini cites this paper.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini DOCCI: Descriptions of Connected and Contrasting Images

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.309525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:d5ae0ab2ccdc40389e462e049d5d0130626c9028e3e99cd4cf0a2a64a664a059