Pith. sign in

Paper Citation Record · LEDGER

CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2311.03354.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.03354 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:07:33.048750Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.091209Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ab28b99f-416d-4fc3-9526-e3528ecc82bd · inbound

3D-VLA: A 3D Vision-Language-Action Generative World Model cites this paper.

3D-VLA: A 3D Vision-Language-Action Generative World Model CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.278672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:9c9dec42bd5922012252151b0495b6860bdae574513541a374dc09bd11b56e9b

Observation bb030ae3-1284-4822-a24a-017b5584ad1a · inbound

EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM cites this paper.

EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T22:07:33.048750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:07:33.048750Z digest=sha256:ba4a66a9ea140bb8ca4fd34cc4931004c66fd74d29c5f39821e49daead488d59

Observation cac99366-4373-490d-94e6-68415c833b96 · inbound

Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor cites this paper.

Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T20:24:58.539838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:24:58.539838Z digest=sha256:14f640fd7f8b46c527245d6fd3592f02da70dd71c9ea505fb55faf599c0fc936

Observation 70171532-8b38-4b87-800a-21addb893a3e · inbound

Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models cites this paper.

Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T18:17:06.194020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:17:06.194020Z digest=sha256:800eb46ce28aec145da820a895c5025cc4c0ab55aeca99f95c605cdf72a457e0

Observation 07dd34a8-6821-4dd7-bb14-7233e708c526 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.093837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:902c3fcaf2f7a7a8dea688cdb3839e2ae64e91e645e2b05b35fc36222b8d748e