Pith. sign in

Paper Citation Record · LEDGER

CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2311.03354.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.03354 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:07:33.048750Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.091209Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ab28b99f-416d-4fc3-9526-e3528ecc82bd · inbound

3D-VLA: A 3D Vision-Language-Action Generative World Model cites this paper.

3D-VLA: A 3D Vision-Language-Action Generative World Model CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.278672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:076de1dd548b40267de87f9a2eb80d2f7eeebd7014367bb7cea1bb5d76fc79b3

Observation bb030ae3-1284-4822-a24a-017b5584ad1a · inbound

EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM cites this paper.

EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T22:07:33.048750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:07:33.048750Z digest=sha256:a4d8fa2f1ad5e2236097e61e32c35345c4531f98c0965f17c401b6d8f04a51a1

Observation cac99366-4373-490d-94e6-68415c833b96 · inbound

Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor cites this paper.

Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T20:24:58.539838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:24:58.539838Z digest=sha256:0080b856508eecdcdcb9b0a32ad7a74d85c2b8499362d4bc10b0741778f4d15e

Observation 70171532-8b38-4b87-800a-21addb893a3e · inbound

Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models cites this paper.

Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T18:17:06.194020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:17:06.194020Z digest=sha256:8dc4c255b2c584ef9664a80c3c3da6573370cfa5f03d590b15561922ca5427c2

Observation 07dd34a8-6821-4dd7-bb14-7233e708c526 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.093837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:186e47c881e8238ee2bb1fd738c2a200401e54de6b1d6c90ead6aeae6d23ecc4