Pith. sign in

Paper Citation Record · LEDGER

mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2205.12005.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2205.12005 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T01:34:38.597959Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:29:56.632867Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fe4fe666-abea-4baf-a398-efc8c3d7cea8 · inbound

GIT: A Generative Image-to-text Transformer for Vision and Language cites this paper.

GIT: A Generative Image-to-text Transformer for Vision and Language mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:54:07.669979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T20:54:07.572136Z digest=sha256:3224d81454e6fdcaee0482da6496e84e07f3700ce88a1ea1e838a32f4323df47

Observation cbc7e6ea-4b69-49b7-9690-b57795862f74 · inbound

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models cites this paper.

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:52:01.283540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T01:52:01.163900Z digest=sha256:291cfdc8564850b359ae5393421e4f0f595ddbd7cd2a7a58da4c74f948eee7dd

Observation 0547117b-2545-43d3-8918-a74655c7f44c · inbound

ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment cites this paper.

ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:43:03.449820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T19:43:03.310237Z digest=sha256:5bc759d579bb515aaafa92dff9cc60d8698165b7a03b8eb8c74543fc9fd24558

Observation a22d27b4-0d92-4ae2-a3e9-4c5822de9a63 · inbound

The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding cites this paper.

The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:25:28.425941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T07:24:01.527093Z digest=sha256:d278c3b1b542f50726be38d62ab6bc294996acc3f1d2dadacd076f2f00919f47

Observation 23b3c1fa-12d1-4479-8187-630a5c314b33 · inbound

HyperCap: Hyperspectral Land Cover Captioning Dataset for Vision Language Models cites this paper.

HyperCap: Hyperspectral Land Cover Captioning Dataset for Vision Language Models mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:16:43.715851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T15:15:19.523035Z digest=sha256:7ff8608bd4c2fbea5143267d67b0ec033607083c0b3cf52508bffcbbe869ec1b

Observation c7014bc0-822a-4cee-8cad-1eca19200620 · inbound

StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback cites this paper.

StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:55.380981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:26:42.034972Z digest=sha256:4effa8aac953f39adc9c950cdec45707d2ab3d5d18776dc4716ad8076bcfbfd1

Observation eeb09fb3-51ba-4736-a08f-288dded6a40d · inbound

AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning cites this paper.

AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:40:19.818843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:46:22.891034Z digest=sha256:0aa0e93f50e622eea0eb7f20d0906f870753c49cf9af07a839c8f44d385df148

Observation 55733916-8a99-4c15-871f-0e32c9a9a304 · inbound

VisChronos: Revolutionizing Image Captioning Through Real-Life Events cites this paper.

VisChronos: Revolutionizing Image Captioning Through Real-Life Events mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:29:56.634915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T01:34:38.597959Z digest=sha256:2212e5332f137a017e7209387b4ce95b374b35bfbc4eb0827ab7f738aa297380