Pith. sign in

Paper Citation Record · LEDGER

Masked Vision and Language Modeling for Multi-modal Representation Learning

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2208.02131.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2208.02131 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:32:54.859675Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:28:31.488306Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fc7a416a-d0d6-4b29-8305-a111b9e8910b · inbound

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration cites this paper.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:18:51.664930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:3ea9f67127e80c5ac3e90f447542a8065c5d95fd8984b2ce98dc0f0516da6bfe

Observation 55cf0588-1a63-43f3-a548-c69e5be1de70 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.402341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:88870ef6e3ffaedd0fc06b44e28cc9efaf63702081e9bb800d6916fd8bf9beff

Observation 1d92d986-5a26-4512-af1c-a75b8062d3db · inbound

Visual question answering: from early developments to recent advances -- a survey cites this paper.

Visual question answering: from early developments to recent advances -- a survey Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T21:46:28.276131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:46:28.276131Z digest=sha256:6a5e58ee7b2502f089d60375d357abc58fef2dd3f36d95cef7c0fd27b50b6226

Observation 0516b83b-13fc-48ca-ae1a-f4ac452978f4 · inbound

Multimodal Large Language Models for Medicine: A Comprehensive Survey cites this paper.

Multimodal Large Language Models for Medicine: A Comprehensive Survey Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 236

Resolution
unresolved
no resolver link, observed 2026-08-16T05:32:54.859675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:32:54.859675Z digest=sha256:ff3d3c139590da83b1f5a19e92f579a8929173bf9b3fc6eca16a46239cb078da

Observation 63114a85-0dcd-4e80-be63-30d9cd94b3b0 · inbound

Investigating the Effect of Parallel Data in the Cross-Lingual Transfer for Vision-Language Encoders cites this paper.

Investigating the Effect of Parallel Data in the Cross-Lingual Transfer for Vision-Language Encoders Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:34.049294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:34.049294Z digest=sha256:532256d00fbb2322407733c1c0aa16f571ce4216df72b0aee37bdc6f5975085a

Observation 3da3e4b0-70a4-4e6e-bbe0-64ac22cf5410 · inbound

Adapting Vision-Language Models Without Labels: A Comprehensive Survey cites this paper.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 299

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:13.284220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:13.284220Z digest=sha256:d0d0db34abef3e8c701749e3e577fa0f0e20dc0602eacc200b555d7de79ce9f2

Observation e98e1c93-59ce-4345-9e37-677a8259489d · inbound

Neural Scene Designer: Self-Styled Semantic Image Manipulation cites this paper.

Neural Scene Designer: Self-Styled Semantic Image Manipulation Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T12:39:02.069302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:39:02.069302Z digest=sha256:76dc0414557ee67edfac3b7e52705b32071fe26de0baf0b8b1235bdd1fdd733d

Observation 3da34da8-9c71-4841-a503-df1e2579246c · inbound

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning cites this paper.

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:52:59.957661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T16:51:48.705876Z digest=sha256:f26c97f9e3f1625b233001cd6fff52dcfaed00ec3a6ed31b221452bebd0a31e4

Observation 684266cc-ee7f-4249-a730-2d398050ffb8 · inbound

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning cites this paper.

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T12:10:53.720348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:10:53.720348Z digest=sha256:9e2d16f6b7f95496f1699e766813c319966fe92e2ba36a5f68305f6cab7c5950

Observation d34a72d1-1c9a-4974-8f61-2a507d245894 · inbound

Personalized Cross-Modal Emotional Correlation Learning for Speech-Preserving Facial Expression Manipulation cites this paper.

Personalized Cross-Modal Emotional Correlation Learning for Speech-Preserving Facial Expression Manipulation Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:26:13.747190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T17:02:44.776943Z digest=sha256:025dfbaef6bfce2d6354bb39a20e8d5c3670a49ce80e8f69f802e0fb6684e5cd

Observation 9b3109b4-55fc-4ed7-a713-e8c580ed799e · inbound

Rethinking Electro-Optical Vision Foundation Models for Remote Sensing Retrieval: A Controlled Comparison with Generalist VFM cites this paper.

Rethinking Electro-Optical Vision Foundation Models for Remote Sensing Retrieval: A Controlled Comparison with Generalist VFM Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:08.649653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T15:59:01.512516Z digest=sha256:b06c695be5b4e30d5146367749f604e8d9a89de6fbbc26d6db3656d03f28f9aa

Observation 55a71721-1a86-43a8-a3df-9732bfb242fe · inbound

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality cites this paper.

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:28:31.489624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T07:03:50.311891Z digest=sha256:6b4c642c75db79813e608dbc08147e229deb19c3136160230af39d8c55a3c9a4