Pith. sign in

Paper Citation Record · LEDGER

Masked Vision and Language Modeling for Multi-modal Representation Learning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2208.02131.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2208.02131 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:17:13.284220Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:28:31.488306Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fc7a416a-d0d6-4b29-8305-a111b9e8910b · inbound

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration cites this paper.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:18:51.664930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:d8f61cc919e51f6245569326bd4d4063042ebdad2f61ffc51afbfde619e04488

Observation 55cf0588-1a63-43f3-a548-c69e5be1de70 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.402341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:a2714248196c7da66a7800f0f4ff4efc13f5ca3cc40d1ec0943a2ea1b320c4cd

Observation 3da3e4b0-70a4-4e6e-bbe0-64ac22cf5410 · inbound

Adapting Vision-Language Models Without Labels: A Comprehensive Survey cites this paper.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 299

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:13.284220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:13.284220Z digest=sha256:e1951d6d0c8293eadbee465a1b7e54276e4d59b4ba455df16e8d52a27a90164a

Observation e98e1c93-59ce-4345-9e37-677a8259489d · inbound

Neural Scene Designer: Self-Styled Semantic Image Manipulation cites this paper.

Neural Scene Designer: Self-Styled Semantic Image Manipulation Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T12:39:02.069302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:39:02.069302Z digest=sha256:e9861e4cd4f7f84c414f189c9a8b151eaec3838684bbb5a04c27bd2f83f66b23

Observation 3da34da8-9c71-4841-a503-df1e2579246c · inbound

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning cites this paper.

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:52:59.957661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T16:51:48.705876Z digest=sha256:b975ba3125f39ecae281ce2fcafeac44dd8c81f4ed2e5a44a0a69b3760778ec1

Observation 684266cc-ee7f-4249-a730-2d398050ffb8 · inbound

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning cites this paper.

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T12:10:53.720348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:10:53.720348Z digest=sha256:59893242d52cc22ddc6dfe053baa43c196a5105a6f095ca032c7107f1936b31f

Observation d34a72d1-1c9a-4974-8f61-2a507d245894 · inbound

Personalized Cross-Modal Emotional Correlation Learning for Speech-Preserving Facial Expression Manipulation cites this paper.

Personalized Cross-Modal Emotional Correlation Learning for Speech-Preserving Facial Expression Manipulation Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:26:13.747190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T17:02:44.776943Z digest=sha256:45f752292c8185d9ed125266d41de075c4f1004d3e21455d1d07f36c6f835081

Observation 9b3109b4-55fc-4ed7-a713-e8c580ed799e · inbound

Rethinking Electro-Optical Vision Foundation Models for Remote Sensing Retrieval: A Controlled Comparison with Generalist VFM cites this paper.

Rethinking Electro-Optical Vision Foundation Models for Remote Sensing Retrieval: A Controlled Comparison with Generalist VFM Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:08.649653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T15:59:01.512516Z digest=sha256:7ec54c3cf97a40cc494dd2e4951358148edde15ebf5ec382b30678b538777f8a

Observation 55a71721-1a86-43a8-a3df-9732bfb242fe · inbound

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality cites this paper.

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:28:31.489624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T07:03:50.311891Z digest=sha256:cd6d25ac84284744d3a4515edfcfe6b7bcf5abc49cb3cb644dd88ecd38f85b62