Pith. sign in

Paper Citation Record · LEDGER

A Survey of Vision-Language Pre-Trained Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2202.10936.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2202.10936 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:53:58.692319Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:07:55.952015Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9cd3c01f-6c9a-4de7-9a68-3105ac6d19a2 · inbound

Mitigating Object Hallucination via Robust Local Perception Search cites this paper.

Mitigating Object Hallucination via Robust Local Perception Search A Survey of Vision-Language Pre-Trained Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:53:58.692319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:53:58.692319Z digest=sha256:352784bbe1466a5ce4fb9e079a3ed72bfb907954f11fcb0a5610b07c4ae397cd

Observation 13e46ba8-bdfe-4e22-aea4-0dc4a23c42a9 · inbound

MrM: Black-Box Membership Inference Attacks against Multimodal RAG Systems cites this paper.

MrM: Black-Box Membership Inference Attacks against Multimodal RAG Systems A Survey of Vision-Language Pre-Trained Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:16.296555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:16.296555Z digest=sha256:58d6c02f94fbe9bd71156ca96e2643b7344b18da303aa373ea9c7968e441df58

Observation dfbd2932-fb48-49b2-b13d-92c274c9d958 · inbound

CF-VLM:CounterFactual Vision-Language Fine-tuning cites this paper.

CF-VLM:CounterFactual Vision-Language Fine-tuning A Survey of Vision-Language Pre-Trained Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.949603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.949603Z digest=sha256:6c9d6fd76170cc830ca6bbf7741fe1fe5f96aeacc040cb459005605a56acb3f5

Observation e4a9cd1b-e7b5-4309-9357-4c2ee836d135 · inbound

Generalizing vision-language models to novel domains: A comprehensive survey cites this paper.

Generalizing vision-language models to novel domains: A comprehensive survey A Survey of Vision-Language Pre-Trained Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:40.995438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:40.995438Z digest=sha256:24b0bd2ac5701f280ff2bb6358352aaa1ed876b391fb80ae1a661d6c884f2729

Observation ffdce5f9-038f-4fa9-83e1-578385bda8de · inbound

LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants cites this paper.

LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants A Survey of Vision-Language Pre-Trained Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:29:39.359722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:29:39.359722Z digest=sha256:97d2d9d5d87638f667a49e9ebc9441c838591bf935997023e0eb91bf5a8429ac

Observation 186cde67-7b07-4077-919c-ee7175e5e096 · inbound

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models cites this paper.

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models A Survey of Vision-Language Pre-Trained Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:25:54.920538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:25:54.920538Z digest=sha256:e45898107b569387e5b4d8cd2a7446e37766ff5c962736c454760efdc301726b

Observation 902feb0e-ce7e-4d32-bab5-23c0b610e022 · inbound

Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting cites this paper.

Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting A Survey of Vision-Language Pre-Trained Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:44:26.880553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T23:41:54.385864Z digest=sha256:0933d3f04398b0373b6333d7db9d53939e02a8112dee8cf75950d5b1fd403256

Observation 36838e3d-44fc-4ce7-956e-97a6ccccf560 · inbound

Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting cites this paper.

Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting A Survey of Vision-Language Pre-Trained Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:52.046324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:51:52.046324Z digest=sha256:967cc59d2b6ab60ae66f24ec7d89605a2d5427f1b51f67cff1c9a551b2e6bd4b

Observation 3e761e56-ebad-4d17-ac87-ec51b196b33a · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration A Survey of Vision-Language Pre-Trained Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:12:54.098165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T00:12:39.834892Z digest=sha256:b506097dabf46c9b98752859ea0c02c289e354b0f0cd4593a9d5823af742952a

Observation 65e4103a-0287-481d-8498-e29585f1edac · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration A Survey of Vision-Language Pre-Trained Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:05:30.720002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T08:02:15.950975Z digest=sha256:b06b6992e5206804d4e97f566e401093dc8c8d6893d8d47fd9e42ff68144dafe

Observation f339b4fc-c974-4f55-a6c2-54a2ebadcc84 · inbound

Temporal Inversion for Learning Interval Change in Chest X-Rays cites this paper.

Temporal Inversion for Learning Interval Change in Chest X-Rays A Survey of Vision-Language Pre-Trained Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:00:50.514589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:23:46.184189Z digest=sha256:b0d4021449197bc0e658b2924f3d52db4ecf01a6a1bd871ed074039de1a7512a

Observation 24eb3a63-25db-4324-aa8a-eb97406f0233 · inbound

Towards Long-horizon Agentic Multimodal Search cites this paper.

Towards Long-horizon Agentic Multimodal Search A Survey of Vision-Language Pre-Trained Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:01:04.591634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:40:32.137708Z digest=sha256:a4bdcb7fb842aa41104b541beac1a1261ee459269fc898faf40712b32fb1f84f

Observation 3bc4273e-a612-4460-9f37-0f2a9c09dec1 · inbound

Federated Cross-Modal Retrieval with Missing Modalities via Semantic Routing and Adapter Personalization cites this paper.

Federated Cross-Modal Retrieval with Missing Modalities via Semantic Routing and Adapter Personalization A Survey of Vision-Language Pre-Trained Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:06:09.279154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T12:44:55.026879Z digest=sha256:d9eee84c19754e60e131ffa438d6f4ed41c35916a5179b5984ca5c381022d6ef

Observation a666579a-421f-4350-b4ae-ae2ab06de875 · inbound

Structural Ranking of the Cognitive Plausibility of Computational Models of Analogy and Metaphors with the Minimal Cognitive Grid cites this paper.

Structural Ranking of the Cognitive Plausibility of Computational Models of Analogy and Metaphors with the Minimal Cognitive Grid A Survey of Vision-Language Pre-Trained Models

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:56:06.216704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T14:33:11.033906Z digest=sha256:95aa6b1d445dd6cc5e3dd5a560d1a5d25bd5cf575ed8454f9da110056def317e

Observation cd80b9d0-947a-4308-b686-af6fbeebf19a · inbound

Efficient Prompt Learning for Traffic Forecasting cites this paper.

Efficient Prompt Learning for Traffic Forecasting A Survey of Vision-Language Pre-Trained Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:06:26.722342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:20:37.192907Z digest=sha256:e13c415ed214ece8b2e466b0931bcbe19deb2f3b163100e1eebdb97732b6897c

Observation 1ef376ca-e6ba-4686-8f92-8eba6c2e91fd · inbound

Horizontal and Longitudinal Comparisons Among AI Subfields: A Bibliometric Perspective cites this paper.

Horizontal and Longitudinal Comparisons Among AI Subfields: A Bibliometric Perspective A Survey of Vision-Language Pre-Trained Models

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:56:28.056860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:29:18.418975Z digest=sha256:24b73055c83e2b1c3f8089196e65565d77fc92afcf3618e18ef1267f432ae12f

Observation 10f370ce-81f6-4164-bfa4-29f6e9f9f733 · inbound

TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation cites this paper.

TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation A Survey of Vision-Language Pre-Trained Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:07:55.953745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T10:15:59.516498Z digest=sha256:171883cbf1ad0c19ac7debeea24778431ff6dd553a7efe4516cac7fc7aab3ef2

Observation cecabeb0-c43d-4b02-903b-316868b99a22 · inbound

TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation cites this paper.

TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation A Survey of Vision-Language Pre-Trained Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T18:06:38.115031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:06:38.115031Z digest=sha256:6dd52143c30739f54f51938bfb71a1b208c8ce6092abe9dbbe53acb1ed831b61