Pith. sign in

Paper Citation Record · LEDGER

Medical Vision Language Pretraining: A survey

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2312.06224.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.06224 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:08.521199Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T00:57:41.989125Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5dfeaaf2-9e48-4498-bfb7-b944696c7e31 · inbound

Multimodal Federated Learning With Missing Modalities through Feature Imputation Network cites this paper.

Multimodal Federated Learning With Missing Modalities through Feature Imputation Network Medical Vision Language Pretraining: A survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:08.521199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:08.521199Z digest=sha256:0548cadd2ac90f661ae182ab3215be0aad7d844f1ab62db1ac8c81fef127c70e

Observation cf92113c-2d18-48dd-ab23-d108399a51d8 · inbound

Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models cites this paper.

Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models Medical Vision Language Pretraining: A survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:20.981003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:02:20.981003Z digest=sha256:ac01fed8b56e1e175df7f6c3fad268ec4b697862c19cfbef5b7711c090a705cf

Observation 07a65afb-90f3-4dce-8907-9151fa3d4814 · inbound

Knowledge to Sight: Reasoning over Visual Attributes via Knowledge Decomposition for Abnormality Grounding cites this paper.

Knowledge to Sight: Reasoning over Visual Attributes via Knowledge Decomposition for Abnormality Grounding Medical Vision Language Pretraining: A survey

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T23:56:50.028792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:56:50.028792Z digest=sha256:feb640ef69e813dca94ba4b2e0e92a69dc2d9b6288616e4e60f586bd8f3aed3e

Observation e033c3f5-f993-4dad-bd62-8a7b42b71c2d · inbound

Estimating 2D Keypoints of Surgical Tools Using Vision-Language Models with Low-Rank Adaptation cites this paper.

Estimating 2D Keypoints of Surgical Tools Using Vision-Language Models with Low-Rank Adaptation Medical Vision Language Pretraining: A survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T14:49:39.680842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:49:39.680842Z digest=sha256:7ce541e2283435664ff730aee343e090a07b88c8be88a9d0a6eb51c0e7710b70

Observation b03d14c2-1e7d-4718-aa44-65820eaae2a0 · inbound

6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models cites this paper.

6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models Medical Vision Language Pretraining: A survey

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T18:40:37.173198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:40:37.173198Z digest=sha256:6dc221919f3ecced9e7f3654364b51b81aaa91c23ceb571b05d714603321c87c

Observation 30a4be07-fa28-404d-96b5-88fbe8b79077 · inbound

ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities cites this paper.

ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities Medical Vision Language Pretraining: A survey

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-07-11T00:57:42.010550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T00:54:18.897346Z digest=sha256:d5a22d7721a2f1fcc7800eae4dbd7dbbe41805d4e802769ca8b31fc6fe0b81e9