Pith. sign in

Paper Citation Record · LEDGER

Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2110.05208.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2110.05208 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:01:38.625976Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T15:47:23.261860Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c3088681-764d-47e2-bfed-00f40ccd0737 · inbound

DetailCLIP: Injecting Image Details into CLIP's Feature Space cites this paper.

DetailCLIP: Injecting Image Details into CLIP's Feature Space Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-24T11:09:22.475201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-24T11:08:20.298043Z digest=sha256:280793bdd4cd6e608cbd46994141a0e311239f37e445c1402d253d95892c6d08

Observation 43973210-4fda-49f0-b1e1-5407148d264f · inbound

Demystifying CLIP Data cites this paper.

Demystifying CLIP Data Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:20:20.309226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-16T09:20:20.143143Z digest=sha256:261b9cac74d057c3852fdf7d6432deeb9ebfd9761f599a4c1d80237cc6dcc44d

Observation 2b4ca0b7-0483-459c-8a3a-09de53af94b3 · inbound

Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey cites this paper.

Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 201

Resolution
verified exact
arxiv_id, observed 2026-05-13T11:32:36.972218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T11:32:36.738536Z digest=sha256:e6574ac82917d0e6e8fcdc912eb05142d5f038ad60b23ea24cca8e698aec3adb

Observation 9066309a-9e9b-43bb-8d46-653d2303d4e6 · inbound

AstroM$^3$: A self-supervised multimodal model for astronomy cites this paper.

AstroM$^3$: A self-supervised multimodal model for astronomy Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T21:23:18.755350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:23:18.755350Z digest=sha256:e13247f7a143efd9159dbc881b6eb52cafd73e9b17ac6bcd20bde5ed4e750224

Observation e3248fd1-ec62-4470-8ee1-387d7052c612 · inbound

HNCSE: Advancing Sentence Embeddings via Hybrid Contrastive Learning with Hard Negatives cites this paper.

HNCSE: Advancing Sentence Embeddings via Hybrid Contrastive Learning with Hard Negatives Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T17:55:23.799740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:55:23.799740Z digest=sha256:fb080b8ffc3471627d52cd552cd5177fe6153022610cd2224177fd3cafbe5a44

Observation 9e9d7ef6-4019-4fdb-ad6a-63b163c52140 · inbound

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers cites this paper.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.135248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.135248Z digest=sha256:8d36f903aeaf340906756c064af376918e1c25c4c4ba8ad143818acfd103bb7d

Observation 5a182fec-8b24-4af0-aba6-5b67af6419af · inbound

OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining cites this paper.

OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:55.429777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:55.429777Z digest=sha256:769c060b0662143ad5f3c3f14dd3bd1548b4181a80baa85012027c7cbf4d1342

Observation 7827851f-bd7d-4a94-8440-0bda523518d5 · inbound

FLAIR: VLM with Fine-grained Language-informed Image Representations cites this paper.

FLAIR: VLM with Fine-grained Language-informed Image Representations Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:06.735220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:06.735220Z digest=sha256:5e2ede75f10546ec1acbcc027242a34b964d749bf79c4d46855fe41be446b8b3

Observation eaed1dd4-93c9-428d-b2f0-f0d9497deb5b · inbound

VladVA: Discriminative Fine-tuning of LVLMs cites this paper.

VladVA: Discriminative Fine-tuning of LVLMs Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:26.418704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:26.418704Z digest=sha256:3adea6f7b33bc01cc86f10a9181bbf6a348a8af9df7da22a1db0f0bf9484f8f9

Observation 55cfc67b-9615-45a6-b1c6-22fd8b29633a · inbound

DiffCLIP: Few-shot Language-driven Multimodal Classifier cites this paper.

DiffCLIP: Few-shot Language-driven Multimodal Classifier Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T19:11:02.387573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:11:02.387573Z digest=sha256:a0521643d1ae137542d4680c8c1203b6873f6ae681bfbe3965b11e7eb9f1d7f3

Observation 36871028-b91e-4da9-a0f0-6ea3e3c58058 · inbound

Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples cites this paper.

Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T16:32:24.360452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:32:24.360452Z digest=sha256:f4756cea90fef53273a2d45dbb5e56457e49382d4dddbd6c5f671a66173db4ba

Observation c8d1c715-6f9e-4bdd-a4f7-801b17483251 · inbound

COEF-VQ: Cost-Efficient Video Quality Understanding through a Cascaded Multimodal LLM Framework cites this paper.

COEF-VQ: Cost-Efficient Video Quality Understanding through a Cascaded Multimodal LLM Framework Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T18:10:44.688935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:10:44.688935Z digest=sha256:7870478ca3aac5953a1f25fa47deb0250974eb17c9ae99bb4b790a94aa139848

Observation 4ff8a9fe-0978-4103-95c2-21f22b981205 · inbound

Sensorformer: Cross-patch attention with global-patch compression is effective for high-dimensional multivariate time series forecasting cites this paper.

Sensorformer: Cross-patch attention with global-patch compression is effective for high-dimensional multivariate time series forecasting Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:10:19.572614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:10:19.572614Z digest=sha256:5861ed3d38d96e6365f98af57a7e0796499401a2345ff172fabb87b75a427cfa

Observation e7fbefa1-7d61-472e-9f7b-ca5749a063c4 · inbound

ProKeR: A Kernel Perspective on Few-Shot Adaptation of Large Vision-Language Models cites this paper.

ProKeR: A Kernel Perspective on Few-Shot Adaptation of Large Vision-Language Models Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T18:40:27.171479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:40:27.171479Z digest=sha256:42dc3cfcde168f214c5011b260bc1488bbd2c7df794f4fd10b843e03d9660017

Observation 1af0c70c-95e4-4bc5-ae4f-476ff02b4339 · inbound

AquaticCLIP: A Vision-Language Foundation Model for Underwater Scene Analysis cites this paper.

AquaticCLIP: A Vision-Language Foundation Model for Underwater Scene Analysis Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T14:32:36.832435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:32:36.832435Z digest=sha256:91b4111d61efbf4de59d241d0e6a4805b210adb82927423c10e6f738fc4130bc

Observation 7eaf461c-ec4e-4ac5-bf04-695ec715f0a5 · inbound

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion cites this paper.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.238464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.238464Z digest=sha256:39a341d926100f9d78c7ccfaecd2c14552d0666bf016d4353455cd08630eb30a

Observation b18e0303-3dcd-4727-b2b8-c843df5c473e · inbound

GeoMM: On Geodesic Perspective for Multi-modal Learning cites this paper.

GeoMM: On Geodesic Perspective for Multi-modal Learning Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.625976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.625976Z digest=sha256:ab944c9e185cdf58941c9890f8a41c4bde3a43bf9b96f41d0f0e1f401a15fbf1

Observation 9cc2e9ac-1017-478e-a544-49dab47ca017 · inbound

Generalizable Vision-Language Few-Shot Adaptation with Predictive Prompts and Negative Learning cites this paper.

Generalizable Vision-Language Few-Shot Adaptation with Predictive Prompts and Negative Learning Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:54.703738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:53:54.703738Z digest=sha256:dbf0543fdc9866aeff2fca26ac241d3ba97f5a6bea9a25a83313a11be7169353

Observation d77a6c48-9565-4f39-98f8-80179d8c550b · inbound

Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical Analysis cites this paper.

Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical Analysis Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:11.100535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:11.100535Z digest=sha256:8d250518b9a1293fa4f7f434085af7c66b483a7cc459af8971130847e4bd8539

Observation 016af8a9-8f2a-410a-80f5-dbd356d88a0f · inbound

Aligning Proteins and Language: A Foundation Model for Protein Retrieval cites this paper.

Aligning Proteins and Language: A Foundation Model for Protein Retrieval Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:49.294185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:50:49.294185Z digest=sha256:d4e73d26688a0e46d97f64c167ff32f1b9019602ce59c49d8abe865c5a2103f7

Observation 34035c79-8394-499e-ab5e-55eef93e6e98 · inbound

Visual Pre-Training on Unlabeled Images using Reinforcement Learning cites this paper.

Visual Pre-Training on Unlabeled Images using Reinforcement Learning Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:20.470954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:06:20.470954Z digest=sha256:f34554da10fdb7d8d309c6858936881602e19f02918e5f0d5862187ba25ff868

Observation a9d6ab1f-8d29-4de2-b802-510eeafda76c · inbound

Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation cites this paper.

Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:07:31.213119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:07:31.213119Z digest=sha256:a91256a8d769430c5208db47af3ca38ffb760a85bc7ffd2f33b101faa55e7c3d

Observation 7568b478-a665-4c51-b333-02e35e594616 · inbound

MobileCLIP2: Improving Multi-Modal Reinforced Training cites this paper.

MobileCLIP2: Improving Multi-Modal Reinforced Training Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T14:59:19.571959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:59:19.571959Z digest=sha256:dacb90201ef2f755c59a9a9efeddda36146f401f3cb3d631a1e8f90f02c399a8

Observation 31375c32-c378-4279-8aa1-deb280288a65 · inbound

MApLe: Multi-instance Alignment of Diagnostic Reports and Large Medical Images cites this paper.

MApLe: Multi-instance Alignment of Diagnostic Reports and Large Medical Images Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.369563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T13:33:02.946839Z digest=sha256:920c78746894ca93b2e81097cbe84ea0247aea39fb628f6daa63310913d47f08

Observation b90db992-73a4-4d65-b59c-e8ee759f25d4 · inbound

Neutral-Reference Prompting for Vision-Language Models cites this paper.

Neutral-Reference Prompting for Vision-Language Models Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:18:54.591922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T19:15:54.152498Z digest=sha256:abc7fd7cdefea117e871a9c1ff922d3a1ce8e95adc46abcd5a2e41e25f0d6a74

Observation 40426464-81bd-4cec-aac9-4c6a226316bf · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:50.752092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:a6dfcbccf42139e90f8280817179fb3d2e7b1bd8607f23539f86f5e72b113ca6

Observation e444a6ec-ce9f-4c29-9e81-52cb1ed6fd91 · inbound

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks cites this paper.

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-10T15:47:23.263021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-10T15:38:58.361411Z digest=sha256:2f67aaeac1daf599ad95eb37841b95da392d6328670fdb0ef3c24ceea3e828c3

Observation 89e3af63-ede3-4e24-89ff-40271e8def59 · inbound

Test-Time Adaptation with Online Personalized Energy-Based Cache for Fine-Grained Video Expression Recognition cites this paper.

Test-Time Adaptation with Online Personalized Energy-Based Cache for Fine-Grained Video Expression Recognition Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:07.932304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:37:07.932304Z digest=sha256:1b9faf073903ddfb7f74acc91283135758d177ce24ab71c621c9edc680e7f9b4