Pith. sign in

Paper Citation Record · LEDGER

CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2203.07190.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.07190 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:40:27.682010Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T09:50:00.732703Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1dc487ff-e6c1-4cc3-99e6-44a31205782b · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.734828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:aa1d379bffe44bbd8ae2c4e1fe22175c8d9ca6fea6b5dd8076a85d88ac681fb4

Observation 57ded672-e299-492f-b903-43903541fc5f · inbound

Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training cites this paper.

Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T05:27:14.865838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:27:14.865838Z digest=sha256:12f51dfc302540c343d18e4217022e7122435b8997bef05995e84951a885cf36

Observation 15b8295a-9ff7-47af-bd7f-be6e5e19c222 · inbound

Human Action CLIPs: Detecting AI-generated Human Motion cites this paper.

Human Action CLIPs: Detecting AI-generated Human Motion CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.674694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.674694Z digest=sha256:3c3148fed1a5160c76d0414d5169148dfbca3211cd700ea06d46ba9da6ae5990

Observation ea5dd575-ea5b-44c1-a9c8-32ed09ae7ccc · inbound

Nearly Solved? Robust Deepfake Detection Requires More than Visual Forensics cites this paper.

Nearly Solved? Robust Deepfake Detection Requires More than Visual Forensics CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T20:31:57.251334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:31:57.251334Z digest=sha256:e65825f73a092d586f2574416251f34dbf634f6be8c1f1916a9cccd0ff3cdda7

Observation be67acef-79f1-4123-8589-88189b65c44e · inbound

DiffCLIP: Few-shot Language-driven Multimodal Classifier cites this paper.

DiffCLIP: Few-shot Language-driven Multimodal Classifier CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T19:11:02.456558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:11:02.456558Z digest=sha256:7753e51dfaa37b6bdf956017299ae01f72a53d2251ec6b3ca6acec1a0901b1fe

Observation 33f3847e-a73c-4d95-b38b-ffb1e24975a7 · inbound

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey cites this paper.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.478596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.478596Z digest=sha256:42580ac3aff3de07c258892fe6e481dc1a97101e0c16ddc067d8b95b6c2c6093

Observation adad59bb-9f9f-406c-9c95-8386147827c9 · inbound

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion cites this paper.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.318675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.318675Z digest=sha256:dd83446203a408514b6df2b1c1a66dce6fcc0d048aeccb0090ec8da554447a1c

Observation 6cacea6f-70fa-4dcc-b2fe-fbd51df36e67 · inbound

Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search cites this paper.

Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T17:27:25.803771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:27:25.803771Z digest=sha256:da9d8b8d5add2de5d57d7ae62c482b6aa7f2633e847df7228ced12adb63f6ca4

Observation fbcb3662-ca8e-4610-946d-517c1f28888d · inbound

Logits DeConfusion with CLIP for Few-Shot Learning cites this paper.

Logits DeConfusion with CLIP for Few-Shot Learning CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:27.682010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:27.682010Z digest=sha256:2e9521d9b03fc8fd39710d9bfdbdb562209790408d5df8933c32865cf1db784f

Observation f4bf2ebf-9f7f-49f2-8a34-66c2ce042822 · inbound

(Almost) Free Modality Stitching of Foundation Models cites this paper.

(Almost) Free Modality Stitching of Foundation Models CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:23.709123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:47:23.709123Z digest=sha256:89414803911086ff9d652b2df7242f3429de95f61ed0af066dd9c97487fd3641

Observation b20cc5d5-8999-419f-b50d-de5067b26b04 · inbound

Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval cites this paper.

Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T17:25:01.024416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:25:01.024416Z digest=sha256:a37b35956564595ca1fa08baa8bc4e8c5f7f57e880f2c25f56a9320ac25059bc

Observation 56c330f5-d58b-42c0-a4f4-b76187d16700 · inbound

O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation cites this paper.

O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T23:55:43.175206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:55:43.175206Z digest=sha256:0114f776ae3a670eec89a4a56d6e6ed8882e251bead9ca3e191236077f48d477

Observation 849ce55d-f81e-411d-a613-d97ed986faaf · inbound

WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval cites this paper.

WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:50:50.089156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T19:31:53.371412Z digest=sha256:d74394018f1fef66850172d7630ab3bc67cefe2f3b90e33604ed9d24ad9d5c53