Pith. sign in

Paper Citation Record · LEDGER

CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2203.07190.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.07190 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:47:23.709123Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T09:50:00.732703Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1dc487ff-e6c1-4cc3-99e6-44a31205782b · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.734828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:43f4e9030907d192cb99c4d4cba7323b65510cc5a9c5167bacd63cb1940f57f2

Observation f4bf2ebf-9f7f-49f2-8a34-66c2ce042822 · inbound

(Almost) Free Modality Stitching of Foundation Models cites this paper.

(Almost) Free Modality Stitching of Foundation Models CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:23.709123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:47:23.709123Z digest=sha256:aa92c29220925e8e93030fa13a003eb550a1f74f32d423e0a099d79493b39867

Observation b20cc5d5-8999-419f-b50d-de5067b26b04 · inbound

Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval cites this paper.

Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T17:25:01.024416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:25:01.024416Z digest=sha256:24a6c122844901eea17c1e29bf265c79b66193f818fcbf22d5ac9df618bb6f26

Observation 56c330f5-d58b-42c0-a4f4-b76187d16700 · inbound

O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation cites this paper.

O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T23:55:43.175206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:55:43.175206Z digest=sha256:2598663911212c7c02082602c94bb633691083235f25af041537b93626070c6e

Observation 849ce55d-f81e-411d-a613-d97ed986faaf · inbound

WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval cites this paper.

WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:50:50.089156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:31:53.371412Z digest=sha256:9364b1c622490e94638b9a20fd8b1718720ffd5420fb927aece3432800e3f7a8