Pith. sign in

Paper Citation Record · LEDGER

CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2209.06430.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2209.06430 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:02:42.840512Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T15:17:07.354594Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8658039f-ff8d-44e3-b8bb-05a068784109 · inbound

A Survey on Foundation Models for Personalized Federated Intelligence cites this paper.

A Survey on Foundation Models for Personalized Federated Intelligence CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:34:57.747854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T15:32:15.293888Z digest=sha256:151749edb7182d78f31305ba1fa7f8d5d24ad635a6430b1529ca5e4ca7ccd6ee

Observation 20ca38b7-a66a-469a-9522-af7642c6a203 · inbound

Adversarial Video Promotion Against Text-to-Video Retrieval cites this paper.

Adversarial Video Promotion Against Text-to-Video Retrieval CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:06:55.131999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T00:05:07.182361Z digest=sha256:688d127ad5ad8f377049662dd6114263ec6b61992b43278c7332e491b40f179a

Observation 56d93407-97c0-4006-adbd-0514b02ba657 · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.941379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:bec985427c1fd6bc1f70774075fe170a9bf2878f2e1747669e7cf621757506f5

Observation de5b80b0-bbc7-4067-a663-182999067c1c · inbound

CoVR-R:Reason-Aware Composed Video Retrieval cites this paper.

CoVR-R:Reason-Aware Composed Video Retrieval CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T21:37:55.887477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T21:37:55.887477Z digest=sha256:852ca764a02bd6ee09af945a9a49679327823c3f5e0a682cbb98881938af0e81

Observation 22e6ccf5-b049-414e-b36a-35f80997fd69 · inbound

VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement cites this paper.

VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:17:07.356298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-07-02T15:09:47.855795Z digest=sha256:4226ce3c671a21d39effb40d85bba82517488b125d1bc8ea7082892ce7625b93

Observation 7c807420-93b1-4ad1-948b-697b97b30293 · inbound

Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing cites this paper.

Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T08:55:44.783824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T08:55:44.783824Z digest=sha256:8c4448b044bf7c5b3a81a760b7bcf3ea6584d881cc759356adc09a99575a0777

Observation e0bf0880-0c9a-4610-a8d4-1c18e2e63f2b · inbound

Knowledge-guided Disentanglement with Atomic Actions for Action Recognition cites this paper.

Knowledge-guided Disentanglement with Atomic Actions for Action Recognition CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:42.840512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:02:42.840512Z digest=sha256:4ba713ac78222d27398f2d15086955223fab121c8c2aabb44fafe9a6bbbcab4c