Pith. sign in

Paper Citation Record · LEDGER

VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2212.04979.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.04979 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:42:52.024136Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T07:32:42.593033Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 270c8b6f-dbbe-4301-b895-4afee2432973 · inbound

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment cites this paper.

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 212

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:27:59.148501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T03:27:58.952076Z digest=sha256:0d7b7e97b1db52c5ce3b38f81504340a42ae4a660c8b3e4ba7349f3f3c077cd2

Observation f8c09a73-9b09-45fc-a634-70fa86f44c15 · inbound

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models cites this paper.

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T13:43:11.155513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T13:43:11.024069Z digest=sha256:3aeb1aed519895e511d29642feb444e86c47ed80c19b7fd6a51923b30ac80661

Observation f11a6134-c2e5-4735-9bcf-4d2ba70a9829 · inbound

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning cites this paper.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.112688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.112688Z digest=sha256:447df07a5bc594e5aa8b99fdc62d207c8afb1d32a8f2d81a50c069e3ca20e954

Observation 770b187d-d44d-457a-9bd3-b82401464665 · inbound

Health AI Developer Foundations cites this paper.

Health AI Developer Foundations VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:30:51.343704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:30:51.343704Z digest=sha256:cb13107714b3ed9af00434a963407c16cff34c2d605aa4885138b51d34d348e6

Observation 3d775e17-4d23-4001-854c-97df62bdf8e3 · inbound

LIVE-GS: LLM Powers Interactive VR Experience with Physics-Aware Gaussian Splatting cites this paper.

LIVE-GS: LLM Powers Interactive VR Experience with Physics-Aware Gaussian Splatting VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:32:42.596400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T07:32:39.047695Z digest=sha256:ca85d6dbe4cf12e4d5415c1ee153a126e41df9e963a3e25ad3aa93c768568110

Observation eb13a480-ead0-4234-ab84-56823a11744e · inbound

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering cites this paper.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.583896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.583896Z digest=sha256:462ab8d1b24c33dd640801185ad9f7e3124f6d82376c017855044c8229f5a7e9

Observation 33e06ae5-572b-4d9c-9a21-50afbc3fc684 · inbound

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment cites this paper.

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T00:47:12.459953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:47:12.459953Z digest=sha256:ec69dacff2af4ac67bf97eb41f0c1625154081de22cf4aa2983a4b73a1675761

Observation 4a0a0142-df48-4de6-b8ef-25a4c2753ec6 · inbound

3D Foundation Model for Generalizable Disease Detection in Head Computed Tomography cites this paper.

3D Foundation Model for Generalizable Disease Detection in Head Computed Tomography VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:27:27.035261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T03:26:46.665351Z digest=sha256:e72843ad27d2b9b5ae1dacb1e1895fb54aca7ce8f2c5b39d2a6acf759826789a

Observation dcf4fac7-81d6-49b8-9783-5da36827db87 · inbound

Learning Streaming Video Representation via Multitask Training cites this paper.

Learning Streaming Video Representation via Multitask Training VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-16T05:42:52.024136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:42:52.024136Z digest=sha256:9fc294997df024ee0086887d3797854809b50d44732db7873354b3e39460f3c9

Observation a2b7853d-f3cf-48e4-8dec-42757fedd10f · inbound

Vision Generalist Model: A Survey cites this paper.

Vision Generalist Model: A Survey VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 189

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.778078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.778078Z digest=sha256:59922d28ab78a5549082fb4580edea18c4d088fbdbe4d1bd33f5d66010cac64f

Observation 62afdd46-7c4f-4f69-83ac-f7cebed5a43b · inbound

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents cites this paper.

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T14:10:15.142182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T14:10:14.929207Z digest=sha256:a8090003b789a7ba41e83fc35641028fcb874516a95842fde2ce1a843307161a

Observation be5f889f-9d9e-4c31-bc0c-d8ab0347161f · inbound

Group Relative Augmentation for Data Efficient Action Detection cites this paper.

Group Relative Augmentation for Data Efficient Action Detection VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T12:55:40.391207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:55:40.391207Z digest=sha256:ef36863953fdecf2f528b18470073569348ab1ebf621545550036a2bb7f02f53

Observation dc6828fc-2087-4544-b2c8-cc2177230532 · inbound

OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward cites this paper.

OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T16:58:52.991179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:58:52.991179Z digest=sha256:0b13ee0c9afec7a04dc3a07b16ffed37a06a454d41e6dfee2a2dbbf9751c5b5b

Observation bdaaa35b-e120-41ed-8c49-3b6ca64cfa71 · inbound

Video Understanding by Design: How Datasets Shape Video Models cites this paper.

Video Understanding by Design: How Datasets Shape Video Models VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 225

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:42.190206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:42.190206Z digest=sha256:bc168961e5dcbdb36a21265d744b524ffb2cfceb8f8534dd1649d06fb4c8f08f

Observation d6745e37-4790-458e-a79b-80ad713b278e · inbound

UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations cites this paper.

UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T18:14:13.735386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:14:13.735386Z digest=sha256:23672be74bc64bd41c0adbc523e83a173d71d388133e5390e021c0f98827adbf