Pith. sign in

Paper Citation Record · LEDGER

Learning Spatiotemporal Features via Video and Text Pair Discrimination

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2001.05691.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2001.05691 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:17:39.665769Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T04:53:57.836625Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f6427790-add4-4399-ba0e-9dd4f2a67756 · inbound

InternVideo: General Video Foundation Models via Generative and Discriminative Learning cites this paper.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Learning Spatiotemporal Features via Video and Text Pair Discrimination

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.277470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:e192df9982e8363234c9cf0f35b1412e7d971fe6357c06494ac005043f116292

Observation 78b1790f-89ba-42ad-9701-6e8853b2cb3b · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding Learning Spatiotemporal Features via Video and Text Pair Discrimination

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.585193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:c31621affbdfe39cd3e33ff84466634b051de7a98632576f96384250a7a2ef0e

Observation 48d86566-db18-4702-8d62-034dec52029b · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Learning Spatiotemporal Features via Video and Text Pair Discrimination

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.546904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:8b83238d9ec90e5b379a4b51cd6125cafe90bd9b62c324842b32c0b53978f4b6

Observation b6fa720b-5e50-44cb-8783-ce9ae101ed0f · inbound

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders cites this paper.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Learning Spatiotemporal Features via Video and Text Pair Discrimination

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.665769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.665769Z digest=sha256:16afa21b9bf9ff0d1e3ef5bf01de9578829c1670a6b74e4ed369f81397da5e68

Observation 41f40409-73ee-4aed-b6f5-d977c26c61c3 · inbound

USV: Towards Understanding the User-generated Short-form Videos cites this paper.

USV: Towards Understanding the User-generated Short-form Videos Learning Spatiotemporal Features via Video and Text Pair Discrimination

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:53:57.838074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:701ad6f0e0e7ed1bb0a3c45603a5dc55f708fe931659cb55a25468ad154f48ac