Pith. sign in

Paper Citation Record · LEDGER

VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2105.09996.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2105.09996 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:48:50.910615Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:09:37.771946Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bcff7980-216c-4fa2-aaa5-e9791c2fd649 · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

Reference 133

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.401847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:dc781d81c5891683683c202c24e22fb92d3a8f90cf72c6c6c80e5da5f3d47579

Observation 2f274156-4873-4ef9-ab0f-4a87e6bc7137 · inbound

What Matters in Building Vision-Language-Action Models for Generalist Robots cites this paper.

What Matters in Building Vision-Language-Action Models for Generalist Robots VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.795057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:0ce942c8b5dc9ad130c41efa892ff856ca40103e3dbb8fb0fe28894d06b67a0d

Observation 44be1a16-d6c4-49e7-8b77-9b3ced128894 · inbound

AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian Splatting cites this paper.

AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian Splatting VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T20:48:50.910615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:48:50.910615Z digest=sha256:a5f7b51b50dcc93e184608ceda28ee142780e6839e22ca553a9a1b09a45f8cde

Observation 3818203e-bbf9-46f3-acea-dc402f5d8ddc · inbound

The Eye of Sherlock Holmes: Uncovering User Private Attribute Profiling via Vision-Language Model Agentic Framework cites this paper.

The Eye of Sherlock Holmes: Uncovering User Private Attribute Profiling via Vision-Language Model Agentic Framework VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:11.869261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:11.869261Z digest=sha256:cc0262be4553f5186ab69d92ba26359052b31b25c73ef1956e37d51a8439666a

Observation 88658e83-acac-4813-ab85-d4451a982d5f · inbound

AutoBridge: Automating Smart Device Integration with Centralized Platform cites this paper.

AutoBridge: Automating Smart Device Integration with Centralized Platform VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T11:04:04.650324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:04:04.650324Z digest=sha256:4aabc9f8e8f99633512110afa8f1e4cb1056495f9ba77f5acaea794587830ff8

Observation 58daff7e-7c57-42bc-80f8-aba61d315be9 · inbound

Dual-Anchoring: Addressing State Drift in Vision-Language Navigation cites this paper.

Dual-Anchoring: Addressing State Drift in Vision-Language Navigation VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

Reference 135

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:51:45.890548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T06:50:34.310831Z digest=sha256:7d37e595ce75c61f8cb3f0dda8387b94a66d7947cdc816aaed8c2ff370f2ca6c

Observation 4ac62342-5768-465e-9bce-a7cdb77fdc49 · inbound

T-MOR: Learning Motion-Aware Skeleton Representations for Human Action Recognition cites this paper.

T-MOR: Learning Motion-Aware Skeleton Representations for Human Action Recognition VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:09:37.788886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T14:44:41.975446Z digest=sha256:d0eae4122f6476ec0e99bcf4e8a8b901f4367ded779c7663e523c586d1a04c6e

Observation e3200cfc-ed0f-4223-9f19-4c0725070bae · inbound

A Modular Vision-Language-Action Robotics Framework for Indoor Environments cites this paper.

A Modular Vision-Language-Action Robotics Framework for Indoor Environments VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:15:44.446427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T05:41:48.645193Z digest=sha256:2209d53c79ff05d06b4ef23baeb884f8cb3026ae3bb9167c20556d2e3c69cb75