Pith. sign in

Paper Citation Record · LEDGER

VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2105.09996.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2105.09996 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:48:50.910615Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:09:37.771946Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bcff7980-216c-4fa2-aaa5-e9791c2fd649 · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

Reference 133

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.401847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:d967aca3b9f590122a6a61a055b7d6bd63f6a419317d4f930dc4df8dd5df8abe

Observation 2f274156-4873-4ef9-ab0f-4a87e6bc7137 · inbound

What Matters in Building Vision-Language-Action Models for Generalist Robots cites this paper.

What Matters in Building Vision-Language-Action Models for Generalist Robots VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.795057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:73d14752936d10629577e137700bca70571599d2d2b4f3e3fb7dfd9d0b8e8841

Observation 44be1a16-d6c4-49e7-8b77-9b3ced128894 · inbound

AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian Splatting cites this paper.

AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian Splatting VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T20:48:50.910615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:48:50.910615Z digest=sha256:3d26a9115f6ad731b33383bec907d78919483afaca9e392ee4bfa186bbba096c

Observation 3818203e-bbf9-46f3-acea-dc402f5d8ddc · inbound

The Eye of Sherlock Holmes: Uncovering User Private Attribute Profiling via Vision-Language Model Agentic Framework cites this paper.

The Eye of Sherlock Holmes: Uncovering User Private Attribute Profiling via Vision-Language Model Agentic Framework VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:11.869261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:11.869261Z digest=sha256:ad3a5f657f1adda2e0969f0d906b51665546ae088225d182170d9de29e6e46bc

Observation 88658e83-acac-4813-ab85-d4451a982d5f · inbound

AutoBridge: Automating Smart Device Integration with Centralized Platform cites this paper.

AutoBridge: Automating Smart Device Integration with Centralized Platform VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T11:04:04.650324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:04:04.650324Z digest=sha256:6f715e2965337375eaab9df702caecc279c8611f89cf6d02ceea58938d448c98

Observation 58daff7e-7c57-42bc-80f8-aba61d315be9 · inbound

Dual-Anchoring: Addressing State Drift in Vision-Language Navigation cites this paper.

Dual-Anchoring: Addressing State Drift in Vision-Language Navigation VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

Reference 135

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:51:45.890548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T06:50:34.310831Z digest=sha256:d01e47da4f86414f1090a737795c71562bf94f1accf9801839520ec4a0911247

Observation 4ac62342-5768-465e-9bce-a7cdb77fdc49 · inbound

T-MOR: Learning Motion-Aware Skeleton Representations for Human Action Recognition cites this paper.

T-MOR: Learning Motion-Aware Skeleton Representations for Human Action Recognition VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:09:37.788886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T14:44:41.975446Z digest=sha256:32f7a40c438a773cd58fc61cd7235ff45bd6167c4e8daf143d473732974c0af9

Observation e3200cfc-ed0f-4223-9f19-4c0725070bae · inbound

A Modular Vision-Language-Action Robotics Framework for Indoor Environments cites this paper.

A Modular Vision-Language-Action Robotics Framework for Indoor Environments VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:15:44.446427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T05:41:48.645193Z digest=sha256:5817e3725817f432dbc3422ce860f88352f1edc1cf7db92555adad33bd747f5a