Pith. sign in

Paper Citation Record · LEDGER

Unmasked Teacher: Towards Training-Efficient Video Foundation Models

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2303.16058.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.16058 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:32:15.474943Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T03:27:59.199947Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ab5c4fc4-b242-4bda-93f8-b7df99450714 · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding Unmasked Teacher: Towards Training-Efficient Video Foundation Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T23:30:00.572703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:f81bd43bc729ef344537d7d675a39404d7ee5e2190ed3ce9e06f523be14cbd1c

Observation c0869013-6a87-41ef-a2fa-73804324fa1b · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Unmasked Teacher: Towards Training-Efficient Video Foundation Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:30:22.486478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:77d3cb7e4cdb72135e009b46bb2d36b4de64105958124052ced67314a16b403d

Observation 6f283b83-764f-4bab-9f4d-1a0e81969c34 · inbound

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment cites this paper.

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment Unmasked Teacher: Towards Training-Efficient Video Foundation Models

Reference 153

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:27:59.205034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-17T03:27:58.952076Z digest=sha256:333be8b965db319c6cfaca33be9f03ff21fe6d992b67676786f7adc1843df984

Observation dbb38b48-8acd-4720-92d5-e473c380ada2 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks Unmasked Teacher: Towards Training-Efficient Video Foundation Models

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:46:09.978550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:1306b43fe9fc7885b8be8bc85512f02353f69cfa74f05bc584f03aec9c7e4a9e

Observation 20f72546-6c43-4c18-be9a-2fe99573e7d9 · inbound

Revisiting Feature Prediction for Learning Visual Representations from Video cites this paper.

Revisiting Feature Prediction for Learning Visual Representations from Video Unmasked Teacher: Towards Training-Efficient Video Foundation Models

Reference 188

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:40:23.998257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T12:40:23.709098Z digest=sha256:4145b2ab7faa53dbc34a599aac262596c1653af8a678c965f6eea53f4626e5dd

Observation 9a1be07f-ca20-4d33-941e-20cf47ff3536 · inbound

SMART-Vision: Survey of Modern Action Recognition Techniques in Vision cites this paper.

SMART-Vision: Survey of Modern Action Recognition Techniques in Vision Unmasked Teacher: Towards Training-Efficient Video Foundation Models

Reference 160

Resolution
unresolved
no resolver link, observed 2026-08-10T16:32:15.474943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:32:15.474943Z digest=sha256:bc9c6c855d5249dda86ec3e181f34769dd0ebf95032b1c75d89f4a27634771d8

Observation 3bd19b2e-3c20-4bcf-95e5-c67e41ab1844 · inbound

A Survey on Efficiency Optimization Techniques for DNN-based Video Analytics: Process Systems, Algorithms, and Applications cites this paper.

A Survey on Efficiency Optimization Techniques for DNN-based Video Analytics: Process Systems, Algorithms, and Applications Unmasked Teacher: Towards Training-Efficient Video Foundation Models

Reference 141

Resolution
unresolved
no resolver link, observed 2026-08-06T15:30:20.062278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:30:20.062278Z digest=sha256:508f82b7bc53fa3dbebfc86bf09ce27cd5da7412406b401179d4825841c8a9e0