Pith. sign in

Paper Citation Record · LEDGER

VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2203.12602.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.12602 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:46:14.151441Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

436
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1a18dcdb-9497-4d96-8ac2-20f2d22e7a76 · inbound

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment cites this paper.

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 124

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:27:59.064233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T03:27:58.952076Z digest=sha256:c0ce894c8c2f09f0f38a5990946166d0c68f24eff0abde392132dbd5f11b456c

Observation 1ee0adec-e107-440f-9b09-0ad25428836c · inbound

Fine-Tuning Video Transformers for Word-Level Bangla Sign Language: A Comparative Analysis for Classification Tasks cites this paper.

Fine-Tuning Video Transformers for Word-Level Bangla Sign Language: A Comparative Analysis for Classification Tasks VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:46:14.151441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:46:14.151441Z digest=sha256:91ace581b37aa5aa984dd7db74cf717ebe6f90a23eca998afe9d4617c9a4d292

Observation 3a3ce278-b861-4ab0-a19b-46b89a94b4c6 · inbound

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory cites this paper.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:59.459159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:59.459159Z digest=sha256:67e94348f893a4c8fa2da8105495c808fbba9e09a815985a5af099a8bfb147cf

Observation 596e9c2a-34a5-4024-9e7c-e77d82852611 · inbound

One Video to Steal Them All: 3D-Printing IP Theft through Optical Side-Channels cites this paper.

One Video to Steal Them All: 3D-Printing IP Theft through Optical Side-Channels VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:19.049850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:19.049850Z digest=sha256:d3e873f136a984608454916233fc59df6dae0a820f3b91a66c75eda9618d6c24

Observation af449d3e-434b-4f99-ada6-63a5585c1387 · inbound

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition cites this paper.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:57.583666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:57.583666Z digest=sha256:f439d81e0db47efe435f6dde6bfb5fef1ac0fd58a162f9f14cf8a666291d10c3

Observation 1343f094-5c33-4b3f-b3ad-46901b5913de · inbound

MVP: Winning Solution to SMP Challenge 2025 Video Track cites this paper.

MVP: Winning Solution to SMP Challenge 2025 Video Track VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:56.193790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:56.193790Z digest=sha256:a5db7b6112bdec2f78558c631659d0b0fc85bea402e858f3de456763953e4c79

Observation 13d5428e-6ba5-44f3-9677-ad9af44217c8 · inbound

Multimodal Framework for Explainable Autonomous Driving: Integrating Video, Sensor, and Textual Data for Enhanced Decision-Making and Transparency cites this paper.

Multimodal Framework for Explainable Autonomous Driving: Integrating Video, Sensor, and Textual Data for Enhanced Decision-Making and Transparency VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:33:25.677404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:33:25.677404Z digest=sha256:f031db21f8af3eef5a2633448624fa708ac4d97018484f8dad691525d11e9981

Observation 8627080a-8146-4fd5-893f-d0b40c9a5e1c · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.574181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.574181Z digest=sha256:453a59422dcca91cf4229e9de596c3750840b1a7a72bda9dc21decaaca0cd5f6

Observation 097ec171-d4f5-48ea-8624-e42e5edbdbd3 · inbound

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders cites this paper.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:65953e7a16c495d058579847f8470dcf062ee56e0cf05c7226ffdd113b8d7054

Observation bd06e627-9cbf-48b4-9d99-0592fa36bc05 · inbound

Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer cites this paper.

Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:01.661100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:07:41.426142Z digest=sha256:d1994fe9ef8f1bf951b3dd21c3e0f3148f5fe55f8a2444e054d9866e19044f52

Observation 32eda57c-720b-401e-8905-3813f7ac41bd · inbound

Zero-shot World Models Are Developmentally Efficient Learners cites this paper.

Zero-shot World Models Are Developmentally Efficient Learners VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:21:00.069998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:33:39.342672Z digest=sha256:c1465da2717ae0251e48ee1f68cbbdb350b9026fc69b6345e51d874a4b08de23

Observation 436e807f-cc52-4e35-b968-5a6601d722d9 · inbound

Beyond Independent Frames: Latent Attention Masked Autoencoders for Multi-View Echocardiography cites this paper.

Beyond Independent Frames: Latent Attention Masked Autoencoders for Multi-View Echocardiography VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:00:22.231240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T11:56:05.271784Z digest=sha256:2dff8acf055623626d72d489724904f48688661a92a61f6eb7880c2c930ab987

Observation 10c72199-24da-4342-ad3b-d63142d15e99 · inbound

Mask World Model: Predicting What Matters for Robust Robot Policy Learning cites this paper.

Mask World Model: Predicting What Matters for Robust Robot Policy Learning VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:11:05.988492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:14:17.676675Z digest=sha256:0b2d0b1288c1daa884f3b06327e137844bdfec090cd489ca766e57cfedc6994b

Observation 5c24090c-e261-4807-9828-8c5940807490 · inbound

SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition cites this paper.

SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:30.819738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T19:23:07.258575Z digest=sha256:d8ad0e6e83a44f597e97027ca394a9f5556688797ecfd6afdd092b291ed856fe

Observation 012078d3-2e08-4cc7-9d3d-bd8b845b2204 · inbound

SpecSem-Net: Integrating Spectral and Semantic Features for Robust AI-generated Video Detection cites this paper.

SpecSem-Net: Integrating Spectral and Semantic Features for Robust AI-generated Video Detection VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:23:21.618120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T14:19:36.263592Z digest=sha256:ce93d46e4feac40638f6aa1b6817725d7d32fab768b0619e62c46ff476f19e3b

Observation eb425c87-a1ff-4375-8e13-c999bb6f4426 · inbound

Prognostic Value of Lung Ultrasound Biomarkers for Readmission Risk in Congestive Heart Failure: A Pilot Data-Driven Analysis cites this paper.

Prognostic Value of Lung Ultrasound Biomarkers for Readmission Risk in Congestive Heart Failure: A Pilot Data-Driven Analysis VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T15:43:26.997735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T15:41:31.181025Z digest=sha256:d026f9d9f1fba0897da44ff2a3bf8b5246eacfd3d549503f5a8e421b88fdc53d

Observation 666c8eb3-efaf-45fa-bfc3-08b6c9fdab57 · inbound

FAST-ME: Foundation-aware Adaptive Stopping for Motion Estimation for Efficient IoT Video Analysis cites this paper.

FAST-ME: Foundation-aware Adaptive Stopping for Motion Estimation for Efficient IoT Video Analysis VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:50:21.230946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T04:47:22.381247Z digest=sha256:4300cdd4bd6292b02d1685ed1b3ca269bc78ca41fd07cfc982fdc312f5d3c141

Observation a4990119-6ac8-4626-9ad9-6aedb9c67083 · inbound

EVA-Net: Subject-Independent EEG Motor Decoding with Video-Derived Motor Priors cites this paper.

EVA-Net: Subject-Independent EEG Motor Decoding with Video-Derived Motor Priors VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:20.318495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T14:59:24.109366Z digest=sha256:9f74cb06e0d82ad4528d3bfc1683f25d656b167a2d9f1a336c7b860a13017d5e

Observation 60ed711d-f744-49c6-8541-5677007360d6 · inbound

EVA-Net: Subject-Independent EEG Motor Decoding with Video-Derived Motor Priors cites this paper.

EVA-Net: Subject-Independent EEG Motor Decoding with Video-Derived Motor Priors VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T15:23:58.657504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:23:58.657504Z digest=sha256:aa05dbc6f72d862b17d6d65f885f17932c46a2c654e656b2a717783bed3d54e4

Observation ec84fac8-62f7-4d77-82c0-93596628fa38 · inbound

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning cites this paper.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:56.824828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T01:52:44.785582Z digest=sha256:df5215f487f2bc21510f23ef8bbaf8b6d12a4ee46864ffa7be7d48fb4142b9ff

Observation d398b846-83d2-478c-a491-b75ce976e56f · inbound

BioVid: Autoregressive Video Generation with Biological Behavior Semantic Comprehension cites this paper.

BioVid: Autoregressive Video Generation with Biological Behavior Semantic Comprehension VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-27T18:41:07.978312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T18:35:14.331659Z digest=sha256:a248c6ef32c4ea129e156f1cebd02b0da59bce00bb3145b1e0a2c39b0b81ec77

Observation b91c577c-86eb-4d8e-87eb-fcc1e100c55a · inbound

Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis cites this paper.

Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:07:28.693423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:25:15.493925Z digest=sha256:f0ae30704f9130bf703c2b264512328810326d5b76df9891309b7ca014b2cccb

Observation 12dc5c0b-df9a-48d7-a577-d48b2aa1b896 · inbound

Deep Learning-Based Sign Language Recognition from Videos and Cross-Lingual Translation to Indian Vernaculars cites this paper.

Deep Learning-Based Sign Language Recognition from Videos and Cross-Lingual Translation to Indian Vernaculars VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:49:42.467011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T10:55:03.728012Z digest=sha256:93dffd7fa1eb2bb158ecef67e5d5a7ad81d2da555ff9c697edc4373834ff4c4e

Observation ab485a0f-7d57-4e89-84c5-b056706ae7a3 · inbound

Empirical Evaluation of Multi-Modal Touch Detection in Over-the-Shoulder Video Surveillance cites this paper.

Empirical Evaluation of Multi-Modal Touch Detection in Over-the-Shoulder Video Surveillance VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:04:20.959297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T07:02:23.993806Z digest=sha256:64d102a8eefa2cb16598a2ba91f4cb6831c8b603e262916c3076c9dc87afdc34

Observation beb3c69b-1797-4c66-853a-d40f3b182606 · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.632640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.632640Z digest=sha256:6515d5cb74aa642814ac26e595e0d887ece0556455c1b2d73ba778745a7206b5