Pith. sign in

Paper Citation Record · LEDGER

MaskViT: Masked Visual Pre-Training for Video Prediction

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2206.11894.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2206.11894 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:17:39.573335Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

45
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 57cc9a3f-f86a-43d2-adff-6b8386c1e5af · inbound

Imagen Video: High Definition Video Generation with Diffusion Models cites this paper.

Imagen Video: High Definition Video Generation with Diffusion Models MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:31:08.208010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T03:31:08.155347Z digest=sha256:7a4d0e06640ba6d0cd5d564c874090043ea89f04a51548b699042029e3ae2b96

Observation 322c753b-1eb5-40d0-83f3-73dfb8ca9b53 · inbound

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory cites this paper.

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 279

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:03:58.131936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T13:03:57.828598Z digest=sha256:2b8ae13c72cea90916a7bfcaef383fb0e10c86d77faa1f823a7fd74654c7cf22

Observation dd7a038e-511f-4579-860f-c97877746cf7 · inbound

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution cites this paper.

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 127

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:12:31.533823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T08:12:30.984870Z digest=sha256:0dae086cd04e79f534804210199cc7f11437ac1193dc10cf22787a609f06a312

Observation 325101e5-1056-41ce-99e0-eb6b8f9b823d · inbound

Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation cites this paper.

Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 140

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T16:32:05.705964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T16:32:05.538507Z digest=sha256:22eebe4541abb36623d504e6b8c28ef90d13c2b048d9d750c470d16eb441845a

Observation 58b1e455-212d-4e5b-8210-49090e264496 · inbound

VideoPoet: A Large Language Model for Zero-Shot Video Generation cites this paper.

VideoPoet: A Large Language Model for Zero-Shot Video Generation MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:51:05.690665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:cfe7f09d9c4a9c4e8b840f3f738741011a25c2a2452e96ba41c5d2db46f74529

Observation 29e36350-4b57-4f41-bfde-68b3ed2c54fd · inbound

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models cites this paper.

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 173

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T13:43:11.244642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T13:43:11.024069Z digest=sha256:c3980d7144d4c8e27bfd96ad18d064e09c9d24cdc194008ff311e7eee0136e08

Observation d20b0b42-f77d-4a59-96a1-a9d30b90dd33 · inbound

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation cites this paper.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:09:34.008304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:32dda5ba2f4bd8f4538e129cdbb2799cb214ef97ef7b95b159fe4c880be981f2

Observation f230e108-8784-43cc-9750-f5125c4111f8 · inbound

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders cites this paper.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.573335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.573335Z digest=sha256:2334cde0df7932adeb49d0333880bc9b60b40837288647e15fe35bfb3ea0ab50

Observation 72503797-2ce4-4cca-ac27-8733d13e0d42 · inbound

Pre-Trained Video Generative Models as World Simulators cites this paper.

Pre-Trained Video Generative Models as World Simulators MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T15:14:22.326566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:14:22.326566Z digest=sha256:f2fe4108d61204a4741c029e82e3606ff94aa4377b9f5499c49a3a704e42dbee

Observation 17c4a7e3-3912-4ab2-ad57-526f15ef0ecc · inbound

Boosting Bot Detection via Heterophily-Aware Representation Learning and Prototype-Guided Cluster Discovery cites this paper.

Boosting Bot Detection via Heterophily-Aware Representation Learning and Prototype-Guided Cluster Discovery MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:47.859287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:47.859287Z digest=sha256:0c9f352da838d0fac4bf7ffc548bf6e090b9e67a89af34cf1e1be35df58af438

Observation c7a600d1-840b-4ff7-86ef-153bff29820d · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:33:50.835310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:14fd26b696cf3def19f692d28ce711546c1ee4a5b6ebd2d6eade5e08d571f7f2

Observation 87e16809-6d0e-4eda-b546-ff98c30ed626 · inbound

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning cites this paper.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.617134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.617134Z digest=sha256:b45c1d5ddc82796fd9262b68b52a6c620414d8398ccecbd0fa1cd38e10b10c1e

Observation e38e71c2-4420-4ca7-bd22-554c24bb8f86 · inbound

Self-Guided Masked Autoencoder cites this paper.

Self-Guided Masked Autoencoder MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:53.899947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:13:53.899947Z digest=sha256:1bbec2d74e8addcc493dfe7b0bafb0ebed1ad0849c78a9f5f78830733dce010d

Observation 338d3494-bf24-40cc-ad35-4201a475fc80 · inbound

RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation cites this paper.

RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T10:43:36.981496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:43:36.981496Z digest=sha256:7231b049a7da081e9d9975a8fe9c519d1c77b9cc341b5bd1b5891908deadac81

Observation 477e305c-7ee3-4735-99a5-081ee04fde00 · inbound

Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Grouping cites this paper.

Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Grouping MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:56:33.933613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T19:51:58.756605Z digest=sha256:5e569a47a2150702eaa7afeeb3500cd7569f02cd28d9260e2719f8c8562e33a1

Observation a6a93d62-b548-42c8-b193-049e660006c1 · inbound

WestWorld: A Knowledge-Encoded Scalable Trajectory World Model for Diverse Robotic Systems cites this paper.

WestWorld: A Knowledge-Encoded Scalable Trajectory World Model for Diverse Robotic Systems MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:44:09.343612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T11:40:28.002606Z digest=sha256:8152b9e77be9de5afdf761178d3ff05d300828bfafc4f26c8de4ed3c96c855ad

Observation e7346973-d3c1-4f50-916f-c54499c48639 · inbound

Physically Viable World Models: A Case for Query-Conditioned Embodied AI cites this paper.

Physically Viable World Models: A Case for Query-Conditioned Embodied AI MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T09:13:16.484324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T06:55:57.801162Z digest=sha256:c22b9eb622dfd342c9fc738ea712dc4695f2914020b0cf553f8e745cb0fdb965