Pith. sign in

Paper Citation Record · LEDGER

Extending Video Masked Autoencoders to 128 frames

As of 22 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 0 inbound Pith citation observations for arXiv:2411.13683.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13683 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:22:25.355326Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

82 of 82 outbound references displayed

  • verified exact0
  • verified fuzzy71
  • unresolved10
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 84cf1d58-7cf6-4e59-b24a-18365b0499be · outbound

This paper cites Towards long-form video understanding.

Extending Video Masked Autoencoders to 128 frames Towards long-form video understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.014117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.014117Z digest=sha256:df4b9cc82fdc4a581019a0717f456cc000c94eb7d535897179dcf0b3d87b9db9

Observation 57d954e0-e57e-4b91-834e-9520d8a00794 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

Extending Video Masked Autoencoders to 128 frames Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.018939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.018939Z digest=sha256:e78b7f74fcd78aa4bfd49ad33e99fc1fcf48604401e6f624be7e21db6bd3c463

Observation 1e2090b2-dcc2-4824-91ac-5108d39886e7 · outbound

This paper cites Long movie clip classification with state-space video models.

Extending Video Masked Autoencoders to 128 frames Long movie clip classification with state-space video models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.442059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.023182Z digest=sha256:649eb126a9933212a6251684014a01e8bdaff73878ba4e32add1dd6dd0d521aa

Observation 3abf5ab4-b5ec-4dfe-a7e7-e0e7e9f6434f · outbound

This paper cites Selective structured state-spaces for long-form video understanding.

Extending Video Masked Autoencoders to 128 frames Selective structured state-spaces for long-form video understanding

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.428871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.027369Z digest=sha256:e6200123d38b873db8a2408e89774102a3e8e2f0ed5b689a441ffd91fac08ffb

Observation e1d2eb4b-a9ef-4679-8eb6-ccfb70df90db · outbound

This paper cites Memory consolidation enables long-context video understanding.

Extending Video Masked Autoencoders to 128 frames Memory consolidation enables long-context video understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.415408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.031425Z digest=sha256:4e39c3025c15ec5c5c264fad052a1d3a2c7dd5a430a4b838bf75af32d9679d09

Observation fad0160e-a622-46f0-8cbe-1ba5c085f687 · outbound

This paper cites Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition.

Extending Video Masked Autoencoders to 128 frames Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.401384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.036030Z digest=sha256:d5e40bdc2629cfdae0911263cb693f9f4612fc4e10d9464a023508f2192babd5

Observation 9136e534-9098-4a59-9d9b-2b43c8578db5 · outbound

This paper cites Token turing machines.

Extending Video Masked Autoencoders to 128 frames Token turing machines

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.386697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.040506Z digest=sha256:6fd7d12b12983574f607490ac96540d906a2ba70a0055d7dbb9520050890bcc7

Observation 1ee8d817-2f90-4da5-b5ff-b16c8255302b · outbound

This paper cites Video recap: Recursive captioning of hour-long videos.

Extending Video Masked Autoencoders to 128 frames Video recap: Recursive captioning of hour-long videos

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.374304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.045808Z digest=sha256:7ac89e6213093568a55d9041ade2ec5ccc876887bcd4197eeef090a5d745b04e

Observation 15282ec6-ac03-48a8-bb04-1f898e7a3bc4 · outbound

This paper cites A simple recipe for contrastively pre-training video-first encoders beyond 16 frames.

Extending Video Masked Autoencoders to 128 frames A simple recipe for contrastively pre-training video-first encoders beyond 16 frames

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.360794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.050401Z digest=sha256:c9e3e215d06db96df6cc0fcb232fad211f8461e5ff0a6633d195dd96019db9ff

Observation 96c04c42-cb02-4ce3-a8e5-aba75807fbeb · outbound

This paper cites A simple llm framework for long-range video question-answering.

Extending Video Masked Autoencoders to 128 frames A simple llm framework for long-range video question-answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.054724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.054724Z digest=sha256:c93345087037783725f166de49c0ef2a251d4e28c7d83ff670223c770decac52

Observation b28fee82-238e-451f-a888-05e92652f15a · outbound

This paper cites Long-form video- language pre-training with multimodal temporal contrastive learning.

Extending Video Masked Autoencoders to 128 frames Long-form video- language pre-training with multimodal temporal contrastive learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.341094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.058702Z digest=sha256:90dc64b5b08bc3dbbbab1261b0cd45970c7f77d3d1ac54fdf8f1dec69e4b7745

Observation 4dc4e05a-4e81-45f9-ba36-ef5bdde63ccb · outbound

This paper cites Koala: Key frame-conditioned long video-llm.

Extending Video Masked Autoencoders to 128 frames Koala: Key frame-conditioned long video-llm

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.328284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.062400Z digest=sha256:c59878ffe0d78175100e01a25776d62b51fd0d9ce1243d37757ead9e47193e2b

Observation 9104b8a1-ea81-4e5d-b003-51f7c116eb30 · outbound

This paper cites Masked autoencoders as spatiotemporal learners.

Extending Video Masked Autoencoders to 128 frames Masked autoencoders as spatiotemporal learners

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.316949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.066391Z digest=sha256:05ba9ce387607842d24c22649f359b86a13416df9943b7136f1f44eaf932a539

Observation 30a10d9c-6a53-456b-9d89-fc6543a31d37 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

Extending Video Masked Autoencoders to 128 frames Videomae v2: Scaling video masked autoencoders with dual masking

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.304662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.071077Z digest=sha256:9546ce3d32841c5d45f59650455ee7df9a08bc81f5acc3beb95238bbbf659e99

Observation 61aa62bd-3850-4120-ade0-88283e9a3003 · outbound

This paper cites VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Extending Video Masked Autoencoders to 128 frames VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.290950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.074927Z digest=sha256:ad0f79518fe886a55f28af1d970be5fe790eb75f12dade2a51491431a40d8058

Observation cdd82006-37e8-4192-a1fe-5725f23d39cc · outbound

This paper cites How can objects help action recognition? In CVPR, 2023.

Extending Video Masked Autoencoders to 128 frames How can objects help action recognition? In CVPR, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.278784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.078492Z digest=sha256:ce9ebebc93ea28dca1eeaed639545a518341c9d7d3a17713d825e78b6192af46

Observation 2d832d9b-485e-4aca-8091-c506b8885877 · outbound

This paper cites Everest: Efficient masked video autoencoder by removing redundant spatiotemporal tokens.

Extending Video Masked Autoencoders to 128 frames Everest: Efficient masked video autoencoder by removing redundant spatiotemporal tokens

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.266111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.082930Z digest=sha256:fd8cd9d446bfb766b647e50b4fc2837380839bf4aba133befb7c73803bbb14b9

Observation 9d1350a1-3813-4f80-bbf4-6a55f4baba20 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100.

Extending Video Masked Autoencoders to 128 frames Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.253869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.086594Z digest=sha256:9158756ee6441f1abf9adf6831a418c19caaff98ad9a6eb6b18a069d4ca9adb2

Observation 6c0d1bea-e94f-46af-ab04-b4c051bd18c1 · outbound

This paper cites Resound: Towards action recognition without representation bias.

Extending Video Masked Autoencoders to 128 frames Resound: Towards action recognition without representation bias

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.241287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.090422Z digest=sha256:a414e205b9424504bfac80aee53d03c268ade7ef482fbc78f5384dabbf127cab

Observation 21f1223f-011f-4832-ade6-2ee718f0afb4 · outbound

This paper cites an unresolved cited work.

Extending Video Masked Autoencoders to 128 frames Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:22:26.229238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.094637Z digest=sha256:386e7e719ec0f9810790db405306ba9ae322f4e1cb6608b2de77228a10b469c5

Observation 05497a92-60ff-443e-9f04-71f253a54044 · outbound

This paper cites Bevt: Bert pretraining of video transformers.

Extending Video Masked Autoencoders to 128 frames Bevt: Bert pretraining of video transformers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.218024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.098703Z digest=sha256:a47b054c726f135c07ba06158221bab0b30ce2bf3cc1ce9bf092338923820c27

Observation e5e7babb-2187-497a-af9d-76469a1373b0 · outbound

This paper cites Girdhar, A.

Extending Video Masked Autoencoders to 128 frames Girdhar, A

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.205639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.103922Z digest=sha256:2d001430b78f87e61265910d7f571aa4aad2ac7d2688a808e926e33fe25978f0

Observation 713e89ef-4263-417a-b095-74a19c92e633 · outbound

This paper cites Zero-shot text-to-image generation.

Extending Video Masked Autoencoders to 128 frames Zero-shot text-to-image generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.193574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.107960Z digest=sha256:d13e99d2596f2fd1cec38b12f01cd587dd88f288b67d5cb1332cb232f322c4d8

Observation bc7e4fd2-2979-4311-a58b-b08529e687ad · outbound

This paper cites Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning.

Extending Video Masked Autoencoders to 128 frames Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.182211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.111942Z digest=sha256:f950d5110879b6c4a52a2bb52db3910d53e9fd88fa7c7276e59d3832315e8be4

Observation e5511110-75f1-465e-bd26-2f2d846c7286 · outbound

This paper cites Magvit: Masked generative video transformer.

Extending Video Masked Autoencoders to 128 frames Magvit: Masked generative video transformer

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.170158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.115504Z digest=sha256:a30e47f2b8f98f1d03d92c26d01cea3eef7d2cdc59e21169ea87f2d3e590ad03

Observation 929db5b8-04b5-4e3d-a6cc-91d48a12431c · outbound

This paper cites Tokenlearner: What can 8 learned tokens do for images and videos? In NeurIPS, 2021.

Extending Video Masked Autoencoders to 128 frames Tokenlearner: What can 8 learned tokens do for images and videos? In NeurIPS, 2021

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.143450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.123089Z digest=sha256:3670e1a1e4bb685e59b2d5dbf657a9ef79c31ddea407bc7f6d7731b2b7957cd7

Observation 9cd15a53-f043-4a68-9079-f5e6871c87d9 · outbound

This paper cites something something.

Extending Video Masked Autoencoders to 128 frames something something

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.130927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.126787Z digest=sha256:6c910122190248796aab31275b79347c8f9eee74c79de1526d1e51be766b18cd

Observation 319a40db-cab6-42ab-86cd-b86c83095958 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Extending Video Masked Autoencoders to 128 frames An image is worth 16x16 words: Transformers for image recognition at scale

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.118986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.130449Z digest=sha256:3b3a02647fa9db47ede7c36af36c51867c3e149794d35179f8e559827652409b

Observation d87ef826-a40a-4764-958d-63623af7b0cc · outbound

This paper cites Mgmae: Motion guided masking for video masked autoencoding.

Extending Video Masked Autoencoders to 128 frames Mgmae: Motion guided masking for video masked autoencoding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.105953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.134329Z digest=sha256:3e00ac9c86b7b16b0e2c898c79205f9773cd8bd549c9511ba4a8814696e28db2

Observation b5b0e6f3-fc2d-408f-a232-3ca3c19afbf2 · outbound

This paper cites Motion-guided masking for spatiotemporal representation learning.

Extending Video Masked Autoencoders to 128 frames Motion-guided masking for spatiotemporal representation learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.091040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.138258Z digest=sha256:b1ba18d5e8d7b57eb73b647000d0a0f732cc8fc08807ae9fed58f1078cbd9e5c

Observation e32096bb-4f64-4512-ac99-52032f45532d · outbound

This paper cites Video codec design: developing image and video compression systems.

Extending Video Masked Autoencoders to 128 frames Video codec design: developing image and video compression systems

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.077412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.142170Z digest=sha256:15d49c3b600b1dc98b12c42e0bc453bdd161a9c155528c5038f55a1cfaf7eb1c

Observation 9a8ae6a4-ee82-4266-bced-f94d9ea68a02 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

Extending Video Masked Autoencoders to 128 frames Raft: Recurrent all-pairs field transforms for optical flow

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.062271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.146065Z digest=sha256:20d0d4852ade8e67cf67e6fb10d1f36f8bd82b1884c28d62a24716c64c614c8a

Observation d71d0fc2-7964-48d3-9875-5cb23baf2090 · outbound

This paper cites Videoprism: A foundational visual encoder for video understanding.

Extending Video Masked Autoencoders to 128 frames Videoprism: A foundational visual encoder for video understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.050812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.149577Z digest=sha256:41a0c72d412f5931e5293c6fcc9c46ce56d677b5f0193590c6c7946199e41a11

Observation 0c81c503-ca16-426b-8f61-e24d49b7afd8 · outbound

This paper cites Internvideo2: Scaling video foundation models for multimodal video understanding.

Extending Video Masked Autoencoders to 128 frames Internvideo2: Scaling video foundation models for multimodal video understanding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.037738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.153989Z digest=sha256:f4193e9175aac734703b0541fce30169eacefd930dcbdad2aa87e30a671ebf9f

Observation 55a81fcf-1c28-4640-b3d8-ce4c2b551f17 · outbound

This paper cites Vivit: A video vision transformer.

Extending Video Masked Autoencoders to 128 frames Vivit: A video vision transformer

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.024691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.157567Z digest=sha256:6066b03d40660b8a817ef9022ef7459d8ed51432219bc69b2e290d9a6ef66563

Observation 07e3e8ce-9ff2-4a61-9d79-e761d97b7bfb · outbound

This paper cites Finite scalar quantization: VQ-V AE made simple.

Extending Video Masked Autoencoders to 128 frames Finite scalar quantization: VQ-V AE made simple

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.011181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.161177Z digest=sha256:908137188fa4e36b8c01f740ca9f33b4495720271effa7a49d1c45294cbfe89e

Observation 02398460-7a6d-40c7-bf63-b314a15cce3f · outbound

This paper cites A Short Note about Kinetics-600.

Extending Video Masked Autoencoders to 128 frames A Short Note about Kinetics-600

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.164843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.164843Z digest=sha256:8d0596e2750763c33dac54ec8bac5c093ff6f5d7ac36114d57cec98930942cd0

Observation 95ecca2f-f416-4873-9e90-80143ff10102 · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

Extending Video Masked Autoencoders to 128 frames A Short Note on the Kinetics-700 Human Action Dataset

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.168857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.168857Z digest=sha256:33914c7360e771e188f7f9ac1678715408da5156fef4d99ccbe4052aff444ac1

Observation a196cfee-b1a4-49b0-b4a2-0d09a82b791d · outbound

This paper cites Multiview transformers for video recognition.

Extending Video Masked Autoencoders to 128 frames Multiview transformers for video recognition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.997209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.172795Z digest=sha256:f0d2a49b327c0c679681e221326e34c512db4ea54cc7b8f2c17b7befb0b77f9a

Observation 33cfc5de-95ae-40d3-8b38-2227648e89d7 · outbound

This paper cites Temporally-Adaptive Models for Efficient Video Understanding.

Extending Video Masked Autoencoders to 128 frames Temporally-Adaptive Models for Efficient Video Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.176362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.176362Z digest=sha256:965a3eece1b4540e9d227fd6dc24d4452f8d8b623678b34b0ca3b4cc25c67960

Observation 2f028e33-ae02-4d0b-8d47-ea889ce8eda2 · outbound

This paper cites Training a Large Video Model on a Single Machine in a Day.

Extending Video Masked Autoencoders to 128 frames Training a Large Video Model on a Single Machine in a Day

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.180917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.180917Z digest=sha256:7319eb5e7b82053fd2dab378c29e480c34fee95ef6329b23ca7a4235e6b06596

Observation 72d5d421-0073-40e7-8db3-f1349385d633 · outbound

This paper cites Imagenet-21k pretraining for the masses.

Extending Video Masked Autoencoders to 128 frames Imagenet-21k pretraining for the masses

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.983302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.185019Z digest=sha256:13282dca364d85d93c5ef2d3efff1c38e63d8145a20df9b655d7318db0d1f66a

Observation 45a79a2a-274c-496e-b7cf-87f33caffe1d · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Extending Video Masked Autoencoders to 128 frames Ego4d: Around the world in 3,000 hours of egocentric video

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.972248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.188571Z digest=sha256:a24e1754e698c702c4e74b9ab80859e887a7e4e9cbe949818cd3d8034d4b11a3

Observation bdb7e0cb-169d-4786-bfcf-e6475c709066 · outbound

This paper cites Verbs in action: Improving verb understanding in video-language models.

Extending Video Masked Autoencoders to 128 frames Verbs in action: Improving verb understanding in video-language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.961247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.192158Z digest=sha256:a78d92920140440519d6cf6f44918ccd15ca625f8e26997d119917a8b0defe2f

Observation ec83fec6-7af5-4875-aefd-eb8f725487be · outbound

This paper cites Slowfast networks for video recognition.

Extending Video Masked Autoencoders to 128 frames Slowfast networks for video recognition

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.949510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.196373Z digest=sha256:a10fe4f9cc187285ab063fcd0c76beb031ecc8965027f1e0f7f42c9cc0cabe56

Observation 24b0f31e-184e-43ca-9067-4636c6be323f · outbound

This paper cites Interactive prototype learning for egocentric action recognition.

Extending Video Masked Autoencoders to 128 frames Interactive prototype learning for egocentric action recognition

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.939197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.200298Z digest=sha256:8fb745bdceadf5066cfd538f8a05be0da3b7ab1b538c61b8adf0bc3a84c26518

Observation 0be7e6c2-5f50-4e9f-aa24-cd85927732bb · outbound

This paper cites Movinets: Mobile video networks for efficient video recognition.

Extending Video Masked Autoencoders to 128 frames Movinets: Mobile video networks for efficient video recognition

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.928022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.203795Z digest=sha256:505733ba955a4a7d93153e6d98dd0746266bf6b2eb61d2d16dc9d71b8f6b5440

Observation a08bd723-c388-4787-9478-9e4cdde59e2d · outbound

This paper cites Omnivore: A single model for many visual modalities.

Extending Video Masked Autoencoders to 128 frames Omnivore: A single model for many visual modalities

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.916079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.207333Z digest=sha256:4d1125e0a9ac688c0efb6ada460c838c72918a2e3fbaa0d283f8bf40fd2bdaba

Observation 08528d82-d877-41a5-ba58-e4c47280dda9 · outbound

This paper cites Learning video representations from large language models.

Extending Video Masked Autoencoders to 128 frames Learning video representations from large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.903493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.211110Z digest=sha256:d6110a548d14ce12966487633ee7b46d29fe9c62708cb81060d8f2b2ba010dab

Observation 994e1ecd-afde-4267-9eff-b2a21a9e9a74 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, 2021.

Extending Video Masked Autoencoders to 128 frames Is space-time attention all you need for video understanding? In ICML, 2021

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.891843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.214507Z digest=sha256:9fbcacd0c9525579ae96a25eccc908176ac31fb468434dfbc38cedcd1e39f842

Observation 62167356-5e55-4fb7-8620-e87d6e8f161f · outbound

This paper cites Video swin transformer.

Extending Video Masked Autoencoders to 128 frames Video swin transformer

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.879718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.218269Z digest=sha256:f8acd578b93b53b3791ff38945b965aeff31162b587e418d9a012a8a6309859e

Observation be25e975-2432-4536-8ce7-8d7427f0d842 · outbound

This paper cites Can an image classifier suffice for action recognition? In ICLR, 2022.

Extending Video Masked Autoencoders to 128 frames Can an image classifier suffice for action recognition? In ICLR, 2022

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.868902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.221993Z digest=sha256:a46b3b3eb273f1a38d14e01f10d3ef1704dabfc537f21f282b4c76f2f15cf2e8

Observation 5efa996f-7043-42bb-bb20-13945d3dc966 · outbound

This paper cites Object-region video transformers.

Extending Video Masked Autoencoders to 128 frames Object-region video transformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.857839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.225841Z digest=sha256:f1cb17e5f57f8cb7b45767316d9eeaeeab1a0dadd0534908139f84c752f0fbf7

Observation 738d1098-9f5d-4766-8627-53e4d24a5d64 · outbound

This paper cites Aim: Adapting image models for efficient video action recognition.

Extending Video Masked Autoencoders to 128 frames Aim: Adapting image models for efficient video action recognition

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.845804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.229883Z digest=sha256:dc79fc31e11798629f066e2d8b6b42643adba52886a1f42bc53788ee2020acaa

Observation 26189b8b-33f3-4750-9990-c658006ca20a · outbound

This paper cites Video-focalnets: Spatio-temporal focal modulation for video action recognition.

Extending Video Masked Autoencoders to 128 frames Video-focalnets: Spatio-temporal focal modulation for video action recognition

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.834222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.233575Z digest=sha256:23d101aebe152ea896a5c267b2f4b1ecc7da594e028c5a97d4a1195f6582dd50

Observation 6d0e83f9-ab6a-4b0e-910c-dc8b000828e3 · outbound

This paper cites Language model beats diffusion–tokenizer is key to visual generation.

Extending Video Masked Autoencoders to 128 frames Language model beats diffusion–tokenizer is key to visual generation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.819649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.237643Z digest=sha256:aac40e7ae7a6e685d7e91c25d78173e747ed846bfa6d3c0b98a05669ae8b5eee

Observation 35bf2940-3183-4085-8371-02c89b5ee713 · outbound

This paper cites Finegym: A hierarchical video dataset for fine-grained action understanding.

Extending Video Masked Autoencoders to 128 frames Finegym: A hierarchical video dataset for fine-grained action understanding

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.804450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.241462Z digest=sha256:023693e11ba528e7bfe8ad8ad2995e68d239cd1fc37289d921e07860cdf89210

Observation c0e06bc4-9991-4702-8521-b65e8aa77aa3 · outbound

This paper cites Learning temporal cues for fine-grained action recognition.

Extending Video Masked Autoencoders to 128 frames Learning temporal cues for fine-grained action recognition

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.790238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.245443Z digest=sha256:5096c08443e92295c2d600bff1cd07b0478108014d5b12f6e3da7eaff6eaa926

Observation b2aa5f83-862d-4df6-b1f7-ebe7564d0c09 · outbound

This paper cites Tsm: Temporal shift module for efficient video understanding.

Extending Video Masked Autoencoders to 128 frames Tsm: Temporal shift module for efficient video understanding

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.775928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.249266Z digest=sha256:a7c3679d00685f836ac797e6f1f77db4643eaa0cf4907989a4f08340c5209aa0

Observation 33ddac3d-e7d2-4c61-85b0-c7df2c2651a0 · outbound

This paper cites Temporal query networks for fine-grained video understanding.

Extending Video Masked Autoencoders to 128 frames Temporal query networks for fine-grained video understanding

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.762566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.253411Z digest=sha256:961aa307879424a62fbaa1e1ee6a3026d04c08b3c7f466bd4db75726d0dc4e97

Observation d5c37d89-8b6f-4e04-b893-76292cb9fbaf · outbound

This paper cites Combined cnn transformer encoder for enhanced fine-grained human action recognition.

Extending Video Masked Autoencoders to 128 frames Combined cnn transformer encoder for enhanced fine-grained human action recognition

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.751023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.257990Z digest=sha256:cd1b4b6a14478619d097627f65a785fcb80e784813e8b62acb71fe1134ff5dd8

Observation c1217cf0-e260-4354-a23a-649cf1548550 · outbound

This paper cites Going deeper with image transformers.

Extending Video Masked Autoencoders to 128 frames Going deeper with image transformers

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.738041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.262106Z digest=sha256:945d7658b2ff74a62e7d13c200c90621852b8d808a4e933b8e14fdc44e33d71f

Observation e4c247d4-831e-46c6-a25f-64cfd87e3220 · outbound

This paper cites Scenic: A jax library for computer vision research and beyond.

Extending Video Masked Autoencoders to 128 frames Scenic: A jax library for computer vision research and beyond

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.722729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.266075Z digest=sha256:72e015da30e81582da1c2783aad54bf8fd530f648e7179f96403793158b3c87b

Observation 0a74a736-5133-4c9c-a196-f891b24fb370 · outbound

This paper cites We found this slightly improved accuracy ( +0.5 points on EPIC-Kitchens-100 Verbs and +1.2 points on Diving48).

Extending Video Masked Autoencoders to 128 frames We found this slightly improved accuracy ( +0.5 points on EPIC-Kitchens-100 Verbs and +1.2 points on Diving48)

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.700975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.270584Z digest=sha256:b91a96acf760b66afd9130095b0aa87621f2bdf7d46bc2897c566e3bb26f4b1e

Observation e8883ef5-e7c8-4e61-a54c-8143413ce151 · outbound

This paper cites an unresolved cited work.

Extending Video Masked Autoencoders to 128 frames Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:22:25.681314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.274933Z digest=sha256:5adf0060aac761bde71aeb498fbc71ac4d45838e5c0aead70448343448caf4d7

Observation 7121ac07-0f58-47f2-ad01-88e579276083 · outbound

This paper cites 17 Table 9: Model size vs frames.

Extending Video Masked Autoencoders to 128 frames 17 Table 9: Model size vs frames

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.667471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.279010Z digest=sha256:24fe7fe80755c2a5263a2df97d80d0985ba0cebb4b6ef192aa9767d56a7fe21d

Observation 13270ed7-aa41-42bd-a1da-25f5ef192c76 · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.653781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.283416Z digest=sha256:c57f914076f30c6d72bbbcf39a427e1e1ff9b205e2150508975cd3860fad1998

Observation fd7f5d79-0999-4578-ba91-897e965c755c · outbound

This paper cites Limitations.

Extending Video Masked Autoencoders to 128 frames Limitations

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.636688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.287958Z digest=sha256:90476650d61352fb15c3fb9b5970e643ff9c8046f3fefba09e517ec6a1e46e8a

Observation 10456fc6-93cc-44fe-982e-3888d20813f1 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.619738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.292611Z digest=sha256:534dd554560f36f8fc2bd4d8ae1a3771f255fb7058a24d34f1643dfb6704c029

Observation 955057ce-da5a-4eab-b8a1-3bb1e743a1f6 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include experiments

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.605281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.296935Z digest=sha256:3561ab6f5b7a24e9f7f81c96c9401e9f2a1baa884c322ecd0e0c0b1b060d98a5

Observation 05b446cf-6eb3-4330-affd-25d7dd01dfe0 · outbound

This paper cites Guidelines: • The answer NA means that paper does not include experiments requiring code.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that paper does not include experiments requiring code

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.588626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.302631Z digest=sha256:21f2866c549276918ab4bca3e78dd0af47f38fbf878f52c39066eae11a209721

Observation 903717b9-95c5-4a0b-ab71-e54f72805d1b · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include experiments

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.572605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.307592Z digest=sha256:be70d19105d0986f3daf28567554fc6bc46db9ede13c415b597e9130845620ef

Observation 7504251a-762c-41d2-ad08-4e797b83ce0a · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include experiments

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.559349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.312618Z digest=sha256:c5dcc32b472d1b68135a2432302f7f0cbabffe0e3d7388f35fd413b1c9b3e362

Observation e409290e-ad8e-4ba9-95c5-8ff826b82976 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include experiments

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.543694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.318955Z digest=sha256:25d3ad546fa86299380f142eb8f4958f89d509ec2ea8da4e1ea4feef6049f37a

Observation 151092d4-fe23-4c42-8e99-94dce5af1702 · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.528011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.323122Z digest=sha256:be920874fcac605c8b4be7c3eb82091f8fd063baef6aa3134bc8f9c64d9c05f4

Observation 27550e91-dc95-4ad8-834c-1d3ca347c615 · outbound

This paper cites • If the authors answer NA or No, they should explain why their work has no societal impact or why the paper does not address societal impact.

Extending Video Masked Autoencoders to 128 frames • If the authors answer NA or No, they should explain why their work has no societal impact or why the paper does not address societal impact

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.510829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.327717Z digest=sha256:43f2eabe5b9d27554c5c32112615500236cfbf1326f29d71f0dd0306c75ef198

Observation 1832e7fb-510a-4c78-80da-e7490525eabf · outbound

This paper cites Guidelines: • The answer NA means that the paper poses no such risks.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper poses no such risks

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.497541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.332877Z digest=sha256:910e18decee886babcd8c12f07a264363344a8ff09281855a25da7db66e13362

Observation e87e9080-6d21-4d13-8e92-c9dc39889c6e · outbound

This paper cites Guidelines: • The answer NA means that the paper does not use existing assets.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not use existing assets

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.484090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.343151Z digest=sha256:6f056666e1423123292255d2b9002a7012d908cbbd1fc05d01a78a2f0155b7df

Observation 0f1ef82c-5084-4726-b5f5-3e7278d9683a · outbound

This paper cites Guidelines: • The answer NA means that the paper does not release new assets.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not release new assets

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.347041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.347041Z digest=sha256:c04eaea5ca6aaf2445b9fb2fc173c82e1998fc904f0c3c6aeb250bb511404f57

Observation 32ea79b1-a7bb-4d37-9d42-eee4bc252c05 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.464389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.351339Z digest=sha256:a3e07ff85c049250ec37e3f323658ad93d39f5270f9950745ffbb0d45830f3b8

Observation 1f04bade-1819-48ac-ab45-49e80ce8e8d8 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.450458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.355326Z digest=sha256:53dd29da13f867f758bce5635a0763e601f09df0b862c240d2f73ab0f2bb6177

Observation 17e2c672-1e70-46a4-a07b-68f28e108efe · outbound

This paper cites an unresolved cited work.

Extending Video Masked Autoencoders to 128 frames Unresolved cited work

Reference 2023

Resolution
parse uncertain
raw_fallback, observed 2026-08-12T16:22:26.154889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:22:25.119455Z digest=sha256:a18182c9e6e6aa81fc39c4ce3b241cf829f446c9b314d21aced8cb971fef05f2

Pith citing papers

No inbound Pith citation observations are available.