Pith. sign in

Paper Citation Record · LEDGER

Extending Video Masked Autoencoders to 128 frames

As of 12 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 0 inbound Pith citation observations for arXiv:2411.13683.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13683 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:22:25.355326Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

82 of 82 outbound references displayed

  • verified exact0
  • verified fuzzy71
  • unresolved10
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 84cf1d58-7cf6-4e59-b24a-18365b0499be · outbound

This paper cites Towards long-form video understanding.

Extending Video Masked Autoencoders to 128 frames Towards long-form video understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.014117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.014117Z digest=sha256:7dbc108c4f2fa632e4153b8f5d3de178008188fa215343bbaa84a21c34846761

Observation 57d954e0-e57e-4b91-834e-9520d8a00794 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

Extending Video Masked Autoencoders to 128 frames Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.018939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.018939Z digest=sha256:0afc2750ba35ac3fb72463814509b74d6ec594c709a0e2eac1948b5f17ff580f

Observation 1e2090b2-dcc2-4824-91ac-5108d39886e7 · outbound

This paper cites Long movie clip classification with state-space video models.

Extending Video Masked Autoencoders to 128 frames Long movie clip classification with state-space video models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.442059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.023182Z digest=sha256:e7ef605b333e30d8ab1f09de7544664192eaa3bbe84eb2a0c3c44e27db3b839a

Observation 3abf5ab4-b5ec-4dfe-a7e7-e0e7e9f6434f · outbound

This paper cites Selective structured state-spaces for long-form video understanding.

Extending Video Masked Autoencoders to 128 frames Selective structured state-spaces for long-form video understanding

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.428871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.027369Z digest=sha256:4eed3de9aefb3d9dae76797f28c84ea6aaef4267da9870c11f920779419c126f

Observation e1d2eb4b-a9ef-4679-8eb6-ccfb70df90db · outbound

This paper cites Memory consolidation enables long-context video understanding.

Extending Video Masked Autoencoders to 128 frames Memory consolidation enables long-context video understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.415408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.031425Z digest=sha256:e13fb15c0237816c0589b3bbd0ab832d32cf2a0ae656e0f666e0c0d8de16bf77

Observation fad0160e-a622-46f0-8cbe-1ba5c085f687 · outbound

This paper cites Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition.

Extending Video Masked Autoencoders to 128 frames Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.401384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.036030Z digest=sha256:743ba8aeca013c59eb5b279d03513f771737c143c1a3a7ca74c1d42119420125

Observation 9136e534-9098-4a59-9d9b-2b43c8578db5 · outbound

This paper cites Token turing machines.

Extending Video Masked Autoencoders to 128 frames Token turing machines

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.386697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.040506Z digest=sha256:739d79118ca124e66725d114fc1d384840a3bb8ee716f213cd944ca5a9f90382

Observation 1ee8d817-2f90-4da5-b5ff-b16c8255302b · outbound

This paper cites Video recap: Recursive captioning of hour-long videos.

Extending Video Masked Autoencoders to 128 frames Video recap: Recursive captioning of hour-long videos

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.374304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.045808Z digest=sha256:5bbe8a0ed994c25ba092bd4a9c1f166567a220437884008311c35e45407f609d

Observation 15282ec6-ac03-48a8-bb04-1f898e7a3bc4 · outbound

This paper cites A simple recipe for contrastively pre-training video-first encoders beyond 16 frames.

Extending Video Masked Autoencoders to 128 frames A simple recipe for contrastively pre-training video-first encoders beyond 16 frames

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.360794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.050401Z digest=sha256:9586c74f8297c89ef053fca26d08a9d069f7f223de62c53925331a8ae9c35d94

Observation 96c04c42-cb02-4ce3-a8e5-aba75807fbeb · outbound

This paper cites A simple llm framework for long-range video question-answering.

Extending Video Masked Autoencoders to 128 frames A simple llm framework for long-range video question-answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.054724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.054724Z digest=sha256:d3ba65a5ae1c32acb450be406a54dbf83d21cc78d0a11e2d9da3dc42740c4dbf

Observation b28fee82-238e-451f-a888-05e92652f15a · outbound

This paper cites Long-form video- language pre-training with multimodal temporal contrastive learning.

Extending Video Masked Autoencoders to 128 frames Long-form video- language pre-training with multimodal temporal contrastive learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.341094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.058702Z digest=sha256:ebb75945d8be89c70b0d4eb103d953716fd6f31bc8dcda9e268eeb3999b5b48c

Observation 4dc4e05a-4e81-45f9-ba36-ef5bdde63ccb · outbound

This paper cites Koala: Key frame-conditioned long video-llm.

Extending Video Masked Autoencoders to 128 frames Koala: Key frame-conditioned long video-llm

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.328284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.062400Z digest=sha256:92078be20c6b997061ea81197560a7eab3fe34e264ce702483ecc491e144cb51

Observation 9104b8a1-ea81-4e5d-b003-51f7c116eb30 · outbound

This paper cites Masked autoencoders as spatiotemporal learners.

Extending Video Masked Autoencoders to 128 frames Masked autoencoders as spatiotemporal learners

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.316949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.066391Z digest=sha256:08a7ea179dcbe82673b07a645f0cdb14164ecf5b3c0ebc5a052317c5e0f077b4

Observation 30a10d9c-6a53-456b-9d89-fc6543a31d37 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

Extending Video Masked Autoencoders to 128 frames Videomae v2: Scaling video masked autoencoders with dual masking

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.304662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.071077Z digest=sha256:0eb922f607073d970b1dc15b6e0be11d2351d1b4f084bc8559ffb4165a918c20

Observation 61aa62bd-3850-4120-ade0-88283e9a3003 · outbound

This paper cites VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Extending Video Masked Autoencoders to 128 frames VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.290950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.074927Z digest=sha256:29a2284a7be74ae226695b5ff6462ec35b88457dd5564ccd4cd31faea0918213

Observation cdd82006-37e8-4192-a1fe-5725f23d39cc · outbound

This paper cites How can objects help action recognition? In CVPR, 2023.

Extending Video Masked Autoencoders to 128 frames How can objects help action recognition? In CVPR, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.278784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.078492Z digest=sha256:9d914b0cd039500763f0db44fce857ce99b437af89c5daa40747de8684f65541

Observation 2d832d9b-485e-4aca-8091-c506b8885877 · outbound

This paper cites Everest: Efficient masked video autoencoder by removing redundant spatiotemporal tokens.

Extending Video Masked Autoencoders to 128 frames Everest: Efficient masked video autoencoder by removing redundant spatiotemporal tokens

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.266111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.082930Z digest=sha256:19d3ebc58c97360c34e65e7d8682c1a7e9db72429c77e65ce23fc5a791d15cc5

Observation 9d1350a1-3813-4f80-bbf4-6a55f4baba20 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100.

Extending Video Masked Autoencoders to 128 frames Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.253869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.086594Z digest=sha256:dc89e4fe32e172aab2160ff81b5a03218fde9d0a7002c3f62f0a5da22052ad03

Observation 6c0d1bea-e94f-46af-ab04-b4c051bd18c1 · outbound

This paper cites Resound: Towards action recognition without representation bias.

Extending Video Masked Autoencoders to 128 frames Resound: Towards action recognition without representation bias

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.241287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.090422Z digest=sha256:878ea058858635548cad6d314019754f495efd949ea1e8cdacc2ce16a878915d

Observation 21f1223f-011f-4832-ade6-2ee718f0afb4 · outbound

This paper cites an unresolved cited work.

Extending Video Masked Autoencoders to 128 frames Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:22:26.229238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.094637Z digest=sha256:91250449a63e0b6479b1807c9e5618f0cabc5405facb2e2a89b8da58578ed9e6

Observation 05497a92-60ff-443e-9f04-71f253a54044 · outbound

This paper cites Bevt: Bert pretraining of video transformers.

Extending Video Masked Autoencoders to 128 frames Bevt: Bert pretraining of video transformers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.218024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.098703Z digest=sha256:a5b098710ab0bac235ab670b4252f82d65c6c7c058afadc614d63cf9f1e1ba7a

Observation e5e7babb-2187-497a-af9d-76469a1373b0 · outbound

This paper cites Girdhar, A.

Extending Video Masked Autoencoders to 128 frames Girdhar, A

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.205639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.103922Z digest=sha256:093dac02808b34c0685650fc66dbdf5aa568445a16b8b8ed61fdca27830e9f98

Observation 713e89ef-4263-417a-b095-74a19c92e633 · outbound

This paper cites Zero-shot text-to-image generation.

Extending Video Masked Autoencoders to 128 frames Zero-shot text-to-image generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.193574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.107960Z digest=sha256:19f1d413ccc1f29f024ddbe7e797ed8e09729bc2087521d5e47f02570143ffdf

Observation bc7e4fd2-2979-4311-a58b-b08529e687ad · outbound

This paper cites Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning.

Extending Video Masked Autoencoders to 128 frames Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.182211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.111942Z digest=sha256:fcfa7b874fd47bc5cd2b79482a6501c258a64511ad81a46159a457426534b1ac

Observation e5511110-75f1-465e-bd26-2f2d846c7286 · outbound

This paper cites Magvit: Masked generative video transformer.

Extending Video Masked Autoencoders to 128 frames Magvit: Masked generative video transformer

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.170158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.115504Z digest=sha256:46867cc501de3b3665ddf61b1737a26d85b61181c3869084d39b5057b0d34b85

Observation 929db5b8-04b5-4e3d-a6cc-91d48a12431c · outbound

This paper cites Tokenlearner: What can 8 learned tokens do for images and videos? In NeurIPS, 2021.

Extending Video Masked Autoencoders to 128 frames Tokenlearner: What can 8 learned tokens do for images and videos? In NeurIPS, 2021

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.143450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.123089Z digest=sha256:53cb1356e8996a702e3259a1f39e9d76f95ee73d99e5f5e6faaefa0a323ad174

Observation 9cd15a53-f043-4a68-9079-f5e6871c87d9 · outbound

This paper cites something something.

Extending Video Masked Autoencoders to 128 frames something something

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.130927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.126787Z digest=sha256:6b05e2d53fb23c533e180ee19f930577c099bbd28dfda20ec14dcad1831d9950

Observation 319a40db-cab6-42ab-86cd-b86c83095958 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Extending Video Masked Autoencoders to 128 frames An image is worth 16x16 words: Transformers for image recognition at scale

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.118986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.130449Z digest=sha256:3015f2709120661b12e4b293d955f6e9197023ea421eea2e6ac27ebda94c94ab

Observation d87ef826-a40a-4764-958d-63623af7b0cc · outbound

This paper cites Mgmae: Motion guided masking for video masked autoencoding.

Extending Video Masked Autoencoders to 128 frames Mgmae: Motion guided masking for video masked autoencoding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.105953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.134329Z digest=sha256:acaa91438654b1ae448fabb7d91703160db87425167a92360b0e07489f225f43

Observation b5b0e6f3-fc2d-408f-a232-3ca3c19afbf2 · outbound

This paper cites Motion-guided masking for spatiotemporal representation learning.

Extending Video Masked Autoencoders to 128 frames Motion-guided masking for spatiotemporal representation learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.091040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.138258Z digest=sha256:f07b3e9eb367c4b70f09cbfe54accb8a6b74429f7572de92b065e042f8d20c19

Observation e32096bb-4f64-4512-ac99-52032f45532d · outbound

This paper cites Video codec design: developing image and video compression systems.

Extending Video Masked Autoencoders to 128 frames Video codec design: developing image and video compression systems

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.077412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.142170Z digest=sha256:9b3c4078cb4c3218f185d7f44bd0f365a76b2a420e00d62d67072bcb50076bb8

Observation 9a8ae6a4-ee82-4266-bced-f94d9ea68a02 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

Extending Video Masked Autoencoders to 128 frames Raft: Recurrent all-pairs field transforms for optical flow

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.062271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.146065Z digest=sha256:40434ed50936b3a07d8a3728b3d5d9994b247c4f4cc0868b0f4b562e373f1b6c

Observation d71d0fc2-7964-48d3-9875-5cb23baf2090 · outbound

This paper cites Videoprism: A foundational visual encoder for video understanding.

Extending Video Masked Autoencoders to 128 frames Videoprism: A foundational visual encoder for video understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.050812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.149577Z digest=sha256:ad77ab146d9738bb419b922dc499cbb3aca04aa2d0f25f1feaae3be5811bb987

Observation 0c81c503-ca16-426b-8f61-e24d49b7afd8 · outbound

This paper cites Internvideo2: Scaling video foundation models for multimodal video understanding.

Extending Video Masked Autoencoders to 128 frames Internvideo2: Scaling video foundation models for multimodal video understanding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.037738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.153989Z digest=sha256:e7376d812c9c5ba38c084946c625837e43a0342c9d1389100f0e80793ab581cf

Observation 55a81fcf-1c28-4640-b3d8-ce4c2b551f17 · outbound

This paper cites Vivit: A video vision transformer.

Extending Video Masked Autoencoders to 128 frames Vivit: A video vision transformer

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.024691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.157567Z digest=sha256:f7d15aff8c81445f03173352ac305e4be2d86a8203fd4c0b2bf589cd4b5268bc

Observation 07e3e8ce-9ff2-4a61-9d79-e761d97b7bfb · outbound

This paper cites Finite scalar quantization: VQ-V AE made simple.

Extending Video Masked Autoencoders to 128 frames Finite scalar quantization: VQ-V AE made simple

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.011181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.161177Z digest=sha256:2da3759e73c952cec2ec10cdbc80c52e1b5f32239793942f2ac1f522ac05ccd4

Observation 02398460-7a6d-40c7-bf63-b314a15cce3f · outbound

This paper cites A Short Note about Kinetics-600.

Extending Video Masked Autoencoders to 128 frames A Short Note about Kinetics-600

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.164843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.164843Z digest=sha256:41a5db2096c533b15a68fcdf8462d07fa0a66eb53863eb77d47fe5488f02bfda

Observation 95ecca2f-f416-4873-9e90-80143ff10102 · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

Extending Video Masked Autoencoders to 128 frames A Short Note on the Kinetics-700 Human Action Dataset

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.168857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.168857Z digest=sha256:f305272be26837037e72fab8f7800f5ffaedd2a8185a5f591409b042fce57207

Observation a196cfee-b1a4-49b0-b4a2-0d09a82b791d · outbound

This paper cites Multiview transformers for video recognition.

Extending Video Masked Autoencoders to 128 frames Multiview transformers for video recognition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.997209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.172795Z digest=sha256:cd3090229baf306a0d12976459ca39cd82441e5d6c3139ae58133efa2a15dcb1

Observation 33cfc5de-95ae-40d3-8b38-2227648e89d7 · outbound

This paper cites Temporally-Adaptive Models for Efficient Video Understanding.

Extending Video Masked Autoencoders to 128 frames Temporally-Adaptive Models for Efficient Video Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.176362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.176362Z digest=sha256:45c4524b4727bee59bc3a00714ff4d0edb71f4025b7ce83f774d3df5df009805

Observation 2f028e33-ae02-4d0b-8d47-ea889ce8eda2 · outbound

This paper cites Training a Large Video Model on a Single Machine in a Day.

Extending Video Masked Autoencoders to 128 frames Training a Large Video Model on a Single Machine in a Day

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.180917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.180917Z digest=sha256:c241c59a8adfd7ba5cd07790f5a94a212fe5d4012e90c14012da5b8d7e9b3bcf

Observation 72d5d421-0073-40e7-8db3-f1349385d633 · outbound

This paper cites Imagenet-21k pretraining for the masses.

Extending Video Masked Autoencoders to 128 frames Imagenet-21k pretraining for the masses

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.983302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.185019Z digest=sha256:9b2bfc6b48dde279d9eec1fcbeb20fcd6e6b43c4023303e2a894980febedbdfc

Observation 45a79a2a-274c-496e-b7cf-87f33caffe1d · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Extending Video Masked Autoencoders to 128 frames Ego4d: Around the world in 3,000 hours of egocentric video

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.972248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.188571Z digest=sha256:be2e9792d04e3d518322234d9eb1b4b694eecb3dfcb118bb6a0efc54e2269430

Observation bdb7e0cb-169d-4786-bfcf-e6475c709066 · outbound

This paper cites Verbs in action: Improving verb understanding in video-language models.

Extending Video Masked Autoencoders to 128 frames Verbs in action: Improving verb understanding in video-language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.961247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.192158Z digest=sha256:6aaaaebc7e09288d53361766ef0836b99656c684f81784e86cf85c3293f91e39

Observation ec83fec6-7af5-4875-aefd-eb8f725487be · outbound

This paper cites Slowfast networks for video recognition.

Extending Video Masked Autoencoders to 128 frames Slowfast networks for video recognition

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.949510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.196373Z digest=sha256:1b788ba275d156553eaf57a89eedfa9a993d0152466266f8d0880e2d6e7a00d2

Observation 24b0f31e-184e-43ca-9067-4636c6be323f · outbound

This paper cites Interactive prototype learning for egocentric action recognition.

Extending Video Masked Autoencoders to 128 frames Interactive prototype learning for egocentric action recognition

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.939197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.200298Z digest=sha256:107983446bd1bc1bb1a759b8e66db48dc2a11f5337c9025dd6d1b17d335b3a31

Observation 0be7e6c2-5f50-4e9f-aa24-cd85927732bb · outbound

This paper cites Movinets: Mobile video networks for efficient video recognition.

Extending Video Masked Autoencoders to 128 frames Movinets: Mobile video networks for efficient video recognition

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.928022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.203795Z digest=sha256:414a310e2f9418d116da3cc260b3f8d29ec97d1072eefd07852f77094568516f

Observation a08bd723-c388-4787-9478-9e4cdde59e2d · outbound

This paper cites Omnivore: A single model for many visual modalities.

Extending Video Masked Autoencoders to 128 frames Omnivore: A single model for many visual modalities

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.916079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.207333Z digest=sha256:44d95dbd4ce31000e3e4a0ceda15464c59e47dbaec81fa5a7d9fcb01e2e0526b

Observation 08528d82-d877-41a5-ba58-e4c47280dda9 · outbound

This paper cites Learning video representations from large language models.

Extending Video Masked Autoencoders to 128 frames Learning video representations from large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.903493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.211110Z digest=sha256:b045b4c2b057fa664023e1262b09ee3d29ded27164775a56eb3a4fb31b05646c

Observation 994e1ecd-afde-4267-9eff-b2a21a9e9a74 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, 2021.

Extending Video Masked Autoencoders to 128 frames Is space-time attention all you need for video understanding? In ICML, 2021

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.891843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.214507Z digest=sha256:74e525c94469eeac78bf792e0a2c9c598899283cbd1c954bf003518ad2a6654a

Observation 62167356-5e55-4fb7-8620-e87d6e8f161f · outbound

This paper cites Video swin transformer.

Extending Video Masked Autoencoders to 128 frames Video swin transformer

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.879718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.218269Z digest=sha256:61904f585afb1019635a65bd11b2a8a83b9314e7d0f0b464370ff1bfea8dc153

Observation be25e975-2432-4536-8ce7-8d7427f0d842 · outbound

This paper cites Can an image classifier suffice for action recognition? In ICLR, 2022.

Extending Video Masked Autoencoders to 128 frames Can an image classifier suffice for action recognition? In ICLR, 2022

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.868902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.221993Z digest=sha256:a9f1d8e08573fb09e62c2d30300a747b7fe3c0bcd4462e66c798be79be41e478

Observation 5efa996f-7043-42bb-bb20-13945d3dc966 · outbound

This paper cites Object-region video transformers.

Extending Video Masked Autoencoders to 128 frames Object-region video transformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.857839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.225841Z digest=sha256:e5c62ebec7482459662639d5fd83cd15ddcadbcb3f6e90023a9a4914ca1a0687

Observation 738d1098-9f5d-4766-8627-53e4d24a5d64 · outbound

This paper cites Aim: Adapting image models for efficient video action recognition.

Extending Video Masked Autoencoders to 128 frames Aim: Adapting image models for efficient video action recognition

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.845804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.229883Z digest=sha256:5340a275f38f6fcf5cdf197bdf985c59d9edf4f6f11481f302710149bd18e0fc

Observation 26189b8b-33f3-4750-9990-c658006ca20a · outbound

This paper cites Video-focalnets: Spatio-temporal focal modulation for video action recognition.

Extending Video Masked Autoencoders to 128 frames Video-focalnets: Spatio-temporal focal modulation for video action recognition

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.834222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.233575Z digest=sha256:ba354ab71114825b43ca8a375e6acdb0176fb4ee12fae3ce657129df7054ccf1

Observation 6d0e83f9-ab6a-4b0e-910c-dc8b000828e3 · outbound

This paper cites Language model beats diffusion–tokenizer is key to visual generation.

Extending Video Masked Autoencoders to 128 frames Language model beats diffusion–tokenizer is key to visual generation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.819649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.237643Z digest=sha256:94ab96816d6575df8ab3715d4edc980f8007dbed1429db00a29677bc0cc82615

Observation 35bf2940-3183-4085-8371-02c89b5ee713 · outbound

This paper cites Finegym: A hierarchical video dataset for fine-grained action understanding.

Extending Video Masked Autoencoders to 128 frames Finegym: A hierarchical video dataset for fine-grained action understanding

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.804450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.241462Z digest=sha256:339560573c077abff5acbf0976df7caa7829c0d56423b4983c38640904cdcd67

Observation c0e06bc4-9991-4702-8521-b65e8aa77aa3 · outbound

This paper cites Learning temporal cues for fine-grained action recognition.

Extending Video Masked Autoencoders to 128 frames Learning temporal cues for fine-grained action recognition

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.790238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.245443Z digest=sha256:8c10355527655f5f0dd7d5edbbcea7bff4586de30faaee36627df34fa8a97a68

Observation b2aa5f83-862d-4df6-b1f7-ebe7564d0c09 · outbound

This paper cites Tsm: Temporal shift module for efficient video understanding.

Extending Video Masked Autoencoders to 128 frames Tsm: Temporal shift module for efficient video understanding

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.775928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.249266Z digest=sha256:6c6f1dad622bd119d56c0d9efc45f02efab37e277e2e2fbba74f8c6ead6f102a

Observation 33ddac3d-e7d2-4c61-85b0-c7df2c2651a0 · outbound

This paper cites Temporal query networks for fine-grained video understanding.

Extending Video Masked Autoencoders to 128 frames Temporal query networks for fine-grained video understanding

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.762566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.253411Z digest=sha256:5542ced59933e6c887dc0947d44fee99e08c6e10196985c7f854b3ff4e2372d3

Observation d5c37d89-8b6f-4e04-b893-76292cb9fbaf · outbound

This paper cites Combined cnn transformer encoder for enhanced fine-grained human action recognition.

Extending Video Masked Autoencoders to 128 frames Combined cnn transformer encoder for enhanced fine-grained human action recognition

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.751023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.257990Z digest=sha256:6f9886d3689b59c67b76cf203da9f0f96af0ff6beeb2d1609f742441470c4868

Observation c1217cf0-e260-4354-a23a-649cf1548550 · outbound

This paper cites Going deeper with image transformers.

Extending Video Masked Autoencoders to 128 frames Going deeper with image transformers

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.738041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.262106Z digest=sha256:c11cc97dc43c6f7409082ce72561b6221022654ac1f42ccb5889f6aece8b28fb

Observation e4c247d4-831e-46c6-a25f-64cfd87e3220 · outbound

This paper cites Scenic: A jax library for computer vision research and beyond.

Extending Video Masked Autoencoders to 128 frames Scenic: A jax library for computer vision research and beyond

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.722729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.266075Z digest=sha256:f55ee67a83fa275d21789bc5d8e9ccb2bfd37b198d33247368c3d444b046cd5c

Observation 0a74a736-5133-4c9c-a196-f891b24fb370 · outbound

This paper cites We found this slightly improved accuracy ( +0.5 points on EPIC-Kitchens-100 Verbs and +1.2 points on Diving48).

Extending Video Masked Autoencoders to 128 frames We found this slightly improved accuracy ( +0.5 points on EPIC-Kitchens-100 Verbs and +1.2 points on Diving48)

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.700975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.270584Z digest=sha256:15dfa97420f51d56d47558cadcf14833f9b603e9ec43a9b6552a0824d1a26a1b

Observation e8883ef5-e7c8-4e61-a54c-8143413ce151 · outbound

This paper cites an unresolved cited work.

Extending Video Masked Autoencoders to 128 frames Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:22:25.681314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.274933Z digest=sha256:08d05bf2f46aeab61ca6ab7652a8ca8a88a5fa860650f4499b0a75b1f4546f8c

Observation 7121ac07-0f58-47f2-ad01-88e579276083 · outbound

This paper cites 17 Table 9: Model size vs frames.

Extending Video Masked Autoencoders to 128 frames 17 Table 9: Model size vs frames

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.667471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.279010Z digest=sha256:cf990f5ea96ceac52d59dc7095356736d444b8bb8714c1121526f78fa3593c94

Observation 13270ed7-aa41-42bd-a1da-25f5ef192c76 · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.653781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.283416Z digest=sha256:615f2fd09a6b20fabbe3057e6bf6d1b4d4a5d16c6c910c8ba884f28b517a0949

Observation fd7f5d79-0999-4578-ba91-897e965c755c · outbound

This paper cites Limitations.

Extending Video Masked Autoencoders to 128 frames Limitations

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.636688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.287958Z digest=sha256:8f5fd342397ed9b028dc2a360969742b2de409e560abf14fb06effacf7307649

Observation 10456fc6-93cc-44fe-982e-3888d20813f1 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.619738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.292611Z digest=sha256:4d11b8f06b3da3691433bacdf388ef66cada7efcc853a4f2be65eea024fc630c

Observation 955057ce-da5a-4eab-b8a1-3bb1e743a1f6 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include experiments

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.605281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.296935Z digest=sha256:2e51c80472837b3213743d8096a8bbb63faccb7f792684348b07aa97af5a295b

Observation 05b446cf-6eb3-4330-affd-25d7dd01dfe0 · outbound

This paper cites Guidelines: • The answer NA means that paper does not include experiments requiring code.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that paper does not include experiments requiring code

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.588626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.302631Z digest=sha256:be7378687065e90e221826a7b1c2e8d37c83748dda5112aaceee3ada5815e161

Observation 903717b9-95c5-4a0b-ab71-e54f72805d1b · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include experiments

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.572605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.307592Z digest=sha256:6a87ba86e11c4957bbe0693aeedebfd12bf0c44cddd87a0e7066ca0e910310bb

Observation 7504251a-762c-41d2-ad08-4e797b83ce0a · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include experiments

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.559349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.312618Z digest=sha256:cf5704c6c944457b0449d14c650cab7a4986209c35b40f9cea4ce1b3ac6572b9

Observation e409290e-ad8e-4ba9-95c5-8ff826b82976 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include experiments

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.543694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.318955Z digest=sha256:7f6c910f7e8f3c7fc085615643506c0f883e36b34cad37b3ab392068ccfd75cd

Observation 151092d4-fe23-4c42-8e99-94dce5af1702 · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.528011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.323122Z digest=sha256:13b97b5f88f0601899d86162e2a8b9da9f27b588e6e07e23e07d446369059d06

Observation 27550e91-dc95-4ad8-834c-1d3ca347c615 · outbound

This paper cites • If the authors answer NA or No, they should explain why their work has no societal impact or why the paper does not address societal impact.

Extending Video Masked Autoencoders to 128 frames • If the authors answer NA or No, they should explain why their work has no societal impact or why the paper does not address societal impact

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.510829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.327717Z digest=sha256:31f95add8e5cc4dcd786a57d53c7419f6707376a67729fd4d8e4f4b707852f43

Observation 1832e7fb-510a-4c78-80da-e7490525eabf · outbound

This paper cites Guidelines: • The answer NA means that the paper poses no such risks.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper poses no such risks

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.497541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.332877Z digest=sha256:b924696458402f27c9fb403e47ea58f268ec3b3db33914ee04c91b2c8b19804e

Observation e87e9080-6d21-4d13-8e92-c9dc39889c6e · outbound

This paper cites Guidelines: • The answer NA means that the paper does not use existing assets.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not use existing assets

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.484090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.343151Z digest=sha256:0f520d5abfbe93bef614baa5303002ae3688dc483bb3fd4bb259c40a94c9be31

Observation 0f1ef82c-5084-4726-b5f5-3e7278d9683a · outbound

This paper cites Guidelines: • The answer NA means that the paper does not release new assets.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not release new assets

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.347041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.347041Z digest=sha256:a3c4af871a2194a945c676718d76e85b12f4b07eb95654f81405cd4545a256c4

Observation 32ea79b1-a7bb-4d37-9d42-eee4bc252c05 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.464389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.351339Z digest=sha256:9c49c13da05da5d18bbb218a40d8121d660ac39c5724bfa7add1fd8ed79d832a

Observation 1f04bade-1819-48ac-ab45-49e80ce8e8d8 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.450458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.355326Z digest=sha256:6fd079a6210125ba123166522c109396c5ce2a5f8842e01583569dd700973c02

Observation 17e2c672-1e70-46a4-a07b-68f28e108efe · outbound

This paper cites an unresolved cited work.

Extending Video Masked Autoencoders to 128 frames Unresolved cited work

Reference 2023

Resolution
parse uncertain
raw_fallback, observed 2026-08-12T16:22:26.154889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.119455Z digest=sha256:24f141f431fb2e822aceb794f0a1b39e2a4d04537cad901b5a2d122c5f51f5b5

Pith citing papers

No inbound Pith citation observations are available.