Pith. sign in

Paper Citation Record · LEDGER

Extending Video Masked Autoencoders to 128 frames

As of 12 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 0 inbound Pith citation observations for arXiv:2411.13683.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13683 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:22:25.355326Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

82 of 82 outbound references displayed

  • verified exact0
  • verified fuzzy71
  • unresolved10
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 84cf1d58-7cf6-4e59-b24a-18365b0499be · outbound

This paper cites Towards long-form video understanding.

Extending Video Masked Autoencoders to 128 frames Towards long-form video understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.014117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.014117Z digest=sha256:b25e716ff5caf137dea05e90baf49edf2ccbf59282278db832055bef7c11f247

Observation 57d954e0-e57e-4b91-834e-9520d8a00794 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

Extending Video Masked Autoencoders to 128 frames Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.018939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.018939Z digest=sha256:3150ff09bcaed6bd5c225f0c3efa5a0f2b4c5e3d4812a6d6e1b56ea170511e94

Observation 1e2090b2-dcc2-4824-91ac-5108d39886e7 · outbound

This paper cites Long movie clip classification with state-space video models.

Extending Video Masked Autoencoders to 128 frames Long movie clip classification with state-space video models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.442059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.023182Z digest=sha256:cf96cf09855cfa70594a7fe54dcd546429ab2943bc11b3f77e9baed45000c27c

Observation 3abf5ab4-b5ec-4dfe-a7e7-e0e7e9f6434f · outbound

This paper cites Selective structured state-spaces for long-form video understanding.

Extending Video Masked Autoencoders to 128 frames Selective structured state-spaces for long-form video understanding

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.428871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.027369Z digest=sha256:758efc038769ca652b805862e80d9ba1c921c48a5eed4e2ad24891809b9519b4

Observation e1d2eb4b-a9ef-4679-8eb6-ccfb70df90db · outbound

This paper cites Memory consolidation enables long-context video understanding.

Extending Video Masked Autoencoders to 128 frames Memory consolidation enables long-context video understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.415408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.031425Z digest=sha256:c244431981b179b4b154389087714bd3eeec3dbc73368199d2c178dea8d4ce0f

Observation fad0160e-a622-46f0-8cbe-1ba5c085f687 · outbound

This paper cites Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition.

Extending Video Masked Autoencoders to 128 frames Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.401384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.036030Z digest=sha256:ce75450a95be14ea37fc9ba2ec1224eebb41cb2f03d2c67090e09a7ab7c6573d

Observation 9136e534-9098-4a59-9d9b-2b43c8578db5 · outbound

This paper cites Token turing machines.

Extending Video Masked Autoencoders to 128 frames Token turing machines

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.386697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.040506Z digest=sha256:88d10a17ae2433e8c2e1c8f6bc5426cdcea1ce92c50fe5a5bda3a570e17018f2

Observation 1ee8d817-2f90-4da5-b5ff-b16c8255302b · outbound

This paper cites Video recap: Recursive captioning of hour-long videos.

Extending Video Masked Autoencoders to 128 frames Video recap: Recursive captioning of hour-long videos

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.374304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.045808Z digest=sha256:bca232f7eed703024a1ab66a0ebc3d6ca9a5c30d69d2cc6c1baeed80c036b3ae

Observation 15282ec6-ac03-48a8-bb04-1f898e7a3bc4 · outbound

This paper cites A simple recipe for contrastively pre-training video-first encoders beyond 16 frames.

Extending Video Masked Autoencoders to 128 frames A simple recipe for contrastively pre-training video-first encoders beyond 16 frames

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.360794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.050401Z digest=sha256:eac004dcaf48fc5c880edce5fab79224151fde8f922fdf022c22bf6b21a91541

Observation 96c04c42-cb02-4ce3-a8e5-aba75807fbeb · outbound

This paper cites A simple llm framework for long-range video question-answering.

Extending Video Masked Autoencoders to 128 frames A simple llm framework for long-range video question-answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.054724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.054724Z digest=sha256:a177352121a727fef042cdc3bf5190e98b973dd28328ff2854d3ab4ad2a1752c

Observation b28fee82-238e-451f-a888-05e92652f15a · outbound

This paper cites Long-form video- language pre-training with multimodal temporal contrastive learning.

Extending Video Masked Autoencoders to 128 frames Long-form video- language pre-training with multimodal temporal contrastive learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.341094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.058702Z digest=sha256:3bf45f09b8ba125b17df00b25e29a63e6b2a1a4c12c6d59a62bccfebf9f8dc25

Observation 4dc4e05a-4e81-45f9-ba36-ef5bdde63ccb · outbound

This paper cites Koala: Key frame-conditioned long video-llm.

Extending Video Masked Autoencoders to 128 frames Koala: Key frame-conditioned long video-llm

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.328284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.062400Z digest=sha256:c6809e2ac3306fc6b5a197aad613eeb3ff2ea932451c89066a48f6b3045cc4c8

Observation 9104b8a1-ea81-4e5d-b003-51f7c116eb30 · outbound

This paper cites Masked autoencoders as spatiotemporal learners.

Extending Video Masked Autoencoders to 128 frames Masked autoencoders as spatiotemporal learners

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.316949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.066391Z digest=sha256:66a00838d2b2b8e1288d999097be55d793e01dc472f4cb8a562438e7137a81d4

Observation 30a10d9c-6a53-456b-9d89-fc6543a31d37 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

Extending Video Masked Autoencoders to 128 frames Videomae v2: Scaling video masked autoencoders with dual masking

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.304662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.071077Z digest=sha256:9845efbc47b057a7009e374399e0d06dfd6303e4d3cd8f1e45cf15113191e0ea

Observation 61aa62bd-3850-4120-ade0-88283e9a3003 · outbound

This paper cites VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Extending Video Masked Autoencoders to 128 frames VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.290950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.074927Z digest=sha256:76cd6abe2b17609f08d3bbd9717dc44a2c07409e0c70ebba4ab3da33c76e8192

Observation cdd82006-37e8-4192-a1fe-5725f23d39cc · outbound

This paper cites How can objects help action recognition? In CVPR, 2023.

Extending Video Masked Autoencoders to 128 frames How can objects help action recognition? In CVPR, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.278784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.078492Z digest=sha256:13af8e4d776be21502806eb7938b89ed97f51535660a5975bfb5e3fb893228cf

Observation 2d832d9b-485e-4aca-8091-c506b8885877 · outbound

This paper cites Everest: Efficient masked video autoencoder by removing redundant spatiotemporal tokens.

Extending Video Masked Autoencoders to 128 frames Everest: Efficient masked video autoencoder by removing redundant spatiotemporal tokens

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.266111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.082930Z digest=sha256:eb2bc63dd586710f39548a246b903d9f905a6d69fb3021f78d0dd623f477fc3f

Observation 9d1350a1-3813-4f80-bbf4-6a55f4baba20 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100.

Extending Video Masked Autoencoders to 128 frames Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.253869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.086594Z digest=sha256:b4a89636cabe36ff9ce2d8fb92e010fb380cb3129d4f8b5d584aa00339555828

Observation 6c0d1bea-e94f-46af-ab04-b4c051bd18c1 · outbound

This paper cites Resound: Towards action recognition without representation bias.

Extending Video Masked Autoencoders to 128 frames Resound: Towards action recognition without representation bias

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.241287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.090422Z digest=sha256:f6349f8b6b8e00ecacc19515945830abf54cd31eb9782817ebf05075096d9be7

Observation 21f1223f-011f-4832-ade6-2ee718f0afb4 · outbound

This paper cites an unresolved cited work.

Extending Video Masked Autoencoders to 128 frames Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:22:26.229238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.094637Z digest=sha256:48e608d3f3269b8a38f5ec2ab2e0b85c205a1cd0fdb69e375438d615ce471d86

Observation 05497a92-60ff-443e-9f04-71f253a54044 · outbound

This paper cites Bevt: Bert pretraining of video transformers.

Extending Video Masked Autoencoders to 128 frames Bevt: Bert pretraining of video transformers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.218024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.098703Z digest=sha256:755e64fcddcfc2eff11c934da26223b15c3aa5bbcbd690052d5d1c98ce967113

Observation e5e7babb-2187-497a-af9d-76469a1373b0 · outbound

This paper cites Girdhar, A.

Extending Video Masked Autoencoders to 128 frames Girdhar, A

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.205639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.103922Z digest=sha256:0e48d168c76793716b69eab622bed20be648923cf4db0b7bcbd3fb661c822c34

Observation 713e89ef-4263-417a-b095-74a19c92e633 · outbound

This paper cites Zero-shot text-to-image generation.

Extending Video Masked Autoencoders to 128 frames Zero-shot text-to-image generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.193574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.107960Z digest=sha256:f8954d17028d7f39bd9f6c536eba93ad87257d881c6c9558949157f7ac2af302

Observation bc7e4fd2-2979-4311-a58b-b08529e687ad · outbound

This paper cites Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning.

Extending Video Masked Autoencoders to 128 frames Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.182211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.111942Z digest=sha256:50108b7e365bf95b2a53e61a75d03fd8d8f820d36251e9428fa845c2bcf626b3

Observation e5511110-75f1-465e-bd26-2f2d846c7286 · outbound

This paper cites Magvit: Masked generative video transformer.

Extending Video Masked Autoencoders to 128 frames Magvit: Masked generative video transformer

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.170158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.115504Z digest=sha256:eb776290259a5fca187683e2c33d3d14bfa82f15f4785a40d815a152a6d626b0

Observation 929db5b8-04b5-4e3d-a6cc-91d48a12431c · outbound

This paper cites Tokenlearner: What can 8 learned tokens do for images and videos? In NeurIPS, 2021.

Extending Video Masked Autoencoders to 128 frames Tokenlearner: What can 8 learned tokens do for images and videos? In NeurIPS, 2021

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.143450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.123089Z digest=sha256:d7777196808935330bd83f432758608e6b6291850436cde438119ce8a679b931

Observation 9cd15a53-f043-4a68-9079-f5e6871c87d9 · outbound

This paper cites something something.

Extending Video Masked Autoencoders to 128 frames something something

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.130927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.126787Z digest=sha256:214127ea5cc4678a8fdb6e0af7586bbb2687b97fccb06d78fd2a4b48e2516d73

Observation 319a40db-cab6-42ab-86cd-b86c83095958 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Extending Video Masked Autoencoders to 128 frames An image is worth 16x16 words: Transformers for image recognition at scale

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.118986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.130449Z digest=sha256:162695a4f0b4544e4f99db09e08532707270c5892b7b7e01fb065d1addee0188

Observation d87ef826-a40a-4764-958d-63623af7b0cc · outbound

This paper cites Mgmae: Motion guided masking for video masked autoencoding.

Extending Video Masked Autoencoders to 128 frames Mgmae: Motion guided masking for video masked autoencoding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.105953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.134329Z digest=sha256:de71589bdac96671210211cb9355317549af86287d2a89bce58515d57bae520b

Observation b5b0e6f3-fc2d-408f-a232-3ca3c19afbf2 · outbound

This paper cites Motion-guided masking for spatiotemporal representation learning.

Extending Video Masked Autoencoders to 128 frames Motion-guided masking for spatiotemporal representation learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.091040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.138258Z digest=sha256:6aad40f014fb300644e486dc75a59c5edce253a690bba61fa9ef3b0fc0c0b213

Observation e32096bb-4f64-4512-ac99-52032f45532d · outbound

This paper cites Video codec design: developing image and video compression systems.

Extending Video Masked Autoencoders to 128 frames Video codec design: developing image and video compression systems

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.077412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.142170Z digest=sha256:514c3e5115324bccc97e1a1951d50e8ca0f42b480ae5ce8663737b285af92138

Observation 9a8ae6a4-ee82-4266-bced-f94d9ea68a02 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

Extending Video Masked Autoencoders to 128 frames Raft: Recurrent all-pairs field transforms for optical flow

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.062271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.146065Z digest=sha256:3c72fa9a05c44676de5eb7a54fde5a3033b4920010f283fdccad79573ac74405

Observation d71d0fc2-7964-48d3-9875-5cb23baf2090 · outbound

This paper cites Videoprism: A foundational visual encoder for video understanding.

Extending Video Masked Autoencoders to 128 frames Videoprism: A foundational visual encoder for video understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.050812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.149577Z digest=sha256:8cc7badebd0dc90403da91a72fa3dfce4312e35bb4a883d0d64e9b4e1d089ebc

Observation 0c81c503-ca16-426b-8f61-e24d49b7afd8 · outbound

This paper cites Internvideo2: Scaling video foundation models for multimodal video understanding.

Extending Video Masked Autoencoders to 128 frames Internvideo2: Scaling video foundation models for multimodal video understanding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.037738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.153989Z digest=sha256:4d4b9b0a2d701aaf447170139184fae5e84944f22eea49a0909a92a0e857a603

Observation 55a81fcf-1c28-4640-b3d8-ce4c2b551f17 · outbound

This paper cites Vivit: A video vision transformer.

Extending Video Masked Autoencoders to 128 frames Vivit: A video vision transformer

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.024691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.157567Z digest=sha256:7d78982e6e19073a30c719bade13beb1e14a581648b5090c7df7788f9df05513

Observation 07e3e8ce-9ff2-4a61-9d79-e761d97b7bfb · outbound

This paper cites Finite scalar quantization: VQ-V AE made simple.

Extending Video Masked Autoencoders to 128 frames Finite scalar quantization: VQ-V AE made simple

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.011181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.161177Z digest=sha256:818445cfd6ca37791e961d563008458879b270f837fe47556b5668572e2d8fd0

Observation 02398460-7a6d-40c7-bf63-b314a15cce3f · outbound

This paper cites A Short Note about Kinetics-600.

Extending Video Masked Autoencoders to 128 frames A Short Note about Kinetics-600

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.164843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.164843Z digest=sha256:410dbc7b5551123366ff4c4bc024e903248ead5c1323f5872ebebf1ac6a05eeb

Observation 95ecca2f-f416-4873-9e90-80143ff10102 · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

Extending Video Masked Autoencoders to 128 frames A Short Note on the Kinetics-700 Human Action Dataset

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.168857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.168857Z digest=sha256:1429657810ac968da0393cca3506c1b7938e2dc8a44d2ce5d1b0aa4e3bf4f434

Observation a196cfee-b1a4-49b0-b4a2-0d09a82b791d · outbound

This paper cites Multiview transformers for video recognition.

Extending Video Masked Autoencoders to 128 frames Multiview transformers for video recognition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.997209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.172795Z digest=sha256:f093e50a35757a1c58d22bb57b4a03d80b1aa11062248da44843f84731ce7ebc

Observation 33cfc5de-95ae-40d3-8b38-2227648e89d7 · outbound

This paper cites Temporally-Adaptive Models for Efficient Video Understanding.

Extending Video Masked Autoencoders to 128 frames Temporally-Adaptive Models for Efficient Video Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.176362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.176362Z digest=sha256:0074e298f69478f2e36db737438862985cf36695706c53c84b3fa3abc94ddf73

Observation 2f028e33-ae02-4d0b-8d47-ea889ce8eda2 · outbound

This paper cites Training a Large Video Model on a Single Machine in a Day.

Extending Video Masked Autoencoders to 128 frames Training a Large Video Model on a Single Machine in a Day

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.180917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.180917Z digest=sha256:73fe7619f75951c2ccf980dc6ea9d98030e8c546f02147f489639e94d7e77c33

Observation 72d5d421-0073-40e7-8db3-f1349385d633 · outbound

This paper cites Imagenet-21k pretraining for the masses.

Extending Video Masked Autoencoders to 128 frames Imagenet-21k pretraining for the masses

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.983302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.185019Z digest=sha256:b9c91aff9b2cbebb845fbd3c5e2407d06af7acd4f09af4fb73c661773f04f1c0

Observation 45a79a2a-274c-496e-b7cf-87f33caffe1d · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Extending Video Masked Autoencoders to 128 frames Ego4d: Around the world in 3,000 hours of egocentric video

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.972248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.188571Z digest=sha256:a4c7b07242d1f1fec6c2e5cd51cc3ecf0231da1a6ac84c4e990439d00e7530bc

Observation bdb7e0cb-169d-4786-bfcf-e6475c709066 · outbound

This paper cites Verbs in action: Improving verb understanding in video-language models.

Extending Video Masked Autoencoders to 128 frames Verbs in action: Improving verb understanding in video-language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.961247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.192158Z digest=sha256:a49625776d003de48d79b8bf96a3c6ec0df87e0bc46d20fb8561b0f92541b710

Observation ec83fec6-7af5-4875-aefd-eb8f725487be · outbound

This paper cites Slowfast networks for video recognition.

Extending Video Masked Autoencoders to 128 frames Slowfast networks for video recognition

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.949510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.196373Z digest=sha256:0051d91802d6151b1f8fd7374ca075e4391be716c830359a1250ec3c8d58edf3

Observation 24b0f31e-184e-43ca-9067-4636c6be323f · outbound

This paper cites Interactive prototype learning for egocentric action recognition.

Extending Video Masked Autoencoders to 128 frames Interactive prototype learning for egocentric action recognition

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.939197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.200298Z digest=sha256:2a732ac927416a109d32423ce4d8cd44939eaff0005346ee14f4f27a8818c7c0

Observation 0be7e6c2-5f50-4e9f-aa24-cd85927732bb · outbound

This paper cites Movinets: Mobile video networks for efficient video recognition.

Extending Video Masked Autoencoders to 128 frames Movinets: Mobile video networks for efficient video recognition

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.928022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.203795Z digest=sha256:d68fec973b779179c375c988b43a0ede66917c2f98d924090277cf906ff25cee

Observation a08bd723-c388-4787-9478-9e4cdde59e2d · outbound

This paper cites Omnivore: A single model for many visual modalities.

Extending Video Masked Autoencoders to 128 frames Omnivore: A single model for many visual modalities

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.916079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.207333Z digest=sha256:b9c9b7f685c7cc92eb95fc323a2b4b19aaab11de52297a7161bdb3df6f315f2f

Observation 08528d82-d877-41a5-ba58-e4c47280dda9 · outbound

This paper cites Learning video representations from large language models.

Extending Video Masked Autoencoders to 128 frames Learning video representations from large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.903493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.211110Z digest=sha256:9d3a6b0f0148d757f38a733e33817376a44559f963059945c82d6e51eecfb85e

Observation 994e1ecd-afde-4267-9eff-b2a21a9e9a74 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, 2021.

Extending Video Masked Autoencoders to 128 frames Is space-time attention all you need for video understanding? In ICML, 2021

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.891843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.214507Z digest=sha256:3666a6b301de3fa7a020e12e2353821787ea1bdf9fe8d6afb81beaf3dba4f771

Observation 62167356-5e55-4fb7-8620-e87d6e8f161f · outbound

This paper cites Video swin transformer.

Extending Video Masked Autoencoders to 128 frames Video swin transformer

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.879718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.218269Z digest=sha256:b641f4ce21c2d434e2e76bfb10b6a6364f4ffa80e073beefd7d2ff34024e687c

Observation be25e975-2432-4536-8ce7-8d7427f0d842 · outbound

This paper cites Can an image classifier suffice for action recognition? In ICLR, 2022.

Extending Video Masked Autoencoders to 128 frames Can an image classifier suffice for action recognition? In ICLR, 2022

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.868902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.221993Z digest=sha256:5e340d4a49457eec55e85a00c873d70cabda08cce0c9b2b03a9cf56ee1a28336

Observation 5efa996f-7043-42bb-bb20-13945d3dc966 · outbound

This paper cites Object-region video transformers.

Extending Video Masked Autoencoders to 128 frames Object-region video transformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.857839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.225841Z digest=sha256:90b88591c207e61b6dc60efc623cfb3bd03d68b3abc91cbc7dead1311e5fd91f

Observation 738d1098-9f5d-4766-8627-53e4d24a5d64 · outbound

This paper cites Aim: Adapting image models for efficient video action recognition.

Extending Video Masked Autoencoders to 128 frames Aim: Adapting image models for efficient video action recognition

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.845804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.229883Z digest=sha256:7136397fb460c40870dacfa3a99bfff83968e5f5db999ed0c65c500a58313d8c

Observation 26189b8b-33f3-4750-9990-c658006ca20a · outbound

This paper cites Video-focalnets: Spatio-temporal focal modulation for video action recognition.

Extending Video Masked Autoencoders to 128 frames Video-focalnets: Spatio-temporal focal modulation for video action recognition

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.834222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.233575Z digest=sha256:7de61d30658e62fb5976ffc858620acc36463bb864a822c6f1ac1d5c73ecadcd

Observation 6d0e83f9-ab6a-4b0e-910c-dc8b000828e3 · outbound

This paper cites Language model beats diffusion–tokenizer is key to visual generation.

Extending Video Masked Autoencoders to 128 frames Language model beats diffusion–tokenizer is key to visual generation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.819649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.237643Z digest=sha256:7d55161f18ec48787f58a146930502923aa8c10f56cf94c323f567f04a86a516

Observation 35bf2940-3183-4085-8371-02c89b5ee713 · outbound

This paper cites Finegym: A hierarchical video dataset for fine-grained action understanding.

Extending Video Masked Autoencoders to 128 frames Finegym: A hierarchical video dataset for fine-grained action understanding

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.804450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.241462Z digest=sha256:6cf4d3d42d8be34433dda4413a1df075ac1566a37f1756269101374fc9c19a80

Observation c0e06bc4-9991-4702-8521-b65e8aa77aa3 · outbound

This paper cites Learning temporal cues for fine-grained action recognition.

Extending Video Masked Autoencoders to 128 frames Learning temporal cues for fine-grained action recognition

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.790238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.245443Z digest=sha256:7947a289767834a7157e1e281b0807b23b990aab28b8c3a3ca5acce6bf22ff4b

Observation b2aa5f83-862d-4df6-b1f7-ebe7564d0c09 · outbound

This paper cites Tsm: Temporal shift module for efficient video understanding.

Extending Video Masked Autoencoders to 128 frames Tsm: Temporal shift module for efficient video understanding

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.775928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.249266Z digest=sha256:a90a0912eb9156dd8fd5a96c92d1a5ca7bfc8faca6e103e247400f2777b2581f

Observation 33ddac3d-e7d2-4c61-85b0-c7df2c2651a0 · outbound

This paper cites Temporal query networks for fine-grained video understanding.

Extending Video Masked Autoencoders to 128 frames Temporal query networks for fine-grained video understanding

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.762566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.253411Z digest=sha256:4d8af5d65bbea59a632c7292acecf82d440c0a04a44b70404067caf7e274e200

Observation d5c37d89-8b6f-4e04-b893-76292cb9fbaf · outbound

This paper cites Combined cnn transformer encoder for enhanced fine-grained human action recognition.

Extending Video Masked Autoencoders to 128 frames Combined cnn transformer encoder for enhanced fine-grained human action recognition

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.751023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.257990Z digest=sha256:dc2da9b99e344c3e2a233a9a2c0ee4fb4a23d7f0e7509853241f701e6c090175

Observation c1217cf0-e260-4354-a23a-649cf1548550 · outbound

This paper cites Going deeper with image transformers.

Extending Video Masked Autoencoders to 128 frames Going deeper with image transformers

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.738041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.262106Z digest=sha256:49f9f6746a8ca55e5ee91650aa30988bd589b9022796d6dc2cb1e197e5d1ed46

Observation e4c247d4-831e-46c6-a25f-64cfd87e3220 · outbound

This paper cites Scenic: A jax library for computer vision research and beyond.

Extending Video Masked Autoencoders to 128 frames Scenic: A jax library for computer vision research and beyond

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.722729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.266075Z digest=sha256:7790235845b58331f0b04e0ca594ae49f145befd6ca7a2fdaf80c38abcab7263

Observation 0a74a736-5133-4c9c-a196-f891b24fb370 · outbound

This paper cites We found this slightly improved accuracy ( +0.5 points on EPIC-Kitchens-100 Verbs and +1.2 points on Diving48).

Extending Video Masked Autoencoders to 128 frames We found this slightly improved accuracy ( +0.5 points on EPIC-Kitchens-100 Verbs and +1.2 points on Diving48)

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.700975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.270584Z digest=sha256:dddb03d54041ca9887699060a0dca0037a4683ca5dad1ea719d02ccf53dfd777

Observation e8883ef5-e7c8-4e61-a54c-8143413ce151 · outbound

This paper cites an unresolved cited work.

Extending Video Masked Autoencoders to 128 frames Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:22:25.681314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.274933Z digest=sha256:d8b2b7dfbea613b1da50b51c16260527bcad3180535fc904a18104eabba2b9fb

Observation 7121ac07-0f58-47f2-ad01-88e579276083 · outbound

This paper cites 17 Table 9: Model size vs frames.

Extending Video Masked Autoencoders to 128 frames 17 Table 9: Model size vs frames

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.667471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.279010Z digest=sha256:7e2c8d099f917c3e69d70c2472800bfe0fd74809821baad969502d486c1c5f01

Observation 13270ed7-aa41-42bd-a1da-25f5ef192c76 · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.653781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.283416Z digest=sha256:7fbc2b324d397dcd765afa3b49c4cce811f5f243db081dc8e51b7f5e6c760759

Observation fd7f5d79-0999-4578-ba91-897e965c755c · outbound

This paper cites Limitations.

Extending Video Masked Autoencoders to 128 frames Limitations

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.636688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.287958Z digest=sha256:d2947482766ae9448975d39ec78bc27c54be3d8b22f7f7d15113ab74ff1b8e26

Observation 10456fc6-93cc-44fe-982e-3888d20813f1 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.619738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.292611Z digest=sha256:ec27003575ed77cf89d631d7744e689fc3f63f2e40cbb448cbfb60b64a40a05e

Observation 955057ce-da5a-4eab-b8a1-3bb1e743a1f6 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include experiments

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.605281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.296935Z digest=sha256:e681dd2ab27f20c8161d1283935b249c262d309f717ee0b504ef0b76ab0f3b56

Observation 05b446cf-6eb3-4330-affd-25d7dd01dfe0 · outbound

This paper cites Guidelines: • The answer NA means that paper does not include experiments requiring code.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that paper does not include experiments requiring code

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.588626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.302631Z digest=sha256:35b431de5923ac0d259583acd08f19b2885a4a19032fea7fdb29ad7b7dd87248

Observation 903717b9-95c5-4a0b-ab71-e54f72805d1b · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include experiments

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.572605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.307592Z digest=sha256:5ce3c0aaf2f72073715a8701023b83e5807c2d9b4f15d22d32d4df17c906c814

Observation 7504251a-762c-41d2-ad08-4e797b83ce0a · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include experiments

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.559349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.312618Z digest=sha256:28f9d277eddaa2196249d26d9118c664cbde047215165e7d9e3b621edbdb290b

Observation e409290e-ad8e-4ba9-95c5-8ff826b82976 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include experiments

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.543694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.318955Z digest=sha256:9ff58b7cf644e24489c80660c161dac1460f3fc6fb7a3bf18ecd8a5eefe2a027

Observation 151092d4-fe23-4c42-8e99-94dce5af1702 · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.528011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.323122Z digest=sha256:4089eb64c59343e311bfcf3e1bc3546746b440e549c4503f8c5d6c6f7d6cb187

Observation 27550e91-dc95-4ad8-834c-1d3ca347c615 · outbound

This paper cites • If the authors answer NA or No, they should explain why their work has no societal impact or why the paper does not address societal impact.

Extending Video Masked Autoencoders to 128 frames • If the authors answer NA or No, they should explain why their work has no societal impact or why the paper does not address societal impact

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.510829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.327717Z digest=sha256:3c2d4b40f826ba5e1679f0c150ca860c57b7a21a7460ca69dddf0f50d76d7145

Observation 1832e7fb-510a-4c78-80da-e7490525eabf · outbound

This paper cites Guidelines: • The answer NA means that the paper poses no such risks.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper poses no such risks

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.497541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.332877Z digest=sha256:3e9652f6da41f1bc0cd6bfa6294de40712e7012a4cc9d2b268c6048b99ef888b

Observation e87e9080-6d21-4d13-8e92-c9dc39889c6e · outbound

This paper cites Guidelines: • The answer NA means that the paper does not use existing assets.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not use existing assets

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.484090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.343151Z digest=sha256:1bceb1414468924c35d94fa7df5c782f83a1d84ec4bfd2a1336907f06a372859

Observation 0f1ef82c-5084-4726-b5f5-3e7278d9683a · outbound

This paper cites Guidelines: • The answer NA means that the paper does not release new assets.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not release new assets

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.347041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.347041Z digest=sha256:6ae0499cc7ecf8a8bf4479ee1b2d750976c4bdfea28b18fc72b67bcfd63b4d78

Observation 32ea79b1-a7bb-4d37-9d42-eee4bc252c05 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.464389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.351339Z digest=sha256:7a961806a0088b99b5bd347de1e2423a39f75df16e6f58e2374fc846177e7d77

Observation 1f04bade-1819-48ac-ab45-49e80ce8e8d8 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.450458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.355326Z digest=sha256:b8999f4c125b0570cb6150c70f970ba2eca1f0ef2d7a075ae4637f2ebfdb534b

Observation 17e2c672-1e70-46a4-a07b-68f28e108efe · outbound

This paper cites an unresolved cited work.

Extending Video Masked Autoencoders to 128 frames Unresolved cited work

Reference 2023

Resolution
parse uncertain
raw_fallback, observed 2026-08-12T16:22:26.154889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:22:25.119455Z digest=sha256:685a0b044b016cfed8759c53e389f2b702cc868dfa8c5412e3d4ba758ac3b8bc

Pith citing papers

No inbound Pith citation observations are available.