Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:59:00.269362Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2411.10503.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:59:00.269362Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:49:42.989079Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T11:49:52.967735Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e38793b4-4c0b-413c-a19f-c16a564bf07f · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction From methods to datasets: A survey on image-caption generators
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 07e62f32-e74b-472f-8214-0da90b355604 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ede4becb-eb9b-4d1d-9244-01c0b8270f6d · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Audiomnist: Exploring explainable artificial intelli- gence for audio analysis on a simple benchmark
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bb47d286-b6a0-48af-854c-3876671d61c9 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Is space-time attention all you need for video understanding? In ICML, page 4, 2021
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1e4e1f29-19c9-4d6c-869a-00d1968a8a6d · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Pcanet: A simple deep learning baseline for image classification? IEEE transactions on image pro- cessing, 24(12):5017–5032, 2015
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ebd41328-7dee-49d8-81d4-4ef700d44ea0 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Generative pre- training from pixels
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be956813-70c8-4907-803b-68d82d645b5f · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction UNITER: UNiversal Image-TExt Representation Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f116849b-3e73-444b-adb7-5411c45c0cea · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction High-performance long- term tracking with meta-updater
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dea018fd-f403-4c3f-8eac-732c7f57a13f · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Tinyvirat: Low-resolution video action recognition
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6e13c638-90f6-46ce-b6cd-b42eb5df37a8 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction An image is worth 16x16 words: Transformers for image recognition at scale
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bded84f4-ef3f-4f61-ba8c-e8b27111a274 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Lasot: A high-quality benchmark for large-scale single ob- ject tracking
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4529372c-b4d8-4775-bb83-368e7da1498d · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Improving Language Understanding from Screenshots
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1edaf6a9-600e-4a08-b89f-46ac57e44bd0 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Photorealistic video generation with diffusion models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 32a06e5c-ea02-4501-8931-9ef37cf18346 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Unit: Multimodal mul- titask learning with a unified transformer
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b3568299-f139-4847-8c3b-24d4ba37b3c7 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Video Question Answering with Spatio-Temporal Reasoning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4b4732d1-b4da-4fe8-b4c3-f5baa2a6ff1d · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Unifying Question Answering, Text Classification, and Regression via Span Extraction
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6886f5d-86be-438c-afaa-8ee0a23febef · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Learning multiple layers of features from tiny images
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7940d5d-b5c9-44f4-b8fe-e9759f30d852 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Temporally consistent video colorization with deep feature propagation and self-regularization learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4360691c-c323-4759-ae1d-7fb0ae0608ab · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Decoupled weight de- cay regularization
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c5327fb-a60e-4265-b2d4-6f95c06ba0af · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction The multi-modal fusion in vi- sual question answering: a review of attention mechanisms
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 36d92b00-531d-4b93-ba7d-06937ba713d6 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction The Natural Language Decathlon: Multitask Learning as Question Answering
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2b3e8e3-6335-4962-8e27-23dcfc6297b2 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Transframer: Arbitrary Frame Prediction with Generative Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8febd18-221e-44e8-870b-771d3b1c81a4 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Gpt-4 technical report, 2024
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28a728f5-9c93-45f8-a5c9-c7221187eec2 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Video (language) modeling: a baseline for generative models of natural videos
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8f9b94e-8d86-4e8c-a9f5-04a2d48dbaeb · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Language Modelling with Pixels
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84a023d8-5388-4b65-8205-75d150f0180a · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Implicit Stacked Autoregressive Model for Video Prediction
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3645c40e-e2f8-41a7-8b0e-74fd0f97f629 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction FLAVA: A Foundational Language And Vision Alignment Model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b5591c1-6ae5-4507-8aba-f7e3079ee87c · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Recursive deep models for semantic compositional- ity over a sentiment treebank
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3bc84ed7-9dfe-438c-a501-a72f9754e8c3 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction PIXAR: Auto-Regressive Language Modeling in Pixel Space
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f7ab565-4240-441f-b0eb-b5330b29b97d · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Transformation-Based Models of Video Sequences
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e09d4d6-8b24-41be-a5a3-25a80c23e98a · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Multimodal llm enhanced cross- lingual cross-modal retrieval
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b589a8ee-9745-4773-88b5-6140da2de8c1 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Bovik, H.R
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42f29cc1-00e4-4b44-a243-a669c6a0c5eb · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Visual question answer- ing: A survey of methods and datasets
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59f3cd8b-2217-46d2-afab-376787778509 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Pixel Sentence Representation Learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2a3e8b8-2ffb-492f-a9ff-e7108e387f74 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Videogpt: Video generation using vq-vae and trans- formers, 2021
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b11476c7-6673-4d84-afec-b5c57ff54520 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Colormnet: A memory-based deep spatial-temporal feature propagation network for video colorization
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8b438b27-b62c-4236-8da5-c7e2ddcf297b · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction CLEVRER: CoLlision Events for Video REpresentation and Reasoning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcc370f7-493e-411a-bdae-cb5f5a8c8bcf · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Akin Yilmaz and Ahmet Murat Tekalp
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6b0c4265-dd2e-4798-bac9-51cb76b4d9b6 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Colorful image colorization
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3902d12c-2589-4e0d-8d3d-082882a42ac4 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction Bag of Tricks for Effective Language Model Pretraining and Downstream Adaptation: A Case Study on GLUE
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2856d329-c2aa-4a9e-be1c-342be7dafe80 · outbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction A survey on vqa: Datasets and approaches
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3c62bd83-a8a4-42bc-a3c9-82446053c66c · inbound
Unraveling Spatio-Temporal Foundation Models via the Pipeline Lens: A Comprehensive Review Everything is a Video: Unifying Modalities through Next-Frame Prediction
Reference 135
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.