Pith. sign in

Paper Citation Record · LEDGER

Everything is a Video: Unifying Modalities through Next-Frame Prediction

As of 13 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2411.10503.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10503 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:59:00.269362Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:49:42.989079Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:49:52.967735Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e38793b4-4c0b-413c-a19f-c16a564bf07f · outbound

This paper cites From methods to datasets: A survey on image-caption generators.

Everything is a Video: Unifying Modalities through Next-Frame Prediction From methods to datasets: A survey on image-caption generators

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.957912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:58:59.985267Z digest=sha256:feeb698fbda37602d31a95086f700f26ac3b928885c8ea9d70ecedd84f8c8010

Observation 07e62f32-e74b-472f-8214-0da90b355604 · outbound

This paper cites data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language.

Everything is a Video: Unifying Modalities through Next-Frame Prediction data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:58:59.991539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:58:59.991539Z digest=sha256:539bf79882d6a0935172357931405f6058612bce86ad18761cf75fb84dc8ad4f

Observation ede4becb-eb9b-4d1d-9244-01c0b8270f6d · outbound

This paper cites Audiomnist: Exploring explainable artificial intelli- gence for audio analysis on a simple benchmark.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Audiomnist: Exploring explainable artificial intelli- gence for audio analysis on a simple benchmark

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.938193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:58:59.997316Z digest=sha256:da5f69929538bd4c4613e820c58faf4dbdda4c861fea76e7662b8046e30f7c9b

Observation bb47d286-b6a0-48af-854c-3876671d61c9 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.919706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:59:00.009812Z digest=sha256:68d46f1eb67851b6ba18948b045cbb689d6d3e4bddf4c2df09e6b7d00b386cff

Observation 1e4e1f29-19c9-4d6c-869a-00d1968a8a6d · outbound

This paper cites Pcanet: A simple deep learning baseline for image classification? IEEE transactions on image pro- cessing, 24(12):5017–5032, 2015.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Pcanet: A simple deep learning baseline for image classification? IEEE transactions on image pro- cessing, 24(12):5017–5032, 2015

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.899756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:59:00.022299Z digest=sha256:5dc64f5551b6d128f80008544268ba315ad116caf15e6d2de69e1afbbffd0908

Observation ebd41328-7dee-49d8-81d4-4ef700d44ea0 · outbound

This paper cites Generative pre- training from pixels.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Generative pre- training from pixels

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.031644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.031644Z digest=sha256:f1f2db56fede57828c2eafae57b0e99137de431d3f6719d8418f5458a3299352

Observation be956813-70c8-4907-803b-68d82d645b5f · outbound

This paper cites UNITER: UNiversal Image-TExt Representation Learning.

Everything is a Video: Unifying Modalities through Next-Frame Prediction UNITER: UNiversal Image-TExt Representation Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.046717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.046717Z digest=sha256:eb74ef4bbe662ce071938969f15b809f1a0902a3cfae19f703bfae7c09a87809

Observation f116849b-3e73-444b-adb7-5411c45c0cea · outbound

This paper cites High-performance long- term tracking with meta-updater.

Everything is a Video: Unifying Modalities through Next-Frame Prediction High-performance long- term tracking with meta-updater

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.865798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:59:00.055661Z digest=sha256:8bbba1974b8fa9a3993092a80971df52796075f47fcf22454caf087180a92394

Observation dea018fd-f403-4c3f-8eac-732c7f57a13f · outbound

This paper cites Tinyvirat: Low-resolution video action recognition.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Tinyvirat: Low-resolution video action recognition

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.848075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:59:00.062318Z digest=sha256:2fbc294eddac4aa031d2048331fa764fdaed70a14f4b27fd4c4893f1985a2324

Observation 6e13c638-90f6-46ce-b6cd-b42eb5df37a8 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Everything is a Video: Unifying Modalities through Next-Frame Prediction An image is worth 16x16 words: Transformers for image recognition at scale

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.831679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:59:00.068835Z digest=sha256:f7f24f62c87f45d7208fad86dc8c12b852979fe9b04b8c257af031ec878b4282

Observation bded84f4-ef3f-4f61-ba8c-e8b27111a274 · outbound

This paper cites Lasot: A high-quality benchmark for large-scale single ob- ject tracking.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Lasot: A high-quality benchmark for large-scale single ob- ject tracking

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.077421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.077421Z digest=sha256:9757fe5e6a9be8aa48466b26bf517627e9f8de5b381edcd677713a6a4c5da387

Observation 4529372c-b4d8-4775-bb83-368e7da1498d · outbound

This paper cites Improving Language Understanding from Screenshots.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Improving Language Understanding from Screenshots

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.085021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.085021Z digest=sha256:3aa7665652d9e0188797ad9fb9f177177b58685fbc71043682b10deaf4c4b4db

Observation 1edaf6a9-600e-4a08-b89f-46ac57e44bd0 · outbound

This paper cites Photorealistic video generation with diffusion models.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Photorealistic video generation with diffusion models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.804522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:59:00.091353Z digest=sha256:a120d87f5a856918758bce69765b00ec5b60322b9ca1bd7c73e9f95a73f2ce71

Observation 32a06e5c-ea02-4501-8931-9ef37cf18346 · outbound

This paper cites Unit: Multimodal mul- titask learning with a unified transformer.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Unit: Multimodal mul- titask learning with a unified transformer

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.787385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:59:00.097416Z digest=sha256:3f5a60d2b5ae3138f761cf3cea226396ebadbd816d5e5528483111e9e6153b92

Observation b3568299-f139-4847-8c3b-24d4ba37b3c7 · outbound

This paper cites Video Question Answering with Spatio-Temporal Reasoning.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Video Question Answering with Spatio-Temporal Reasoning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.770191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:59:00.103573Z digest=sha256:fcef6264ce8ebf075aec06e9410b9ff0f539c30fc9c5a7a831abe950f675764e

Observation 4b4732d1-b4da-4fe8-b4c3-f5baa2a6ff1d · outbound

This paper cites Unifying Question Answering, Text Classification, and Regression via Span Extraction.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Unifying Question Answering, Text Classification, and Regression via Span Extraction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.112379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.112379Z digest=sha256:399f29d56ea529b2413e68eb47f648de8872bb6132941f91b33509b63ba73d8b

Observation c6886f5d-86be-438c-afaa-8ee0a23febef · outbound

This paper cites Learning multiple layers of features from tiny images.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Learning multiple layers of features from tiny images

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.119758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.119758Z digest=sha256:3b82dcf9148cd1c17e547bdadb6a00626deaefafc1e4c7e9050aa3cb6c4fe844

Observation d7940d5d-b5c9-44f4-b8fe-e9759f30d852 · outbound

This paper cites Temporally consistent video colorization with deep feature propagation and self-regularization learning.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Temporally consistent video colorization with deep feature propagation and self-regularization learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.739590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:59:00.126027Z digest=sha256:bf8a30b8312162f8ffe0dddfd2fdc0f7fbd2a8d4b73d1d97a022a6a0314b006b

Observation 4360691c-c323-4759-ae1d-7fb0ae0608ab · outbound

This paper cites Decoupled weight de- cay regularization.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Decoupled weight de- cay regularization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.134549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.134549Z digest=sha256:fe2a5c18d6d80975c2c0fd518de4f9e75eb9d72f5ed982a171d127a42a1f121c

Observation 3c5327fb-a60e-4265-b2d4-6f95c06ba0af · outbound

This paper cites The multi-modal fusion in vi- sual question answering: a review of attention mechanisms.

Everything is a Video: Unifying Modalities through Next-Frame Prediction The multi-modal fusion in vi- sual question answering: a review of attention mechanisms

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.710446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:59:00.140434Z digest=sha256:31ed20bd590aa54fe87228aa5ab4644bc88f7d2f2799fd28b3ffc16136852786

Observation 36d92b00-531d-4b93-ba7d-06937ba713d6 · outbound

This paper cites The Natural Language Decathlon: Multitask Learning as Question Answering.

Everything is a Video: Unifying Modalities through Next-Frame Prediction The Natural Language Decathlon: Multitask Learning as Question Answering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.145573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.145573Z digest=sha256:0443b4aee262f10d5a070fed727aeafd88fb5d2f7a13a4c8ecdf215a074a13ea

Observation d2b3e8e3-6335-4962-8e27-23dcfc6297b2 · outbound

This paper cites Transframer: Arbitrary Frame Prediction with Generative Models.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Transframer: Arbitrary Frame Prediction with Generative Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.151096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.151096Z digest=sha256:ef6d26288ecac4e735af4ed08ad0669329d518829fd79ca9915710c3e5554ce9

Observation e8febd18-221e-44e8-870b-771d3b1c81a4 · outbound

This paper cites Gpt-4 technical report, 2024.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Gpt-4 technical report, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.157591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.157591Z digest=sha256:4bf77119b9544ad0f4587b7ff8174edd148acc93598aa35211987fb36f37c56b

Observation 28a728f5-9c93-45f8-a5c9-c7221187eec2 · outbound

This paper cites Video (language) modeling: a baseline for generative models of natural videos.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Video (language) modeling: a baseline for generative models of natural videos

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.167425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.167425Z digest=sha256:d71a0d1001bfcab4f610e1e33c50e16553a31de6f1e9d26d63e26989779edff9

Observation b8f9b94e-8d86-4e8c-a9f5-04a2d48dbaeb · outbound

This paper cites Language Modelling with Pixels.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Language Modelling with Pixels

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.174488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.174488Z digest=sha256:751f358f06e80932bce11b6fd536dcb6074f14117378142b68fd5d500c7ad255

Observation 84a023d8-5388-4b65-8205-75d150f0180a · outbound

This paper cites Implicit Stacked Autoregressive Model for Video Prediction.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Implicit Stacked Autoregressive Model for Video Prediction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.180810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.180810Z digest=sha256:7f44948f1b5544594cdf887cf82908beee75416f1a732588c5ca847abd163a40

Observation 3645c40e-e2f8-41a7-8b0e-74fd0f97f629 · outbound

This paper cites FLAVA: A Foundational Language And Vision Alignment Model.

Everything is a Video: Unifying Modalities through Next-Frame Prediction FLAVA: A Foundational Language And Vision Alignment Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.186534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.186534Z digest=sha256:fc537fe8e7784819316026c026dd172149a78a5597c4ff2d99a1b6acb3f1476b

Observation 9b5591c1-6ae5-4507-8aba-f7e3079ee87c · outbound

This paper cites Recursive deep models for semantic compositional- ity over a sentiment treebank.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Recursive deep models for semantic compositional- ity over a sentiment treebank

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.684208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:59:00.192904Z digest=sha256:d4b7f38134cac61c0b3819c7a5aa99f059246d6bd09b6e7ff328783d11063a65

Observation 3bc84ed7-9dfe-438c-a501-a72f9754e8c3 · outbound

This paper cites PIXAR: Auto-Regressive Language Modeling in Pixel Space.

Everything is a Video: Unifying Modalities through Next-Frame Prediction PIXAR: Auto-Regressive Language Modeling in Pixel Space

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.198753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.198753Z digest=sha256:396d40cef4ab83f008ef8f741cac3c225738306add176f8a64eb1c6a47a484ea

Observation 1f7ab565-4240-441f-b0eb-b5330b29b97d · outbound

This paper cites Transformation-Based Models of Video Sequences.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Transformation-Based Models of Video Sequences

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.204388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.204388Z digest=sha256:06e6fc74d358492bf86243df921fc7ea58ea541669ac09d050ff67acf774eff0

Observation 2e09d4d6-8b24-41be-a5a3-25a80c23e98a · outbound

This paper cites Multimodal llm enhanced cross- lingual cross-modal retrieval.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Multimodal llm enhanced cross- lingual cross-modal retrieval

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.667610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:59:00.210160Z digest=sha256:efcf7a5327d8ca9663d180bd4581d55b18daa6761a4b90bf8ad828e411409634

Observation b589a8ee-9745-4773-88b5-6140da2de8c1 · outbound

This paper cites Bovik, H.R.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Bovik, H.R

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.216372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.216372Z digest=sha256:c1cef56a5d862083f34a793f0c5d1d4280181c83ea17439e1ba472cb4ef5085f

Observation 42f29cc1-00e4-4b44-a243-a669c6a0c5eb · outbound

This paper cites Visual question answer- ing: A survey of methods and datasets.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Visual question answer- ing: A survey of methods and datasets

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.222174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.222174Z digest=sha256:6c456dfb46d4140c5d319d041199fb84440abec19563a8df55c6e73523b7c76a

Observation 59f3cd8b-2217-46d2-afab-376787778509 · outbound

This paper cites Pixel Sentence Representation Learning.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Pixel Sentence Representation Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.228785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.228785Z digest=sha256:fa88f3108dec31f9c476b5eae0a061a41f640895ac3b48cc3379cbbd5137d2ef

Observation b2a3e8b8-2ffb-492f-a9ff-e7108e387f74 · outbound

This paper cites Videogpt: Video generation using vq-vae and trans- formers, 2021.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Videogpt: Video generation using vq-vae and trans- formers, 2021

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.235445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.235445Z digest=sha256:2ac1577b97c125c14012d23ca30f0d9636de9f82c3f611a5497c7ed03bbb2324

Observation b11476c7-6673-4d84-afec-b5c57ff54520 · outbound

This paper cites Colormnet: A memory-based deep spatial-temporal feature propagation network for video colorization.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Colormnet: A memory-based deep spatial-temporal feature propagation network for video colorization

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.621333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:59:00.240859Z digest=sha256:b202e19274ee63527d62a577d929c38414407f5c8b9d8dcba5f2779b0518c879

Observation 8b438b27-b62c-4236-8da5-c7e2ddcf297b · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

Everything is a Video: Unifying Modalities through Next-Frame Prediction CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.245986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.245986Z digest=sha256:f7d0a26ece465d37eff8b044a76802f695ba360e55a9cc0f4f1940502b526b95

Observation bcc370f7-493e-411a-bdae-cb5f5a8c8bcf · outbound

This paper cites Akin Yilmaz and Ahmet Murat Tekalp.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Akin Yilmaz and Ahmet Murat Tekalp

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.604360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:59:00.252101Z digest=sha256:f59b61d209a59740304f21c26954415b420c25b0220a18282d3f05600c1a8bda

Observation 6b0c4265-dd2e-4798-bac9-51cb76b4d9b6 · outbound

This paper cites Colorful image colorization.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Colorful image colorization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.587679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:59:00.257545Z digest=sha256:ba4d1a4ea553ff99b547a220f5cfb1e7528d52e22371b0ac8fd61a796b089dc2

Observation 3902d12c-2589-4e0d-8d3d-082882a42ac4 · outbound

This paper cites Bag of Tricks for Effective Language Model Pretraining and Downstream Adaptation: A Case Study on GLUE.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Bag of Tricks for Effective Language Model Pretraining and Downstream Adaptation: A Case Study on GLUE

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-12T19:59:00.323391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:59:00.263231Z digest=sha256:b414fdfbf297425e0678b7b96586cd1644a20bb12b6c2df9b2394429ad85963c

Observation 2856d329-c2aa-4a9e-be1c-342be7dafe80 · outbound

This paper cites A survey on vqa: Datasets and approaches.

Everything is a Video: Unifying Modalities through Next-Frame Prediction A survey on vqa: Datasets and approaches

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.569716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T19:59:00.269362Z digest=sha256:48e340754df3185b7376a1f7268c156d949aa5e0dbbeb05ea3ead3869735c153

Pith citing papers

Observation 3c62bd83-a8a4-42bc-a3c9-82446053c66c · inbound

Unraveling Spatio-Temporal Foundation Models via the Pipeline Lens: A Comprehensive Review cites this paper.

Unraveling Spatio-Temporal Foundation Models via the Pipeline Lens: A Comprehensive Review Everything is a Video: Unifying Modalities through Next-Frame Prediction

Reference 135

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:49:53.110190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:49:42.989079Z digest=sha256:203caf38e3e35e45f9fd89014dbbaeaa605778823922a3fc19e0115f143e0f0f