Pith. sign in

Paper Citation Record · LEDGER

An Empirical Study of Autoregressive Pre-training from Videos

As of 14 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 6 inbound Pith citation observations for arXiv:2501.05453.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05453 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:18:08.598227Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:01:59.614460Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T05:36:39.313175Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 125e0653-d6f5-458f-8695-c1b08967dcee · outbound

This paper cites Figure 11 µ-Parameterization Learning Rate: We show that µ-Parameterization (Yang et al., 2022), we can train all width Toto models, with an single optimal learning rate of 2−7.

An Empirical Study of Autoregressive Pre-training from Videos Figure 11 µ-Parameterization Learning Rate: We show that µ-Parameterization (Yang et al., 2022), we can train all width Toto models, with an single optimal learning rate of 2−7

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:09.071666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:18:08.598227Z digest=sha256:c5429e5402978f86713c767b73b16ad69163ac431fbb651bcda1904cdf4c3f28

Observation dcb101e3-dee2-4aba-ae17-88454352a2c4 · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

An Empirical Study of Autoregressive Pre-training from Videos A Short Note on the Kinetics-700 Human Action Dataset

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.404695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.404695Z digest=sha256:2703b9f36ed325dfab01111c8d1200228147aa6fba0f76e415ddabacd9e2b57e

Observation 125c856e-165d-48cf-bc99-03264d6ea67a · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

An Empirical Study of Autoregressive Pre-training from Videos Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.440032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.440032Z digest=sha256:f7d645fdc45e59f091d3070b36691e6a027e17b2af82f393a49a846c224cf3df

Observation 28df9a01-00f9-4805-8583-3ea9508be52f · outbound

This paper cites Scaling Laws for Autoregressive Generative Modeling.

An Empirical Study of Autoregressive Pre-training from Videos Scaling Laws for Autoregressive Generative Modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.446082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.446082Z digest=sha256:fb1738ae1c91066f0ec36cebc29712143985e778c63fcf0745d0f63beb2a294a

Observation 173044f7-051f-489e-afce-93349f30517c · outbound

This paper cites Perceptual losses for real-time style transfer and super-resolution.

An Empirical Study of Autoregressive Pre-training from Videos Perceptual losses for real-time style transfer and super-resolution

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:09.111669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:18:08.454097Z digest=sha256:4451308f5503ce0343d2b27389a3716e2e3352291613ca75f7bc8557d6fe8299

Observation c7e9951d-d7cc-4d3d-80e4-1c1c40c710b5 · outbound

This paper cites Network In Network.

An Empirical Study of Autoregressive Pre-training from Videos Network In Network

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.466279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.466279Z digest=sha256:cd97c6557ae3df0037653d1ef67aab5ae2665f63aa8e100df6c1fd10970bf945

Observation de9ade04-c8b0-444b-bf9b-d83bb87b79ba · outbound

This paper cites In-context Learning and Induction Heads.

An Empirical Study of Autoregressive Pre-training from Videos In-context Learning and Induction Heads

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.480459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.480459Z digest=sha256:51a168a28d610c527508181fe20fe06dbd5bb9238901563e98667ba63a2c331d

Observation ba918768-af36-4d98-bca1-acfc8a857c47 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

An Empirical Study of Autoregressive Pre-training from Videos DINOv2: Learning Robust Visual Features without Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.486951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.486951Z digest=sha256:c3af3f69d42208e8a6b6807c3af0130a4adcadbea984882beab52fce5820a8a4

Observation 903a5970-8225-426e-950a-7dc653e6b249 · outbound

This paper cites Video (language) modeling: a baseline for generative models of natural videos.

An Empirical Study of Autoregressive Pre-training from Videos Video (language) modeling: a baseline for generative models of natural videos

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.497076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.497076Z digest=sha256:58138fc44414021670dd412b941b66f10861be3f936ed3bf35142209824e1ca2

Observation 2069eb2d-81af-456d-a1ef-0c1eb791918d · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

An Empirical Study of Autoregressive Pre-training from Videos Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.513202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.513202Z digest=sha256:dfd7559d58c4c0fd98d6d9b15eb3c9644754d12fd418ba26bfa33c884f66fe96

Observation 5e1b2c61-5c8c-477a-91d1-e1a64246ce42 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

An Empirical Study of Autoregressive Pre-training from Videos LLaMA: Open and Efficient Foundation Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.518021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.518021Z digest=sha256:001c9a59c79c4caf9dacb080737d078c5a809db8c0aee6c1b163ef9c6327fdeb

Observation 28ce68f5-b37b-4853-9ffd-55225236cd83 · outbound

This paper cites Scaling Autoregressive Video Models.

An Empirical Study of Autoregressive Pre-training from Videos Scaling Autoregressive Video Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.532079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.532079Z digest=sha256:cd80b01d88628dfcef28d7676adefc286beef97fda8ab6d04c8d80da6eac72ac

Observation 34ded5a1-c0ae-4b11-b6f1-e0e73f66d51e · outbound

This paper cites Masked Visual Pre-training for Motor Control.

An Empirical Study of Autoregressive Pre-training from Videos Masked Visual Pre-training for Motor Control

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.542978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.542978Z digest=sha256:26dcdb49457014ae515c4b1711fc3e44e7658958d219625ac10969823f886ba1

Observation f7754097-acf0-4569-b106-0bc45a05e8a2 · outbound

This paper cites TFCNet: Temporal Fully Connected Networks for Static Unbiased Temporal Reasoning.

An Empirical Study of Autoregressive Pre-training from Videos TFCNet: Temporal Fully Connected Networks for Static Unbiased Temporal Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.557957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.557957Z digest=sha256:dba9b3cb28ceb4bdbcd101ddc7af4a3f39f2e520c9e7ad16f0bdbe35a725c1b9

Observation 7994cab7-3320-4d49-b9b2-b2bbd6e767b5 · outbound

This paper cites Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model.

An Empirical Study of Autoregressive Pre-training from Videos Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.581095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.581095Z digest=sha256:5798bbc95aedb6e64cc47fcff61d4142dc0565d4b395302b7f55b228295070ce

Observation 3ce465dc-f3f4-477f-8982-9f026c46cf69 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Autoregressive Pre-training from Videos Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:18:09.092643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:18:08.591288Z digest=sha256:932076426406ee2e2a35d99ae8563e99c11e7c7dcaf0ed239baa89b3af6ac03b

Observation 19e7aba7-a982-4765-9198-3a9749b7aa97 · outbound

This paper cites GLU Variants Improve Transformer.

An Empirical Study of Autoregressive Pre-training from Videos GLU Variants Improve Transformer

Reference 1951

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.507922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.507922Z digest=sha256:66564b600f7679cb0c4cfd071badc92721a8c0de387a860881694417d4411249

Observation 5a9326d3-f23a-4ee1-96ff-47a0dd24c251 · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

An Empirical Study of Autoregressive Pre-training from Videos BEiT: BERT Pre-Training of Image Transformers

Reference 1954

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.393198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.393198Z digest=sha256:8cd1cd1b8ac4857c559e00ab9a2f83a3a3cd38fc9633dd07be22a98f2786f242

Observation 8d14d8f7-e564-49d2-ae11-52a2216ad168 · outbound

This paper cites Decoupled Weight Decay Regularization.

An Empirical Study of Autoregressive Pre-training from Videos Decoupled Weight Decay Regularization

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.473512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.473512Z digest=sha256:ed03cea0affad6531e16dcf9f0e14ec464eed0e14c730114a0e1d6c94cdd9e9d

Observation 5edd0f72-243b-45dd-91bc-398f37e9ce29 · outbound

This paper cites Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles.

An Empirical Study of Autoregressive Pre-training from Videos Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.503176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.503176Z digest=sha256:ca73c9c25961870ee2e5fbd986087c219aba908bc23297b598eb1ad25c7eefad

Observation 06619f79-214b-48fe-b95c-86df15aa6a53 · outbound

This paper cites The Kinetics Human Action Video Dataset.

An Empirical Study of Autoregressive Pre-training from Videos The Kinetics Human Action Video Dataset

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.460067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.460067Z digest=sha256:1996b8422c6b68f76fdefc2cbfd74d73bb1386385180a389e2de855eb343277e

Observation c63646bf-f532-4508-ba8d-8831ca267467 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

An Empirical Study of Autoregressive Pre-training from Videos InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.522588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.522588Z digest=sha256:b0f91d85d242d67821fd4ff72b0bb9acf5766dc6f2e0475fbfa47d37bb8edcc6

Observation 4c7dd642-e5fc-40ea-be9d-2e2d5d59c165 · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

An Empirical Study of Autoregressive Pre-training from Videos The 2017 DAVIS Challenge on Video Object Segmentation

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.492053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.492053Z digest=sha256:ac20d3c43f4e9b36bfbf9151de4c43788debf8dd6009acde134e423e17c471d5

Observation dfc0a216-49f9-4de8-a788-9aa47949f5c5 · outbound

This paper cites Scalable Pre-training of Large Autoregressive Image Models.

An Empirical Study of Autoregressive Pre-training from Videos Scalable Pre-training of Large Autoregressive Image Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.416521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.416521Z digest=sha256:08436b1f9f6caa0dfd367da4ab79a447fd47c27bca9030104549fa6c4aad0677

Observation eee79cca-c31b-4d1e-bb49-349427a8ebb9 · outbound

This paper cites Data Filtering Networks.

An Empirical Study of Autoregressive Pre-training from Videos Data Filtering Networks

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.428054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.428054Z digest=sha256:3beb46a23fdc82662baf7d5151e9c5b960852d6187966fb4334e18f1fd510a96

Observation 8124c0fa-9850-4d79-9fed-f9856dae87d9 · outbound

This paper cites Language Models are Few-Shot Learners.

An Empirical Study of Autoregressive Pre-training from Videos Language Models are Few-Shot Learners

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.398820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.398820Z digest=sha256:af0d537e9be362a262fed28cd565f6399ebfdfbb93161f3a3a857f1fcd8412ae

Observation ad3ab6d0-712e-41b8-933e-0cba572da75a · outbound

This paper cites CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning.

An Empirical Study of Autoregressive Pre-training from Videos CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.433520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.433520Z digest=sha256:447576323c3392506e76501aacbe3448b40e593a4cdd2be7b9bfed39ae71c81e

Observation c2568aca-c82d-40ee-89ea-7c92c58075f6 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

An Empirical Study of Autoregressive Pre-training from Videos Imagenet: A large-scale hierarchical image database

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.410401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.410401Z digest=sha256:1ad7471d22d33553c378915696e8941572914cd62a56947d05368af9cb7f8fc2

Observation cf000250-c70d-4a25-8d36-818ad539de03 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

An Empirical Study of Autoregressive Pre-training from Videos Taming transformers for high-resolution image synthesis

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:09.130230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:18:08.422538Z digest=sha256:1c113d3f47824f6ddb0e4e49da81b9afcbb1f2b77fcf4882f69e2b9d93d9506b

Pith citing papers

Observation 0ab97ceb-e762-42b1-88c2-25fb482fee79 · inbound

Poly-Autoregressive Prediction for Modeling Interactions cites this paper.

Poly-Autoregressive Prediction for Modeling Interactions An Empirical Study of Autoregressive Pre-training from Videos

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T00:01:59.614460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:01:59.614460Z digest=sha256:8534498da3ceade0fa46f512d2dc52568073b9ce9d2135a3636e8df326370104

Observation 50356802-4c7d-4a8a-a832-980328cf2c9b · inbound

Perception Encoder: The best visual embeddings are not at the output of the network cites this paper.

Perception Encoder: The best visual embeddings are not at the output of the network An Empirical Study of Autoregressive Pre-training from Videos

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:21:15.812012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T22:21:15.681336Z digest=sha256:65cd1d3f32f41d74d557d99704f30c8f05d6236493206b8944f3995a3445d9d9

Observation c075b0bc-fdc5-4161-8d3d-a77885c85539 · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics An Empirical Study of Autoregressive Pre-training from Videos

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:22:37.506893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:9c7d4e3d3221d72cb5cce830d42fa908be0503f9bc16662f301b9ac0ccf783f9

Observation 93043d74-8061-46d9-9ec8-2118051a49ef · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning An Empirical Study of Autoregressive Pre-training from Videos

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:51.155917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:5839275b48839ce2cd55bb569fc254d3318412547d254248be82b07716256985

Observation 43be2391-8b97-4bb9-a60f-a74db1b4da80 · inbound

Frozen Forecasting: A Unified Evaluation cites this paper.

Frozen Forecasting: A Unified Evaluation An Empirical Study of Autoregressive Pre-training from Videos

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T03:42:57.283557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T03:42:54.620069Z digest=sha256:b4f00a0031daefe2c5b2bb044fe64a6e2312a1931c36538c5052d4e7d0c13d3d

Observation 61840861-5c14-40c3-8312-df09a22b62aa · inbound

Uncovering the Latent Potential of Deep Intermediate Representations cites this paper.

Uncovering the Latent Potential of Deep Intermediate Representations An Empirical Study of Autoregressive Pre-training from Videos

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:36:39.314533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-25T05:36:24.743558Z digest=sha256:9a384ee61a59a570c0dea0ed2722c647efe398fba352fb4b66ecc1fdcba95351