Pith. sign in

Paper Citation Record · LEDGER

V-RAE: Rethinking Video Latent Spaces for Generation

As of 15 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2608.13556.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.13556 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:16:07.560939Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1ff04b0-f62e-4edf-8228-4f445fd41fea · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

V-RAE: Rethinking Video Latent Spaces for Generation Cosmos World Foundation Model Platform for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.419865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.419865Z digest=sha256:162c217105ada80a4212536927dfa5ebc0d9a20c95e5acd026b608b1201ecf68

Observation 4447f9fa-a3c3-431d-8097-f730f0dfeb96 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

V-RAE: Rethinking Video Latent Spaces for Generation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.432474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.432474Z digest=sha256:ac75092a2aa47da2110049c5bf275ca5d37acd1f1979110355c8a576eb78599e

Observation bdcfce0e-4edc-4c09-a229-1410ffb3817a · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

V-RAE: Rethinking Video Latent Spaces for Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.437307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.437307Z digest=sha256:1c224bfbd41efafe6684153541b13ae3fbc3a670ea34d3d7a12e050327174fc5

Observation 822e9cf7-499b-4321-8c06-45437b1d7d2c · outbound

This paper cites The latent perturbation is used only as a reconstruction-training augmentation; evaluation, latent-statistics estimation, and latent video generation all use clean latents.

V-RAE: Rethinking Video Latent Spaces for Generation The latent perturbation is used only as a reconstruction-training augmentation; evaluation, latent-statistics estimation, and latent video generation all use clean latents

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:16:08.235615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:16:07.556599Z digest=sha256:9046a05a9e9854f17f5d71d146409bec4c016a27ae1ed68e406ea7b2b41eb984

Observation 7be6db4f-9549-4388-8706-111c5887fa99 · outbound

This paper cites The Kinetics Human Action Video Dataset.

V-RAE: Rethinking Video Latent Spaces for Generation The Kinetics Human Action Video Dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.449932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.449932Z digest=sha256:a3edec38bc6cd926ba3175ede62cc0d6b2f3aa89036605fc21f2f6118508677e

Observation c640daf3-f0a5-4252-9f5c-4bb9c82b55b5 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

V-RAE: Rethinking Video Latent Spaces for Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.453726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.453726Z digest=sha256:ba1226c530f7a2766ce471e4ac5e3393a4e48c62dac64d288b2a95810f0ce60e

Observation 4b88751b-242a-4767-b78d-6090785898a2 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

V-RAE: Rethinking Video Latent Spaces for Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.458705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.458705Z digest=sha256:29076f1847a9989017ee5f679daa8f7d0dbc0e2c37d394c63bbb1141970cda54

Observation 32c377c0-0683-4177-a23f-c8d74756ed76 · outbound

This paper cites Atoken: A unified tokenizer for vision.arXiv preprint arXiv:2509.14476,.

V-RAE: Rethinking Video Latent Spaces for Generation Atoken: A unified tokenizer for vision.arXiv preprint arXiv:2509.14476,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.463799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.463799Z digest=sha256:bad4e595b5041cd68bc2db17a3707c4fad68527afd3249509e61e84a2348ff23

Observation 1715a996-a8eb-4e7c-93d8-7855af826bc1 · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

V-RAE: Rethinking Video Latent Spaces for Generation Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.468746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.468746Z digest=sha256:084f24e2b68f32c23baa9709dca137757c20df457c420555750951d4eecac095

Observation 24aedaf7-f34f-43c1-988b-43d5371d167e · outbound

This paper cites DINOv3.

V-RAE: Rethinking Video Latent Spaces for Generation DINOv3

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.478488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.478488Z digest=sha256:143a50c5b907665c7b7ee588d1518cd7e42b7c82ec37703b4b8802e4b8c3e529

Observation bce3e1be-f1a3-454b-b373-e8c803aa75d7 · outbound

This paper cites Improved Baselines with Representation Autoencoders.

V-RAE: Rethinking Video Latent Spaces for Generation Improved Baselines with Representation Autoencoders

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.483649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.483649Z digest=sha256:0d3901bf1774a69912a0f445b048dc2e651c058b4b4bd5f4a437eacbec52e35e

Observation 4c927a75-48a4-446e-a8db-1c66ec61b5ed · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

V-RAE: Rethinking Video Latent Spaces for Generation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.492798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.492798Z digest=sha256:98e34c1c47f4d19872465f8c93048c9a0c29e76ce9d69e8192c3b15c49b433e7

Observation 1a1d4876-fec9-4f39-b268-392b5243f904 · outbound

This paper cites Scaling text-to-image diffusion transformers with representation autoencoders.arXiv preprint arXiv:2601.16208,.

V-RAE: Rethinking Video Latent Spaces for Generation Scaling text-to-image diffusion transformers with representation autoencoders.arXiv preprint arXiv:2601.16208,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.501394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.501394Z digest=sha256:a5a7b1a5f4f406fe5b1cd548ed0732119698ac09cb99ab4de47963fcea816969

Observation dea9f47f-51ca-4dea-94e4-9db511d75c7a · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

V-RAE: Rethinking Video Latent Spaces for Generation SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.504838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.504838Z digest=sha256:a24531f3224301f17dae120e3f1feb0a2c9797e7a58e02fb0240b272b3d323d7

Observation 6ed848e9-08e3-4044-99f2-b9f7936e713e · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

V-RAE: Rethinking Video Latent Spaces for Generation Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.509532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.509532Z digest=sha256:d1656485f3da5ba3e2ef32998267ce2cfb18af4ef12ff93e06394bacc9cff643

Observation 0d46f0a2-d88e-4157-87d3-ebf0d5b62093 · outbound

This paper cites Larp: Tokenizing videos with a learned autoregressive generative prior.

V-RAE: Rethinking Video Latent Spaces for Generation Larp: Tokenizing videos with a learned autoregressive generative prior

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:16:08.260044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:16:07.518767Z digest=sha256:726ad230e81affe7e72c8c18e467343e758caee0e9e929afc8aaaa486926e737

Observation c0fc0f33-5340-4164-8dc2-b8294f5720c0 · outbound

This paper cites VidTwin: Video VAE with Decoupled Structure and Dynamics.

V-RAE: Rethinking Video Latent Spaces for Generation VidTwin: Video VAE with Decoupled Structure and Dynamics

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.523981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.523981Z digest=sha256:7123b5860bd7cf323ca50ad757d296817f420af095cd348d35d07bf2c4556e05

Observation 5c801366-1054-466b-b96e-bbc3f9721aa7 · outbound

This paper cites Making Reconstruction FID Predictive of Diffusion Generation FID.

V-RAE: Rethinking Video Latent Spaces for Generation Making Reconstruction FID Predictive of Diffusion Generation FID

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.532657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.532657Z digest=sha256:144b2c767fe77f1b8c1f6c4fd97f059c00d2365eee12b20e2d0521a2f9e8444e

Observation e19e2ac3-fd39-4026-88f7-8aa24b5b4913 · outbound

This paper cites Cogvideox: Text-to-video diffusion models with an expert transformer.

V-RAE: Rethinking Video Latent Spaces for Generation Cogvideox: Text-to-video diffusion models with an expert transformer

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:16:08.247634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:16:07.536617Z digest=sha256:d9c88d76fc8b130de767f4e335954d3e5820f173405bba0b1a6fa9053a5699c9

Observation 07ba8789-d4c7-4792-84dd-94679d813625 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

V-RAE: Rethinking Video Latent Spaces for Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.540481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.540481Z digest=sha256:2f5dc4cadce48f7f65b0f8c89206450e19079121383330ca448512d5a0b8b044

Observation 3ce2d571-5fb6-4834-853f-f71141afcf86 · outbound

This paper cites Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think.

V-RAE: Rethinking Video Latent Spaces for Generation Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.544422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.544422Z digest=sha256:d64c038035a5cf4a7594ee35511e813dd919bdf34ae270e01ecd25f870b5d673

Observation d574185f-1288-415a-aa4d-085b3074756b · outbound

This paper cites Diffusion Transformers with Representation Autoencoders.

V-RAE: Rethinking Video Latent Spaces for Generation Diffusion Transformers with Representation Autoencoders

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.548558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.548558Z digest=sha256:0924734b07ab361a158f22f69effa677e8ae817b678db9a40a6ea78bf8afccdc

Observation 668e9aa1-a0a4-4e1e-bc87-0ea517977c57 · outbound

This paper cites Efficient universal perception encoder.arXiv preprint arXiv:2603.22387,.

V-RAE: Rethinking Video Latent Spaces for Generation Efficient universal perception encoder.arXiv preprint arXiv:2603.22387,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.552563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.552563Z digest=sha256:8ddd1d472b5fcb5e25fd97f05c37831c11b28144809a3966279d9b54901828bc

Observation f503f54a-b211-4ffd-9eeb-50438244f92e · outbound

This paper cites Only the input channel count and corresponding time shift change for EUPE-B.

V-RAE: Rethinking Video Latent Spaces for Generation Only the input channel count and corresponding time shift change for EUPE-B

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:16:08.221888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:16:07.560939Z digest=sha256:772f628c3f0a6f22c9431e51cd80bfb6727027af426903f0e752eec8dc8bd9c9

Observation 4a60b787-5796-4cda-9d61-5ab9045089ec · outbound

This paper cites Representation entanglement for generation: Training diffusion transformers is much easier than you think.Advances in Neural Information Processing Systems, 38:7714–7743, 2025a.

V-RAE: Rethinking Video Latent Spaces for Generation Representation entanglement for generation: Training diffusion transformers is much easier than you think.Advances in Neural Information Processing Systems, 38:7714–7743, 2025a

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.528689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.528689Z digest=sha256:e3c3d3d427e5d4505fbfe0ed5de53da089a42f6e17bae0c681e4442b837a1753

Observation caceaa4c-1ffa-4ba9-be4c-a72b7725b96e · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

V-RAE: Rethinking Video Latent Spaces for Generation RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.497419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.497419Z digest=sha256:eaf98818921050b668328158d2b44738955ce5ea00f292184f17227d6a359bc2

Observation a0fe687e-725d-44d6-b5c2-0b3ce4674630 · outbound

This paper cites Dera: Decoupled representation alignment for video tokenization.arXiv preprint arXiv:2512.04483,.

V-RAE: Rethinking Video Latent Spaces for Generation Dera: Decoupled representation alignment for video tokenization.arXiv preprint arXiv:2512.04483,

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.445929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.445929Z digest=sha256:29ed671d28f448ee6ee79a57d09469f9d4729222127055839371e0b241d5609d

Observation ddcbd4b4-db9c-40cf-b329-2bc79b53a10b · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

V-RAE: Rethinking Video Latent Spaces for Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.514157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.514157Z digest=sha256:4dc18a9b3736bf600d73e7b6f698a1f117e408708147e3b8c0844555333bc47d

Observation 6d972246-bfff-44ac-bf89-e1863b9d3ed9 · outbound

This paper cites A Short Note about Kinetics-600.

V-RAE: Rethinking Video Latent Spaces for Generation A Short Note about Kinetics-600

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.441723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.441723Z digest=sha256:58338577046f6638a7717636ab6bf9fe18f7337514d1b19ceaa617c1134603fe

Observation 4a781aca-ad0e-49de-a977-ae39dad674cd · outbound

This paper cites V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning.

V-RAE: Rethinking Video Latent Spaces for Generation V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.473562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.473562Z digest=sha256:6d56fbfd139457f082069a6a34b757c2b8da06967b97b4f2ae57d133cd74ac63

Observation ff700b73-2f1e-4f9f-8c12-40d2fd0a3c5c · outbound

This paper cites CoVLA: Comprehensive vision-language-action dataset for autonomous driving.

V-RAE: Rethinking Video Latent Spaces for Generation CoVLA: Comprehensive vision-language-action dataset for autonomous driving

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:16:08.273629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:16:07.427895Z digest=sha256:a3780b9aec3f7dfca841ffcffabaa00c29b5345479267329ce9049683088558d

Observation 83a764d7-a9ff-4c8b-a812-46ac2f97e315 · outbound

This paper cites Improving the Diffusability of Autoencoders.

V-RAE: Rethinking Video Latent Spaces for Generation Improving the Diffusability of Autoencoders

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.488440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.488440Z digest=sha256:a99974848b2fcaef7b1ec506b3bbdeba4b677a3974f17017ebd91aa33df56553

Pith citing papers

No inbound Pith citation observations are available.