Pith. sign in

Paper Citation Record · LEDGER

V-RAE: Rethinking Video Latent Spaces for Generation

As of 15 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2608.13556.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.13556 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:16:07.560939Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1ff04b0-f62e-4edf-8228-4f445fd41fea · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

V-RAE: Rethinking Video Latent Spaces for Generation Cosmos World Foundation Model Platform for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.419865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.419865Z digest=sha256:0cf1c9f695569f8cf72f98b6d4c152b6f5069bef8ae48cef41ed4f1932f97c43

Observation 4447f9fa-a3c3-431d-8097-f730f0dfeb96 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

V-RAE: Rethinking Video Latent Spaces for Generation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.432474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.432474Z digest=sha256:b7ab051593d3b7ec546097fda8a824ba754aaf6b930a84a1786edb96d99759b6

Observation bdcfce0e-4edc-4c09-a229-1410ffb3817a · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

V-RAE: Rethinking Video Latent Spaces for Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.437307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.437307Z digest=sha256:b9e0a63d10090392bde0c16b932f850be79ed71a1f64329d81cc63d2088c6095

Observation 822e9cf7-499b-4321-8c06-45437b1d7d2c · outbound

This paper cites The latent perturbation is used only as a reconstruction-training augmentation; evaluation, latent-statistics estimation, and latent video generation all use clean latents.

V-RAE: Rethinking Video Latent Spaces for Generation The latent perturbation is used only as a reconstruction-training augmentation; evaluation, latent-statistics estimation, and latent video generation all use clean latents

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:16:08.235615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:16:07.556599Z digest=sha256:d980995295b8cb4fad302dd29562fe60f46a6ac97de08aec640657f74156e99a

Observation 7be6db4f-9549-4388-8706-111c5887fa99 · outbound

This paper cites The Kinetics Human Action Video Dataset.

V-RAE: Rethinking Video Latent Spaces for Generation The Kinetics Human Action Video Dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.449932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.449932Z digest=sha256:b7634f5fb25bb4cf2dc1e565a03268f9a897142613b25f5e563b6e9df0dd8dc1

Observation c640daf3-f0a5-4252-9f5c-4bb9c82b55b5 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

V-RAE: Rethinking Video Latent Spaces for Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.453726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.453726Z digest=sha256:1f68f16b98aeb16f5113b349613f906312c3a4e13e018a5ba8bcbc1443e5e6f3

Observation 4b88751b-242a-4767-b78d-6090785898a2 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

V-RAE: Rethinking Video Latent Spaces for Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.458705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.458705Z digest=sha256:b457dca0678ff1b2ac3ce4dde4e40033ff20ad9616fab419bb580da3d18a973a

Observation 32c377c0-0683-4177-a23f-c8d74756ed76 · outbound

This paper cites Atoken: A unified tokenizer for vision.arXiv preprint arXiv:2509.14476,.

V-RAE: Rethinking Video Latent Spaces for Generation Atoken: A unified tokenizer for vision.arXiv preprint arXiv:2509.14476,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.463799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.463799Z digest=sha256:25643657c6ea030c886515c49984cf9f1ad384ea8b88c78c8964d15daaa52bf5

Observation 1715a996-a8eb-4e7c-93d8-7855af826bc1 · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

V-RAE: Rethinking Video Latent Spaces for Generation Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.468746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.468746Z digest=sha256:ed5bef313d77c36fcbbfd71b1cf8b11060044aa48120758ca67fba11f207655e

Observation 24aedaf7-f34f-43c1-988b-43d5371d167e · outbound

This paper cites DINOv3.

V-RAE: Rethinking Video Latent Spaces for Generation DINOv3

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.478488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.478488Z digest=sha256:107d4e670f5837df50b8ceaa3cb21b65a7397443f0f252d7f3c52c7dbd2145eb

Observation bce3e1be-f1a3-454b-b373-e8c803aa75d7 · outbound

This paper cites Improved Baselines with Representation Autoencoders.

V-RAE: Rethinking Video Latent Spaces for Generation Improved Baselines with Representation Autoencoders

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.483649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.483649Z digest=sha256:fe2c5af0273338708cbcefda78a69c3c5d022f1ec4ea5fdb81ab473eba56f37a

Observation 4c927a75-48a4-446e-a8db-1c66ec61b5ed · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

V-RAE: Rethinking Video Latent Spaces for Generation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.492798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.492798Z digest=sha256:e607aa947bb472378a08db82486deb6eaf364659aba424542674973917a8453a

Observation 1a1d4876-fec9-4f39-b268-392b5243f904 · outbound

This paper cites Scaling text-to-image diffusion transformers with representation autoencoders.arXiv preprint arXiv:2601.16208,.

V-RAE: Rethinking Video Latent Spaces for Generation Scaling text-to-image diffusion transformers with representation autoencoders.arXiv preprint arXiv:2601.16208,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.501394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.501394Z digest=sha256:bda09f8ca401a2f228a32fcfe9861acee4cfe05f4adfcb001c947169da779fa2

Observation dea9f47f-51ca-4dea-94e4-9db511d75c7a · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

V-RAE: Rethinking Video Latent Spaces for Generation SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.504838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.504838Z digest=sha256:e33e252fa0b939560469fcb36c14abadfbbdaa11a59276363b18ab50d1a3e14b

Observation 6ed848e9-08e3-4044-99f2-b9f7936e713e · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

V-RAE: Rethinking Video Latent Spaces for Generation Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.509532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.509532Z digest=sha256:a560b97c20ebff9febfa696ef74d4563766c191cdf34e8fe46f5d26441ecf0ea

Observation 0d46f0a2-d88e-4157-87d3-ebf0d5b62093 · outbound

This paper cites Larp: Tokenizing videos with a learned autoregressive generative prior.

V-RAE: Rethinking Video Latent Spaces for Generation Larp: Tokenizing videos with a learned autoregressive generative prior

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:16:08.260044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:16:07.518767Z digest=sha256:847975d9a524c55015446eddb5cb8d10d024b1a792a812ab2597c783d654b460

Observation c0fc0f33-5340-4164-8dc2-b8294f5720c0 · outbound

This paper cites VidTwin: Video VAE with Decoupled Structure and Dynamics.

V-RAE: Rethinking Video Latent Spaces for Generation VidTwin: Video VAE with Decoupled Structure and Dynamics

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.523981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.523981Z digest=sha256:0117d942b215df04c7de0e782b9b653cc4210bb6d436d317f01f9059cc50b90f

Observation 5c801366-1054-466b-b96e-bbc3f9721aa7 · outbound

This paper cites Making Reconstruction FID Predictive of Diffusion Generation FID.

V-RAE: Rethinking Video Latent Spaces for Generation Making Reconstruction FID Predictive of Diffusion Generation FID

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.532657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.532657Z digest=sha256:9512ca1044ec99d44c89dd907d5028b349f5df91d54c6f93c9e3c4cac6611650

Observation e19e2ac3-fd39-4026-88f7-8aa24b5b4913 · outbound

This paper cites Cogvideox: Text-to-video diffusion models with an expert transformer.

V-RAE: Rethinking Video Latent Spaces for Generation Cogvideox: Text-to-video diffusion models with an expert transformer

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:16:08.247634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:16:07.536617Z digest=sha256:24203295ba5b11372a2f72eaad73efe7115f1e50cb5644137b42a494ff80be26

Observation 07ba8789-d4c7-4792-84dd-94679d813625 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

V-RAE: Rethinking Video Latent Spaces for Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.540481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.540481Z digest=sha256:9f73951b61f591d9868596e3b5709cfc2c7be84c8e8d7650c77240c46bc8688c

Observation 3ce2d571-5fb6-4834-853f-f71141afcf86 · outbound

This paper cites Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think.

V-RAE: Rethinking Video Latent Spaces for Generation Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.544422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.544422Z digest=sha256:b77c18562e2027a3cbcee5bdd3c238a2de4c3608d8d70fe64d57ffcec5e3ca17

Observation d574185f-1288-415a-aa4d-085b3074756b · outbound

This paper cites Diffusion Transformers with Representation Autoencoders.

V-RAE: Rethinking Video Latent Spaces for Generation Diffusion Transformers with Representation Autoencoders

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.548558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.548558Z digest=sha256:9e7fe50fabb771e137399402816f1ab2ce5f7ca7cd507fb05108509b03020d70

Observation 668e9aa1-a0a4-4e1e-bc87-0ea517977c57 · outbound

This paper cites Efficient universal perception encoder.arXiv preprint arXiv:2603.22387,.

V-RAE: Rethinking Video Latent Spaces for Generation Efficient universal perception encoder.arXiv preprint arXiv:2603.22387,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.552563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.552563Z digest=sha256:9d76281880dc9e7081344319613c9c0a6542e102a06539f44f6550ba8e63e0cf

Observation f503f54a-b211-4ffd-9eeb-50438244f92e · outbound

This paper cites Only the input channel count and corresponding time shift change for EUPE-B.

V-RAE: Rethinking Video Latent Spaces for Generation Only the input channel count and corresponding time shift change for EUPE-B

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:16:08.221888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:16:07.560939Z digest=sha256:564f62172186fb1a70c3e716456d5049f522ea8bf406218d6a68272fc2586e9e

Observation 4a60b787-5796-4cda-9d61-5ab9045089ec · outbound

This paper cites Representation entanglement for generation: Training diffusion transformers is much easier than you think.Advances in Neural Information Processing Systems, 38:7714–7743, 2025a.

V-RAE: Rethinking Video Latent Spaces for Generation Representation entanglement for generation: Training diffusion transformers is much easier than you think.Advances in Neural Information Processing Systems, 38:7714–7743, 2025a

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.528689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.528689Z digest=sha256:fb33658705bf45f6f9890f0238ddd9224941c151b85f15a148950b56dd874bc5

Observation caceaa4c-1ffa-4ba9-be4c-a72b7725b96e · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

V-RAE: Rethinking Video Latent Spaces for Generation RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.497419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.497419Z digest=sha256:be11a4c7e7e54491257637ec5b694f0c32a6d049cd27ae44479c9f99f7b2e8b1

Observation a0fe687e-725d-44d6-b5c2-0b3ce4674630 · outbound

This paper cites Dera: Decoupled representation alignment for video tokenization.arXiv preprint arXiv:2512.04483,.

V-RAE: Rethinking Video Latent Spaces for Generation Dera: Decoupled representation alignment for video tokenization.arXiv preprint arXiv:2512.04483,

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.445929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.445929Z digest=sha256:06670e5a4b5368d65755a859893c538c6d03c8be200312758fbf1e9fee8fd0a9

Observation ddcbd4b4-db9c-40cf-b329-2bc79b53a10b · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

V-RAE: Rethinking Video Latent Spaces for Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.514157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.514157Z digest=sha256:7669342e0395865f75e012fe2a00621e17ace3242b2f2fabf9b57049241323d4

Observation 6d972246-bfff-44ac-bf89-e1863b9d3ed9 · outbound

This paper cites A Short Note about Kinetics-600.

V-RAE: Rethinking Video Latent Spaces for Generation A Short Note about Kinetics-600

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.441723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.441723Z digest=sha256:1a856b8cd296eeeee13671af488d44f8de7e92d8e230af64e120f3b65f612959

Observation 4a781aca-ad0e-49de-a977-ae39dad674cd · outbound

This paper cites V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning.

V-RAE: Rethinking Video Latent Spaces for Generation V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.473562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.473562Z digest=sha256:097407ef8e098c73ed9764e1333a29b49824d9bd176e171eb7563fd25dc07fa0

Observation ff700b73-2f1e-4f9f-8c12-40d2fd0a3c5c · outbound

This paper cites CoVLA: Comprehensive vision-language-action dataset for autonomous driving.

V-RAE: Rethinking Video Latent Spaces for Generation CoVLA: Comprehensive vision-language-action dataset for autonomous driving

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:16:08.273629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:16:07.427895Z digest=sha256:8c56fa7e9e24e2fe204871443ecf59fc335b32ad741685e7b74139edfa3f57f6

Observation 83a764d7-a9ff-4c8b-a812-46ac2f97e315 · outbound

This paper cites Improving the Diffusability of Autoencoders.

V-RAE: Rethinking Video Latent Spaces for Generation Improving the Diffusability of Autoencoders

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.488440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.488440Z digest=sha256:cd0c355bf8752defeec8e175aae91f6c4e1f8df7ea1296a5ab0f98a4de0f4bb8

Pith citing papers

No inbound Pith citation observations are available.