Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:48:08.265907Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 11 inbound Pith citation observations for arXiv:2501.09755.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:48:08.265907Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:32:54.714314Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-07T15:43:53.781661Z
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b35607cf-99de-48b5-8b3b-0741941ed404 · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a76d1005-9940-4863-8258-c894aad87378 · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Unified Auto-Encoding with Masked Diffusion
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12bdf26c-6073-4028-ae31-c78688c6010b · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Autoregressive Image Generation using Residual Quantization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dc2b18a-3d6d-4a22-b309-6cf0235b761f · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Decoupled Weight Decay Regularization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba747b7c-4db7-4479-ab14-cbe8c7d9875d · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Finite Scalar Quantization: VQ-VAE Made Simple
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7918047c-6bbd-4215-9854-a9c167ced876 · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation A tokenizer designed for efficient processing of large-scale datasets
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 20f6135c-b6dd-4ffb-bb08-13f31de8116d · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be8d39b6-1658-4929-9f09-f306e7ae52f5 · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Movie Gen: A Cast of Media Foundation Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7f9c7a1-135f-458a-8c35-6e17b82a5169 · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Score-Based Generative Modeling through Stochastic Differential Equations
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b115cdb-cedb-4424-86fc-6f46695f1ae4 · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc066984-6293-4d49-b410-98d8ce1f3cd8 · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64bf2786-e14d-45ad-a00e-e252a9d2fbff · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Video Occupancy Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07b1dca9-a14f-4664-8329-b74f503ec422 · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Towards Accurate Generative Models of Video: A New Metric & Challenges
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d32b1449-baa1-49cb-9d7c-a05999221b60 · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation ElasticTok: Adaptive Tokenization for Image and Video
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6737044d-731d-4629-8795-d114f4c6a526 · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4a97e9a-7635-442f-95a5-32fa1a303fa6 · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85ac8289-37b4-47be-b2f2-2893a1459a15 · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Apollo: An Exploration of Video Understanding in Large Multimodal Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac779b7f-16db-41e3-809e-1d29cf7680aa · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4ec8c220-c5f8-432d-b2a6-3b331f242b95 · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation The architecture consists of Transformer blocks (Vaswani et al.,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f7d49cc9-be21-4894-a707-b58201a7a151 · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Additionally, we integrate video processing code from Apollo (Zohar et al.,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4c64ea08-317e-4f7c-ba52-704646a3792c · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Overall, ViTok leverages advanced training techniques and architectural innovations to achieve state-of-the-art performance in image and video reconstruction and generation tasks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d8f0a97d-cdf9-450a-863a-f6cf45e8874b · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Consequently, it provides an alternative to Simple ViTok with greater control over the latent space
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 68eb0807-21bd-480f-b0eb-372a09a15e8c · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas
Reference 2004
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33d2fd40-5049-4518-a1e4-e766afda89cb · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior
Reference 2008
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6997ee64-89eb-4561-a6a3-3cacb09b1727 · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Classifier-Free Diffusion Guidance
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7abd60d-e829-46b7-a0c5-e97216cee5b6 · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57359005-bf6a-4c54-9d7d-5ad3254b298b · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation GLU Variants Improve Transformer
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca83b6d8-6896-4a8f-a7c8-625e7de5a3cf · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Improving neural networks by preventing co-adaptation of feature detectors
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82e18de8-d31e-4e50-a8a1-57ac643aa548 · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation A Style-Based Generator Architecture for Generative Adversarial Networks
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90d7f7c3-a8f4-4ddc-a334-57124dd7e326 · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Perceptual losses for real-time style transfer and super-resolution
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 49525f17-50ad-417c-a538-189bf0fcf3f4 · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Layer Normalization
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec66f98c-c9cb-4194-be2a-452ff883e8e8 · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Video generation models as world simulators
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0bf26328-b53f-418d-967c-769f8a5515ca · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b843b63e-f2d6-4f91-82ec-6f8eff33911f · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive Transformer
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b55997f5-f1f5-4024-82d7-7b2d4262b30b · outbound
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09dadb60-9087-4483-9174-0060d7155eec · inbound
MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a76ca4bf-c885-4701-ab4f-dcb157ef389b · inbound
Flow marching for a generative PDE foundation model Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2aa396cf-8da4-4bbc-b54a-8e497bd7139a · inbound
TC-AE: Unlocking Token Capacity for Deep Compression Autoencoders Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ac0c1624-78af-40ce-9719-35b212adf4b2 · inbound
Latent-Compressed Variational Autoencoder for Video Diffusion Models Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b0ace6f3-d26d-4699-b281-320374b491fa · inbound
ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parameters Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d0055b1f-9d88-459b-bc31-c64cb3afe703 · inbound
LiVeAction: a Lightweight, Versatile, and Asymmetric Neural Codec Design for Real-time Operation Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7e29255c-bf86-4f97-9053-82a549e05319 · inbound
What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3dbd517c-af80-47d7-b9be-5335072a6892 · inbound
Yeti: A compact protein structure tokenizer for reconstruction and multi-modal generation Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a0bbab4b-173f-4e72-9dca-99de3929f7dc · inbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 51798922-a3bf-4a4f-9b45-9dd38f20c13a · inbound
Balancing Image Compression and Generation with Bootstrapped Tokenization Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c02e9aa8-1cbf-473b-a748-2a9cb2180d83 · inbound
Multiplayer Interactive World Models with Representation Autoencoders Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.