Pith. sign in

Paper Citation Record · LEDGER

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

As of 21 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 11 inbound Pith citation observations for arXiv:2501.09755.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09755 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:48:08.265907Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:32:54.714314Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T15:43:53.781661Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b35607cf-99de-48b5-8b3b-0741941ed404 · outbound

This paper cites an unresolved cited work.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:48:08.654487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T19:48:08.261706Z digest=sha256:6c71cc325ef05dca4f5e952ceffc3369c7c7c1218280ad18fe37e4c1bfc31730

Observation a76d1005-9940-4863-8258-c894aad87378 · outbound

This paper cites Unified Auto-Encoding with Masked Diffusion.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Unified Auto-Encoding with Masked Diffusion

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.144364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.144364Z digest=sha256:e70da66d0ed9a7e6855130db2ee8b6ac3f981c7fafbd5b20e97b5166a395d823

Observation 12bdf26c-6073-4028-ae31-c78688c6010b · outbound

This paper cites Autoregressive Image Generation using Residual Quantization.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Autoregressive Image Generation using Residual Quantization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.165694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.165694Z digest=sha256:3f5edd246e325d175622dcf87fd39da54f0628f8198e88379f01fa9b5eed25f1

Observation 3dc2b18a-3d6d-4a22-b309-6cf0235b761f · outbound

This paper cites Decoupled Weight Decay Regularization.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Decoupled Weight Decay Regularization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.169664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.169664Z digest=sha256:37c75e2cd2e2cc0539608465bd2a187233666265f82c356874ac55d05820c0f6

Observation ba747b7c-4db7-4479-ab14-cbe8c7d9875d · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Finite Scalar Quantization: VQ-VAE Made Simple

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.173627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.173627Z digest=sha256:39e31345b7ce25b2a8b2c3c2d85256a6ad62876cb82c87b03b7f348973b11cad

Observation 7918047c-6bbd-4215-9854-a9c167ced876 · outbound

This paper cites A tokenizer designed for efficient processing of large-scale datasets.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation A tokenizer designed for efficient processing of large-scale datasets

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:08.731010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T19:48:08.177356Z digest=sha256:2b6d0dce084661522aa254eb91d481e5cdda6ee5ee03eafd6bfac40b55fa6cbe

Observation 20f6135c-b6dd-4ffb-bb08-13f31de8116d · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.185393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.185393Z digest=sha256:1fdc976326c7fc4c7968b95e9e304dd9d4e77e877a976cee0ebea71c837ad783

Observation be8d39b6-1658-4929-9f09-f306e7ae52f5 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Movie Gen: A Cast of Media Foundation Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.189670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.189670Z digest=sha256:aabb774568e0797a2a50e5fbb111592e96cfee2bc73216907c165bf6efcd690b

Observation d7f9c7a1-135f-458a-8c35-6e17b82a5169 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Score-Based Generative Modeling through Stochastic Differential Equations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.198018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.198018Z digest=sha256:d1976b9ebf483b1f3bfad2f530f6c53331923cdc5ad0bfca1a47a5817bd0a8db

Observation 8b115cdb-cedb-4424-86fc-6f46695f1ae4 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.202298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.202298Z digest=sha256:fd7747fbfcd2c9ddaa41f27c61d7571035996ddb0bb9490ebd613e1338861e5c

Observation dc066984-6293-4d49-b410-98d8ce1f3cd8 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.206529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.206529Z digest=sha256:7f3701d536ada124abd73a7f616c696d266f2527fb6479b36811bbf261dff8f5

Observation 64bf2786-e14d-45ad-a00e-e252a9d2fbff · outbound

This paper cites Video Occupancy Models.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Video Occupancy Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.210707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.210707Z digest=sha256:68e1272a34d2ab3a3f1c8bffa55cdd538a7a4f1a3a9a103a1e253a41476e4e91

Observation 07b1dca9-a14f-4664-8329-b74f503ec422 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.214934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.214934Z digest=sha256:778606d21d6809ad1e13a52b768dfbbf382ec66d70399a2d383dca2283e4d09e

Observation d32b1449-baa1-49cb-9d7c-a05999221b60 · outbound

This paper cites ElasticTok: Adaptive Tokenization for Image and Video.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation ElasticTok: Adaptive Tokenization for Image and Video

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.226775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.226775Z digest=sha256:15da4aa4f7b1eddcc29d10be4853c0069096147a8d357e89dbca44a443b7e07f

Observation 6737044d-731d-4629-8795-d114f4c6a526 · outbound

This paper cites Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.234771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.234771Z digest=sha256:4b4e3acedb8f0a272ceedf83d76620647c8b27ae1a7150b78198e5f37d516db4

Observation f4a97e9a-7635-442f-95a5-32fa1a303fa6 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.238826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.238826Z digest=sha256:a048ace3e198aa2f9fc6baf99b0ce7ff1899f40638e37cc32fb92a4537c589b3

Observation 85ac8289-37b4-47be-b2f2-2893a1459a15 · outbound

This paper cites Apollo: An Exploration of Video Understanding in Large Multimodal Models.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.242746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.242746Z digest=sha256:0c938a3082c5c05bc47df1108d2fefbd35508a6eeec7de9fde25adb0cffc412e

Observation ac779b7f-16db-41e3-809e-1d29cf7680aa · outbound

This paper cites an unresolved cited work.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:48:08.706258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T19:48:08.246523Z digest=sha256:e3d70c90df5b47ce91eef570d88bdfd2849c2825a118c7534eff99ced70fd258

Observation 4ec8c220-c5f8-432d-b2a6-3b331f242b95 · outbound

This paper cites The architecture consists of Transformer blocks (Vaswani et al.,.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation The architecture consists of Transformer blocks (Vaswani et al.,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:08.692782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T19:48:08.250329Z digest=sha256:5dee688d8ded63af0b3094887aa671a95ffbeae8dde445c00198d7175ead1637

Observation f7d49cc9-be21-4894-a707-b58201a7a151 · outbound

This paper cites Additionally, we integrate video processing code from Apollo (Zohar et al.,.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Additionally, we integrate video processing code from Apollo (Zohar et al.,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:08.680802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T19:48:08.253854Z digest=sha256:63d3c441ef78687a6b77cd88c8d2546234569f450b36cfbd1200decba59fd713

Observation 4c64ea08-317e-4f7c-ba52-704646a3792c · outbound

This paper cites Overall, ViTok leverages advanced training techniques and architectural innovations to achieve state-of-the-art performance in image and video reconstruction and generation tasks.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Overall, ViTok leverages advanced training techniques and architectural innovations to achieve state-of-the-art performance in image and video reconstruction and generation tasks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:08.668859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T19:48:08.257651Z digest=sha256:d586e56ec0e852931dfd13679ec680c2e27d272c4b71e0c3979f5821a4cdde2f

Observation d8f0a97d-cdf9-450a-863a-f6cf45e8874b · outbound

This paper cites Consequently, it provides an alternative to Simple ViTok with greater control over the latent space.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Consequently, it provides an alternative to Simple ViTok with greater control over the latent space

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:08.641138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T19:48:08.265907Z digest=sha256:f7ca263fc86beac5be0d19afdcb3b756e6729c1203b36d1f93db06069e19ef9b

Observation 68eb0807-21bd-480f-b0eb-372a09a15e8c · outbound

This paper cites Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.223066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.223066Z digest=sha256:33842e243c0f199fd7d3948273e9da2eac4fc978984eaa6a84ffb3346bed8377

Observation 33d2fd40-5049-4518-a1e4-e766afda89cb · outbound

This paper cites LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.219120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.219120Z digest=sha256:fa52fafad6d479cc8c8ef977c2cc1ea3b2b8881d9dd741a217a2585927eb0efa

Observation 6997ee64-89eb-4561-a6a3-3cacb09b1727 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Classifier-Free Diffusion Guidance

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.153068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.153068Z digest=sha256:01b52b9fcfa627cd4f414a72584c1510d2bdb1f10e8fd04450e25a9a570fb24a

Observation e7abd60d-e829-46b7-a0c5-e97216cee5b6 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.139919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.139919Z digest=sha256:20697496be729b146f0fc4ae3fa6e3a4780301c3308b32d664e2894bad5bee46

Observation 57359005-bf6a-4c54-9d7d-5ad3254b298b · outbound

This paper cites GLU Variants Improve Transformer.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation GLU Variants Improve Transformer

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.193450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.193450Z digest=sha256:b2c3799f883e48aceda022872ef7f0828e12c34979446100ce4fd36a8aa0d768

Observation ca83b6d8-6896-4a8f-a7c8-625e7de5a3cf · outbound

This paper cites Improving neural networks by preventing co-adaptation of feature detectors.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Improving neural networks by preventing co-adaptation of feature detectors

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.148759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.148759Z digest=sha256:eda0c6add3159215c54a988eb76b465c9663cb40cb5b0af1fd3091e8e74e9489

Observation 82e18de8-d31e-4e50-a8a1-57ac643aa548 · outbound

This paper cites A Style-Based Generator Architecture for Generative Adversarial Networks.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation A Style-Based Generator Architecture for Generative Adversarial Networks

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.161562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.161562Z digest=sha256:9f527e9cf8056a1f5e83cb35b79e7b636d05ce13096e36fd6eb75200b7c9c2ff

Observation 90d7f7c3-a8f4-4ddc-a334-57124dd7e326 · outbound

This paper cites Perceptual losses for real-time style transfer and super-resolution.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Perceptual losses for real-time style transfer and super-resolution

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:08.742934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T19:48:08.157561Z digest=sha256:bc085f2b4b397362892a9cb9bc61f1851b11df95da1614bf1a0f0db9700edadf

Observation 49525f17-50ad-417c-a538-189bf0fcf3f4 · outbound

This paper cites Layer Normalization.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Layer Normalization

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.126056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.126056Z digest=sha256:5df4d9aed941992e4f54083a93562ab3a6f66f976366d48ba28492c1218cecdd

Observation ec66f98c-c9cb-4194-be2a-452ff883e8e8 · outbound

This paper cites Video generation models as world simulators.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Video generation models as world simulators

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:08.754031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T19:48:08.130961Z digest=sha256:4a3a2dcd24cf88c9354e18916a2eeb6dc548b853eb49436076625b448f2a9086

Observation 0bf26328-b53f-418d-967c-769f8a5515ca · outbound

This paper cites Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:08.719302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T19:48:08.181404Z digest=sha256:cbb6518deee0118367b136d57bedb42a811f898b0773addbe9259f6c2d903c20

Observation b843b63e-f2d6-4f91-82ec-6f8eff33911f · outbound

This paper cites Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive Transformer.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive Transformer

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.135556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.135556Z digest=sha256:4771f6432466488044b0d8bf009f52ae0d2ec5d4bb93b323431c426a43c577c8

Observation b55997f5-f1f5-4024-82d7-7b2d4262b30b · outbound

This paper cites Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.230701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.230701Z digest=sha256:f6d2d890667b5be51e2312c0bd8a20336c0d7c70bb970cf363e711057ce4d12d

Pith citing papers

Observation 09dadb60-9087-4483-9174-0060d7155eec · inbound

MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization cites this paper.

MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:54.714314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:54.714314Z digest=sha256:e91d9b11c72a7c148085eec884aa028692df6afe9496b7390aca8a95662d84a8

Observation a76ca4bf-c885-4701-ab4f-dcb157ef389b · inbound

Flow marching for a generative PDE foundation model cites this paper.

Flow marching for a generative PDE foundation model Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:51:25.457933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-18T13:48:14.532529Z digest=sha256:d8fd998eaca1ad14296ce8d6d87fe4b5e299ae19eced28da012600f6f97eb47d

Observation 2aa396cf-8da4-4bbc-b54a-8e497bd7139a · inbound

TC-AE: Unlocking Token Capacity for Deep Compression Autoencoders cites this paper.

TC-AE: Unlocking Token Capacity for Deep Compression Autoencoders Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:15:57.138473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:44:14.636654Z digest=sha256:00584474d44022efbc4cd05bce4285913d20ae82b113c660250b96ac0017db2c

Observation ac0c1624-78af-40ce-9719-35b212adf4b2 · inbound

Latent-Compressed Variational Autoencoder for Video Diffusion Models cites this paper.

Latent-Compressed Variational Autoencoder for Video Diffusion Models Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:03.077326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:23:19.583713Z digest=sha256:c7ce44358faa32033fbd712ce4823ec1e8a8d74a34821009adf0786c92a1e409

Observation b0ace6f3-d26d-4699-b281-320374b491fa · inbound

ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parameters cites this paper.

ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parameters Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:51:06.445578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T17:06:21.438738Z digest=sha256:be3d7eed101683941b779321319caf1fad4e48dc7d2d7183f9555757159b1904

Observation d0055b1f-9d88-459b-bc31-c64cb3afe703 · inbound

LiVeAction: a Lightweight, Versatile, and Asymmetric Neural Codec Design for Real-time Operation cites this paper.

LiVeAction: a Lightweight, Versatile, and Asymmetric Neural Codec Design for Real-time Operation Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:56:30.898606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T03:42:56.202960Z digest=sha256:131dc9b6ed8a6df541998fa3028b98c6cf89da1840967964a61e7f56d9075629

Observation 7e29255c-bf86-4f97-9053-82a549e05319 · inbound

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion cites this paper.

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:05:57.333354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T01:57:24.033068Z digest=sha256:634c64603eb1046eeee0f9b060a74ca383c476ce79678e8a7f0f3a0475cb31e9

Observation 3dbd517c-af80-47d7-b9be-5335072a6892 · inbound

Yeti: A compact protein structure tokenizer for reconstruction and multi-modal generation cites this paper.

Yeti: A compact protein structure tokenizer for reconstruction and multi-modal generation Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:29.699689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:01:05.492969Z digest=sha256:1400dd404479a4e02ce3524ddde696ef2a73fa5a1a1df88de3dcbade21d26979

Observation a0bbab4b-173f-4e72-9dca-99de3929f7dc · inbound

Vision Foundation Models as Generalist Tokenizers for Image Generation cites this paper.

Vision Foundation Models as Generalist Tokenizers for Image Generation Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:03:13.439682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T11:01:24.738195Z digest=sha256:3520666d591f7b2439af2922cf7727bbbaaca1066e7c16b38d6e0e7f5b190557

Observation 51798922-a3bf-4a4f-9b45-9dd38f20c13a · inbound

Balancing Image Compression and Generation with Bootstrapped Tokenization cites this paper.

Balancing Image Compression and Generation with Bootstrapped Tokenization Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:55.456464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T03:07:33.054518Z digest=sha256:cd4db5c1f5612e382ea03ad7863a0f8b9fb5c933d2bba24675caa3ef38161af0

Observation c02e9aa8-1cbf-473b-a748-2a9cb2180d83 · inbound

Multiplayer Interactive World Models with Representation Autoencoders cites this paper.

Multiplayer Interactive World Models with Representation Autoencoders Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 101

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T15:43:53.783116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-07T15:38:36.897705Z digest=sha256:48148f6942fa75a31f178fa8c8f4fcca836cf5af9c77bbef5c5363f479622d8f