Pith. sign in

Paper Citation Record · LEDGER

Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2406.11837.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.11837 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:48:03.489771Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:08:37.450322Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 36306d8e-f675-456d-91f7-11a1e0cecabb · inbound

Deep Learning-Based Image Compression for Wireless Communications: Impacts on Reliability,Throughput, and Latency cites this paper.

Deep Learning-Based Image Compression for Wireless Communications: Impacts on Reliability,Throughput, and Latency Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T19:33:24.947232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:33:24.947232Z digest=sha256:7eb15a5937a33db178ca1e0761bbd0c486372b3e1e9269855aefafe24c4d2d8d

Observation 93290248-c703-4097-aa4b-77e74cd4525f · inbound

Factorized Visual Tokenization and Generation cites this paper.

Factorized Visual Tokenization and Generation Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:03.179680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:53:03.179680Z digest=sha256:48a7d984c0d4fd5af8d28ad45a5851a62f6eb1313c5322f3577384b7743cdd1a

Observation ec26adc0-ca0f-4719-a6af-816c9277ff18 · inbound

Scalable Image Tokenization with Index Backpropagation Quantization cites this paper.

Scalable Image Tokenization with Index Backpropagation Quantization Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:46.631242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:46.631242Z digest=sha256:a5df04c0c129abdb8d791ca71dc086c9e0a90a6191be8da5fa2bd473c207b729

Observation 52a578d8-715e-4592-bc2e-8762c087d474 · inbound

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation cites this paper.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 76

Resolution
malformed identifier
no resolver link, observed 2026-08-11T22:54:20.710373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.710373Z digest=sha256:dfd1ec4501e90cb0766cdc358157ce29684a2f0d43d8d7ee5133e14b8e9b1a27

Observation cc3f6384-aede-40b3-bbc3-037859db19ab · inbound

SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer cites this paper.

SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 124

Resolution
unresolved
no resolver link, observed 2026-08-11T15:33:13.683119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:33:13.683119Z digest=sha256:4f612dff9c0101a436184cd646d1ccfd8ec075a4e676717b1186d4477d7c0939

Observation 582e0f43-eef2-45f7-99cc-7fc8d2641c8c · inbound

Preventing Local Pitfalls in Vector Quantization via Optimal Transport cites this paper.

Preventing Local Pitfalls in Vector Quantization via Optimal Transport Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:29.866761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:29.866761Z digest=sha256:e67ba645750c95f066ddcf66f6278df3495dbf3fbe5199b7f3fcd4ba5f2ee950

Observation ea50a6f4-551d-4743-8670-bc21ae9a44b0 · inbound

Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models cites this paper.

Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:14.574576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:14.574576Z digest=sha256:e4141a932d1f826d9cf2aa95035831bf432cd1db3b922e34bae0e7259bf586d3

Observation ffad08d1-d959-47c7-8ac0-10421ed12af9 · inbound

InterLCM: Low-Quality Images as Intermediate States of Latent Consistency Models for Effective Blind Face Restoration cites this paper.

InterLCM: Low-Quality Images as Intermediate States of Latent Consistency Models for Effective Blind Face Restoration Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T13:02:37.551540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:02:37.551540Z digest=sha256:7e4182e46d15eb26ddbdb2a9a151dc1eff024bdd44bc1761209df2fce8004f18

Observation 6d3811b7-d03b-4fa6-9681-480d2b127e50 · inbound

Masked Autoencoders Are Effective Tokenizers for Diffusion Models cites this paper.

Masked Autoencoders Are Effective Tokenizers for Diffusion Models Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T04:47:30.466469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:47:30.466469Z digest=sha256:62d41a58bb7a1ac48ecedb1499fada84979107e85342f3aaa893d5481eb3c890

Observation 336f5d11-e67d-4f1b-b457-b757e135dace · inbound

Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens cites this paper.

Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-16T11:48:03.489771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:48:03.489771Z digest=sha256:4c40defe55467ca0d8891c3df715d8169468eb4802b725b6247d5c36ac01c4ae

Observation 69267f92-7675-4126-b76f-2b9aa29e46d9 · inbound

Fast Autoregressive Models for Continuous Latent Generation cites this paper.

Fast Autoregressive Models for Continuous Latent Generation Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:52.748177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:38:52.748177Z digest=sha256:03bedfd484bd9ae689055d9c25e364136a167cef4fb745ccb7be5b7d3310522e

Observation dee18e11-8a7f-4bf7-bd7a-55144e5b916e · inbound

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information cites this paper.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:18.403349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:18.403349Z digest=sha256:f431aa6e21a5768c35ac76f0c50e82168094f956277765ba071e23d08269be91

Observation b0468529-a537-472a-aa38-067c6e842efa · inbound

MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation cites this paper.

MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:30.315990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:30.315990Z digest=sha256:dac95e4d31467f912e0bbb03c55f436f659dfe7225d51a70c84e79ed4b8a4075

Observation 6459ee8c-5f34-43d2-9b3e-5fba6bd2d812 · inbound

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation cites this paper.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.941622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.941622Z digest=sha256:5fd9a70c55044fbe5fec64fddc1e5768141e4d3ddcea3bb6fc9c0b5227e764a4

Observation 9739bb18-d759-4093-91f0-3834a5dc1c86 · inbound

Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis cites this paper.

Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T20:49:43.861863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:49:43.861863Z digest=sha256:3c17c58b406e1e59ed82ae5ffa25a4f91d69fd1bbd3e490e22f19966e49e3ab2

Observation 3c5cddd5-da6d-4fec-a0e4-672898b727ba · inbound

Hita: Holistic Tokenizer for Autoregressive Image Generation cites this paper.

Hita: Holistic Tokenizer for Autoregressive Image Generation Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 63

Resolution
malformed identifier
no resolver link, observed 2026-08-06T20:38:59.016813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:38:59.016813Z digest=sha256:fcf2779768e250f74e8d34610b7296e78e17135c7ba31c703c63f8863558017a

Observation 750424a1-162d-4ad8-82f6-b5b90669c3c6 · inbound

Quantize-then-Rectify: Efficient VQ-VAE Training cites this paper.

Quantize-then-Rectify: Efficient VQ-VAE Training Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T17:34:57.408009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:34:57.408009Z digest=sha256:1cdffa15126e995073202615fb793d58e9797dd9284812a417e1856d22c77b47

Observation 8c1944f5-a9c2-48ea-8d90-fe0410a0f41e · inbound

Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization cites this paper.

Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T18:12:48.383882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:12:48.383882Z digest=sha256:d6d3432def474b7ad54e4ce8df1c454bad85f355651bdcca1c21303b81b0542f

Observation 4a92f15a-bb73-4e1a-a97a-1392246486f9 · inbound

Mutual Enhancement Between Global Tokens and Patch Tokens: From Theory to Practice cites this paper.

Mutual Enhancement Between Global Tokens and Patch Tokens: From Theory to Practice Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:43:51.034253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-20T22:41:44.510546Z digest=sha256:9e9b080707ae0de11d0d2eee55170eb0f7072fb20c0b0b10df3dcb9bf820ad51

Observation f2e809a6-712a-4268-aa3a-68b685ee27ce · inbound

Vision Foundation Models as Generalist Tokenizers for Image Generation cites this paper.

Vision Foundation Models as Generalist Tokenizers for Image Generation Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 103

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:03:13.425796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T11:01:24.738195Z digest=sha256:f618ec32ad321181316ffb09bea103bd135674f290e30d26a8b7f2045bce699c

Observation 298bf8bd-f40b-41b5-965e-dc8a8efac37b · inbound

HiTokSR: A Coarse-to-Fine Tokenizer with Hierarchical Codebooks for High-Fidelity Real-World Image Super-Resolution cites this paper.

HiTokSR: A Coarse-to-Fine Tokenizer with Hierarchical Codebooks for High-Fidelity Real-World Image Super-Resolution Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:06:13.489175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T17:32:15.239191Z digest=sha256:a0a31225b9e3402fc405e2d2bb447edc80c888fcff9e8cf51808ff78a33be8c1

Observation 80afaa1a-edb8-48a5-bfd4-a95221c22ac1 · inbound

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder cites this paper.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:37:40.265134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:1e0014c7e3f043d02fd191c885dc721bd4fcdf7c4a061cbdc530563e922cb120

Observation 526adc27-5157-4f46-b302-46da018cc692 · inbound

NSVQ: Mitigating Codebook Collapse by Stabilizing Encoder Drift in Vector Quantization cites this paper.

NSVQ: Mitigating Codebook Collapse by Stabilizing Encoder Drift in Vector Quantization Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:27:39.717943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T13:18:47.472178Z digest=sha256:f297b506c8978c6f8b504eff5cf0f9df745833e9645f08326fa9936668da6ab5

Observation 50a5e554-ca50-4fcc-b637-0521da0cabc8 · inbound

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment cites this paper.

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:08:37.452164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T06:05:26.735340Z digest=sha256:1b38fa007301a9073e91ebddd5b7a689c7bb8f8c88b645de84f9da5c548ca4e4