Pith. sign in

REVIEW 5 cited by

SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic Quantization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.07547 v2 pith:26RQW4MN submitted 2022-05-16 cs.LG cs.CV

classification cs.LGcs.CV
keywords sq-vaequantizationcodebookstochastictrainingvariationalvq-vaeautoencoder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

One noted issue of vector-quantized variational autoencoder (VQ-VAE) is that the learned discrete representation uses only a fraction of the full capacity of the codebook, also known as codebook collapse. We hypothesize that the training scheme of VQ-VAE, which involves some carefully designed heuristics, underlies this issue. In this paper, we propose a new training scheme that extends the standard VAE via novel stochastic dequantization and quantization, called stochastically quantized variational autoencoder (SQ-VAE). In SQ-VAE, we observe a trend that the quantization is stochastic at the initial stage of the training but gradually converges toward a deterministic quantization, which we call self-annealing. Our experiments show that SQ-VAE improves codebook utilization without using common heuristics. Furthermore, we empirically show that SQ-VAE is superior to VAE and VQ-VAE in vision- and speech-related tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    1D binary image latents reduce a 1024x1024 image to 128 discrete tokens and support text-to-image generation with diffusion and autoregressive models.

  2. Discrete JEPA: Learning Discrete Token Representations without Reconstruction

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Discrete-JEPA learns discrete semantic image tokens through latent predictive coding without pixel reconstruction, and achieves stable long-horizon prediction on synthetic symbolic tasks.

  3. A look at adversarial attacks on radio waveforms from discrete latent space

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A VQVAE's discrete reconstruction substantially reduces the effectiveness of adversarial attacks on high-SNR radio waveform classifiers.

  4. Cross-Layer Discrete Concept Discovery for Interpreting Language Models

    cs.LG 2025-06 reject novelty 5.0 of 10

    CLVQ-VAE maps lower-layer transformer activations to higher-layer ones through a discrete codebook, yielding concept vectors evaluated with probe ablation and human annotation.

  5. MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Splitting quantization across multiple small sub-codebooks with nested masking raises VQ-VAE reconstruction fidelity, giving MGVQ rFID 0.49 and PSNR 24.70 on ImageNet at 16 times downsampling.

Pith tools