Pith. sign in

REVIEW 5 cited by

Very Deep VAEs Generalize Autoregressive Models and Can Outperform Them on Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.10650 v2 pith:WGMS5PB4 submitted 2020-11-20 cs.LG cs.CV

classification cs.LGcs.CV
keywords modelsvaesautoregressivedeepdepthfasterhierarchicalimages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a hierarchical VAE that, for the first time, generates samples quickly while outperforming the PixelCNN in log-likelihood on all natural image benchmarks. We begin by observing that, in theory, VAEs can actually represent autoregressive models, as well as faster, better models if they exist, when made sufficiently deep. Despite this, autoregressive models have historically outperformed VAEs in log-likelihood. We test if insufficient depth explains why by scaling a VAE to greater stochastic depth than previously explored and evaluating it CIFAR-10, ImageNet, and FFHQ. In comparison to the PixelCNN, these very deep VAEs achieve higher likelihoods, use fewer parameters, generate samples thousands of times faster, and are more easily applied to high-resolution images. Qualitative studies suggest this is because the VAE learns efficient hierarchical visual representations. We release our source code and models at https://github.com/openai/vdvae.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tractable Representation Learning with Probabilistic Circuits

    cs.LG 2025-07 conditional novelty 7.0 of 10

    Autoencoding probabilistic circuits train a single probabilistic circuit to jointly model data and explicit embedding variables, enabling end-to-end autoencoding with neural decoders and robust encoding under missing data.

  2. SoK: Can Synthetic Images Replace Real Data? A Survey of Utility and Privacy of Synthetic Image Generation

    cs.CR 2025-06 conditional novelty 7.0 of 10

    A systematic survey and benchmark showing that diffusion-based synthetic data can achieve better utility-privacy tradeoffs than DP-SGD on real data for some image classifiers, with the best release strategy depending ...

  3. Unpaired Joint Distribution Modeling via Multi-Scale Image Representations

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A latent graphical model with multi-scale auxiliary maps bounds unpaired joint-distribution error and synthesizes realistic noise pairs that improve real-world and cryo-EM denoising.

  4. Probabilistic cosmological inference on HI tomographic data

    astro-ph.IM 2025-07 conditional novelty 6.0 of 10

    A 3D CNN encoder plus a masked autoregressive flow recovers Ωm and σ8 from simulated HI tomographic data cubes at z=1 with R² ≥ 0.91 on test sets, with reduced but still useful accuracy out of distribution.

  5. Diffusion Counterfactual Generation with Semantic Abduction

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Diffusion-based causal image counterfactuals with semantic abduction improve identity preservation at a small cost in intervention effectiveness, demonstrated on Morpho-MNIST, CelebA-HQ, and mammogram artifact removal.

Pith tools