REVIEW 7 cited by
Very Deep VAEs Generalize Autoregressive Models and Can Outperform Them on Images
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present a hierarchical VAE that, for the first time, generates samples quickly while outperforming the PixelCNN in log-likelihood on all natural image benchmarks. We begin by observing that, in theory, VAEs can actually represent autoregressive models, as well as faster, better models if they exist, when made sufficiently deep. Despite this, autoregressive models have historically outperformed VAEs in log-likelihood. We test if insufficient depth explains why by scaling a VAE to greater stochastic depth than previously explored and evaluating it CIFAR-10, ImageNet, and FFHQ. In comparison to the PixelCNN, these very deep VAEs achieve higher likelihoods, use fewer parameters, generate samples thousands of times faster, and are more easily applied to high-resolution images. Qualitative studies suggest this is because the VAE learns efficient hierarchical visual representations. We release our source code and models at https://github.com/openai/vdvae.
Forward citations
Cited by 7 Pith papers
-
Tractable Representation Learning with Probabilistic Circuits
Autoencoding probabilistic circuits train a single probabilistic circuit to jointly model data and explicit embedding variables, enabling end-to-end autoencoding with neural decoders and robust encoding under missing data.
-
SoK: Can Synthetic Images Replace Real Data? A Survey of Utility and Privacy of Synthetic Image Generation
A systematic survey and benchmark showing that diffusion-based synthetic data can achieve better utility-privacy tradeoffs than DP-SGD on real data for some image classifiers, with the best release strategy depending ...
-
Unpaired Joint Distribution Modeling via Multi-Scale Image Representations
A latent graphical model with multi-scale auxiliary maps bounds unpaired joint-distribution error and synthesizes realistic noise pairs that improve real-world and cryo-EM denoising.
-
Probabilistic cosmological inference on HI tomographic data
A 3D CNN encoder plus a masked autoregressive flow recovers Ωm and σ8 from simulated HI tomographic data cubes at z=1 with R² ≥ 0.91 on test sets, with reduced but still useful accuracy out of distribution.
-
Diffusion Counterfactual Generation with Semantic Abduction
Diffusion-based causal image counterfactuals with semantic abduction improve identity preservation at a small cost in intervention effectiveness, demonstrated on Morpho-MNIST, CelebA-HQ, and mammogram artifact removal.
-
Improving Brain-to-Image Reconstruction via Fine-Grained Text Bridging
Fine-grained text decoded from fMRI with three reward signals improves brain-to-image reconstruction when fused into existing diffusion pipelines.
-
TaxaDiffusion: Progressively Trained Diffusion Model for Fine-Grained Species Generation
A diffusion model trained progressively from Kingdom to Species generates more accurate fine-grained animal images, including rare species with as few as one training sample.
Discussion (0). Sign in to comment.