Pith. sign in

Variational Diffusion Auto-encoder: Latent Space Extraction from Pre-trained Diffusion Models

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

As a widely recognized approach to deep generative modeling, Variational Auto-Encoders (VAEs) still face challenges with the quality of generated images, often presenting noticeable blurriness. This issue stems from the unrealistic assumption that approximates the conditional data distribution, $p(\textbf{x} | \textbf{z})$, as an isotropic Gaussian. In this paper, we propose a novel solution to address these issues. We illustrate how one can extract a latent space from a pre-existing diffusion model by optimizing an encoder to maximize the marginal data log-likelihood. Furthermore, we demonstrate that a decoder can be analytically derived post encoder-training, employing the Bayes rule for scores. This leads to a VAE-esque deep latent variable model, which discards the need for Gaussian assumptions on $p(\textbf{x} | \textbf{z})$ or the training of a separate decoder network. Our method, which capitalizes on the strengths of pre-trained diffusion models and equips them with latent spaces, results in a significant enhancement to the performance of VAEs.

fields

stat.ML 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

CoVAE: Consistency Training of Variational Autoencoders

stat.ML · 2025-07-12 · conditional · novelty 6.0

CoVAE trains a time-dependent VAE with a consistency loss so one or few decoder passes generate images, reaching FID 5.62 on MNIST and 11.69 on CIFAR-10 with adversarial loss, without a learned prior.

citing papers explorer

Showing 1 of 1 citing paper.

  • CoVAE: Consistency Training of Variational Autoencoders stat.ML · 2025-07-12 · conditional · none · ref 6 · internal anchor

    CoVAE trains a time-dependent VAE with a consistency loss so one or few decoder passes generate images, reaching FID 5.62 on MNIST and 11.69 on CIFAR-10 with adversarial loss, without a learned prior.