Pith. sign in

REVIEW 1 cited by

Re-parameterizing VAEs for stability

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.13739 v1 pith:KORTYP52 submitted 2021-06-25 cs.LG cs.CV

classification cs.LGcs.CV
keywords vaesdistributionsapproachcomplexnormaloutputsourcestability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a theoretical approach towards the training numerical stability of Variational AutoEncoders (VAE). Our work is motivated by recent studies empowering VAEs to reach state of the art generative results on complex image datasets. These very deep VAE architectures, as well as VAEs using more complex output distributions, highlight a tendency to haphazardly produce high training gradients as well as NaN losses. The empirical fixes proposed to train them despite their limitations are neither fully theoretically grounded nor generally sufficient in practice. Building on this, we localize the source of the problem at the interface between the model's neural networks and their output probabilistic distributions. We explain a common source of instability stemming from an incautious formulation of the encoded Normal distribution's variance, and apply the same approach on other, less obvious sources. We show that by implementing small changes to the way we parameterize the Normal distributions on which they rely, VAEs can securely be trained.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On the Statistical Capacity of Deep Generative Models

    stat.ML 2025-01 conditional novelty 6.0 of 10

    Push-forwards of Gaussian or log-concave latent variables through Lipschitz neural networks are always sub-Gaussian or sub-exponential, so common deep generative models cannot generate heavy-tailed distributions.

Pith tools