Pith. sign in

REVIEW 2 cited by

Optimizing Hierarchical Image VAEs for Sample Quality

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.10205 v1 pith:AN53EOIC submitted 2022-10-18 cs.LG cs.CV

classification cs.LGcs.CV
keywords imagehierarchicalvaesintroducestrategyachievedadditionallyaddress
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

While hierarchical variational autoencoders (VAEs) have achieved great density estimation on image modeling tasks, samples from their prior tend to look less convincing than models with similar log-likelihood. We attribute this to learned representations that over-emphasize compressing imperceptible details of the image. To address this, we introduce a KL-reweighting strategy to control the amount of infor mation in each latent group, and employ a Gaussian output layer to reduce sharpness in the learning objective. To trade off image diversity for fidelity, we additionally introduce a classifier-free guidance strategy for hierarchical VAEs. We demonstrate the effectiveness of these techniques in our experiments. Code is available at https://github.com/tcl9876/visual-vae.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Nested Diffusion Models Using Hierarchical Latent Priors

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Nested diffusion models that generate images by progressively synthesizing hierarchical semantic latents from a frozen pretrained encoder improve image quality over single-level baselines at modest extra cost.

  2. Scaling Image Tokenizers with Grouped Spherical Quantization

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GSQ combines spherical codebook initialization, normalized lookup, and group-wise latent decomposition to achieve strong reconstruction quality at 16x spatial downsampling in far fewer training steps than prior tokenizers.

Pith tools