REVIEW 8 cited by
Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
To achieve the highest perceptual quality, state-of-the-art diffusion models are optimized with objectives that typically look very different from the maximum likelihood and the Evidence Lower Bound (ELBO) objectives. In this work, we reveal that diffusion model objectives are actually closely related to the ELBO. Specifically, we show that all commonly used diffusion model objectives equate to a weighted integral of ELBOs over different noise levels, where the weighting depends on the specific objective used. Under the condition of monotonic weighting, the connection is even closer: the diffusion objective then equals the ELBO, combined with simple data augmentation, namely Gaussian noise perturbation. We show that this condition holds for a number of state-of-the-art diffusion models. In experiments, we explore new monotonic weightings and demonstrate their effectiveness, achieving state-of-the-art FID scores on the high-resolution ImageNet benchmark.
Forward citations
Cited by 8 Pith papers
-
Flow Matching Policy Gradients
FPO trains flow-based policies with PPO by replacing the likelihood ratio with an exponentiated flow matching loss difference.
-
Unifying Generative Models with Path Integrals
A one-loop correction, computed from two auxiliary ODEs, brings deterministic generative samplers close to the stochastic reference (53% error reduced to 1.6% on a cubic drift), within a path-integral framework that u...
-
WaiT for the Signal: Simple Frequency-Aware Flow-Matching
WaiT delays high-frequency wavelet bands in flow-matching image generation until coarse structure emerges, improving quality and cutting compute, with a reported SOTA FID of 1.30 on ImageNet 512.
-
Bridging Diffusion Pruning and Step Distillation with Teacher-Aligned Repair
A short teacher-alignment repair stage between structured pruning and one-step distillation yields a 20% pruned one-step generator that improves FID from 3.53 to 3.12 on ImageNet-512 while reducing NFE from 63 to 1.
-
ELT: Elastic Looped Transformers for Visual Generation
Weight-shared looped transformers trained with intra-loop self-distillation match MaskGIT-class FID/FVD at roughly 4x fewer parameters and support any-time inference across loop counts.
-
Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments
Flow equivariant world models use a latent memory that shifts with the agent and with inferred object motion, giving stable long-horizon prediction under partial observability.
-
INRFlow: Flow Matching for INRs in Ambient Space
A domain-agnostic PerceiverIO-style transformer, trained with a point-wise flow-matching loss in ambient space, generates competitive images, point clouds, and protein structures without a separate data compressor.
-
Conditional Diffusion Models are Medical Image Classifiers that Provide Explainability and Uncertainty for Free
Conditional diffusion models can classify medical images by comparing reconstruction errors, and the per-noise-level majority vote also yields explanation and uncertainty byproducts.
Discussion (0). Continue with ORCID to comment.