Pith. sign in

REVIEW 33 cited by

Understanding Diffusion Models: A Unified Perspective

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.11970 v1 pith:WK42L2EZ submitted 2022-08-25 cs.LG cs.CV

Understanding Diffusion Models: A Unified Perspective

classification cs.LG cs.CV
keywords modelsdiffusionvariationalinputperspectivearbitraryfunctiongenerative
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Diffusion models have shown incredible capabilities as generative models; indeed, they power the current state-of-the-art models on text-conditioned image generation such as Imagen and DALL-E 2. In this work we review, demystify, and unify the understanding of diffusion models across both variational and score-based perspectives. We first derive Variational Diffusion Models (VDM) as a special case of a Markovian Hierarchical Variational Autoencoder, where three key assumptions enable tractable computation and scalable optimization of the ELBO. We then prove that optimizing a VDM boils down to learning a neural network to predict one of three potential objectives: the original source input from any arbitrary noisification of it, the original source noise from any arbitrarily noisified input, or the score function of a noisified input at any arbitrary noise level. We then dive deeper into what it means to learn the score function, and connect the variational perspective of a diffusion model explicitly with the Score-based Generative Modeling perspective through Tweedie's Formula. Lastly, we cover how to learn a conditional distribution using diffusion models via guidance.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 33 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Transformers Learn the Optimal DDPM Denoiser for Multi-Token GMMs

    cs.LG 2026-04 unverdicted novelty 8.0

    Transformers converge globally to the optimal DDPM denoiser for multi-token GMMs via self-attention mean denoising, with explicit token and iteration requirements.

  2. Bayesian Experimental Design via Score Matching

    stat.ML 2026-07 conditional novelty 7.0

    SCOREBED isolates EIG double intractability in a policy-independent score-matching stage, then trains design policies with a singly intractable gradient estimator, enabling cheap multi-policy selection.

  3. Pathway variability, coat stiffening and mechanical adaptation during clathrin-mediated endocytosis

    q-bio.SC 2026-06 unverdicted novelty 7.0

    Hybrid simulation and non-Euclidean elasticity theory demonstrate that clathrin coats develop adaptive rigidity and memory during growth, producing flat, stalled, or closed outcomes through two energy-landscape gates ...

  4. Physics-Informed Neural PDE Solvers via Spatio-Temporal MeanFlow

    cs.LG 2026-05 unverdicted novelty 7.0

    Spatio-Temporal MeanFlow adapts MeanFlow to PDEs by replacing the generative velocity field with the physical operator and extending the integral constraint to the spatio-temporal domain, yielding a unified solver for...

  5. Causal inference for social network formation

    econ.EM 2026-04 conditional novelty 7.0

    Random team assignments in a professional firm reveal that indirect ties strongly increase new direct tie formation, while effects of degree and local density are smaller and less robust.

  6. Score-based Membership Inference on Diffusion Models

    cs.LG 2025-09 unverdicted novelty 7.0

    Presents SimA, a score-based single-query membership inference attack for diffusion models and LDMs that uses denoiser output norm to reveal training set proximity and outperforms multi-query baselines on eight datasets.

  7. Doloris: Dual Conditional Diffusion Implicit Bridges with Sparsity Masking Strategy for Unpaired Single-Cell Perturbation Estimation

    cs.LG 2025-06 unverdicted novelty 7.0

    Doloris introduces dual conditional diffusion implicit bridges plus a sparsity masking strategy to model unpaired single-cell perturbation responses and reports state-of-the-art results on public datasets.

  8. To Erase, or Not to Erase: Robust Training-Free Concept Erasure with Preservation aware Adaptive Ranked Subspace Expansion

    cs.CV 2026-07 conditional novelty 6.5

    Diffusion-grounded erase/retain retrieval plus retain-orthogonal value projection and trigger-guided subspace expansion erases concepts more robustly than prior CETs while keeping FID/CLIP near the unedited model.

  9. Tightening the Score Matching Gap for Diffusion Models

    stat.ML 2026-07 conditional novelty 6.0

    Tighter score-matching gap bounds for diffusion models via entropy flows, LSI and reflection couplings show that low-noise score accuracy dominates sample quality metrics.

  10. CDCP: Conditional Diffusion Model with Contextual Prompts for Multi-task Offline Safe Reinforcement Learning

    cs.LG 2026-07 conditional novelty 6.0

    A conditional diffusion model with contextual prompts and classifier-free cost guidance learns a shared safe multi-task policy from offline data and meets varying cost limits without retraining.

  11. A Two-Step Ensemble Score Filter for Data Assimilation in Partially Observed Systems

    physics.ao-ph 2026-06 unverdicted novelty 6.0

    EnSF-LR combines nonlinear score-based analysis on observed components with EnKF-style linear regression on unobserved components via ensemble covariance, achieving lower full-state RMSE than EnSF and EnKF in nonlinea...

  12. Latent Block-Diffusion Temporal Point Processes: A Semi-Autoregressive Framework for Asynchronous Event Sequence Generation

    cs.LG 2026-06 unverdicted novelty 6.0

    LBDTPP generates high-quality variable-length event sequences by autoregressing over latent blocks and diffusing within blocks, with Wasserstein bounds claiming reduced error accumulation under local approximation and...

  13. T-CLIP: Enabling Thermal Perception for Contrastive Language-Image Pretraining

    cs.CV 2026-05 unverdicted novelty 6.0

    T-CLIP introduces a physics-aware thermal captioning dataset (IR-Cap) and a decoupled dual-LoRA adaptation of CLIP that improves cross-modal retrieval on thermal benchmarks by separating scene-level and object-level t...

  14. Mixture-of-Experts Diffusion Models for Adaptive Massive MIMO Channel Estimation via Variational Bayesian Inference

    eess.SP 2026-05 unverdicted novelty 6.0

    A mixture-of-experts diffusion model with variational Bayesian inference jointly infers the channel and expert indicator to adapt to different propagation environments in massive MIMO channel estimation.

  15. GenTS: A Comprehensive Benchmark Library for Generative Time Series Models

    cs.LG 2026-05 unverdicted novelty 6.0

    GenTS is a modular benchmark library providing unified data pipelines, generative models, and evaluation metrics for time series synthesis, forecasting, and imputation, with open-source code and initial benchmarking e...

  16. Inference-Time Attribute Distribution Alignment for Unconditional Diffusion

    cs.LG 2026-05 unverdicted novelty 6.0

    An optimal control formulation adds time-dependent perturbations to the reverse diffusion process to match target attribute distributions while preserving sample fidelity.

  17. On the Robustness of Distribution Support under Diffusion Guidance

    cs.LG 2026-05 unverdicted novelty 6.0

    Guided diffusion generates samples near the target distribution support under exact score access, explaining its empirical success in producing plausible outputs.

  18. Brownian Bridge Diffusion for Sequential Recommendation

    cs.IR 2025-07 unverdicted novelty 6.0

    BBDRec applies Brownian bridge diffusion to enable direct item-to-history transitions in sequential recommendation, outperforming prior diffusion and sequential baselines on public datasets.

  19. DMin: Scalable Training Data Influence Estimation for Diffusion Models

    cs.CV 2024-12 unverdicted novelty 6.0

    DMin uses gradient compression to scalably estimate training data influence in billion-parameter diffusion models.

  20. Improved DDIM Sampling with Moment Matching Gaussian Mixtures

    cs.CV 2023-11 unverdicted novelty 6.0

    Moment-matched GMM kernels in DDIM yield lower FID and higher IS than Gaussian kernels at small sampling steps on CelebA-HQ, FFHQ, ImageNet, and Stable Diffusion tasks.

  21. Rethinking Forgery Attacks on Semantic Watermarks in Black-Box Settings: A Geometric Distortion Perspective

    cs.CR 2026-06 unverdicted novelty 5.0

    Semantic watermarks in LDMs have an irreducible geometric distortion floor from proxy-target model mismatches that limits forgery fidelity and supports scheme-agnostic detection via global drift and local deformation.

  22. On the Redundancy of Timestep Embeddings in Diffusion Models

    cs.LG 2026-06 conditional novelty 5.0

    Under high-dimensional concentration conditions, the diffusion denoising objective admits the same global minimizer without timestep embeddings, and time-agnostic U-Nets/DiTs empirically match or improve FID on CelebA...

  23. Self-Regulating Annealing in Heavy-Tailed Diffusion Models

    stat.ML 2026-06 unverdicted novelty 5.0

    Proposes SDE sampler with state-dependent diffusion for HTDMs that induces self-regulating annealing, claimed necessary for heavy-tailed sampling.

  24. Physical Foundation Models: Fixed hardware implementations of large-scale neural networks

    cs.LG 2026-04 unverdicted novelty 5.0

    Physical Foundation Models are fixed physical hardware realizations of foundation-scale neural networks that compute via inherent material dynamics, potentially delivering orders-of-magnitude gains in energy efficienc...

  25. Rethinking the Diffusion Model from a Langevin Perspective

    cs.LG 2026-04 unverdicted novelty 5.0

    Diffusion models are reorganized under a Langevin perspective that unifies ODE and SDE formulations and shows flow matching is equivalent to denoising under maximum likelihood.

  26. Downscaling weather forecasts from Low- to High-Resolution with Diffusion Models

    physics.ao-ph 2026-03 unverdicted novelty 5.0

    A conditional diffusion model downscales global atmospheric forecasts from 100 km to 30 km resolution while improving probabilistic skill, matching power spectra, and preserving physical relationships.

  27. Preserve and Personalize: Personalized Text-to-Image Diffusion Models without Distributional Drift

    cs.CV 2025-05 unverdicted novelty 5.0

    Proposes Lipschitz regularization during fine-tuning to prevent distributional drift in personalized diffusion models, improving subject fidelity and prompt adherence.

  28. A Probabilistic Formulation of Offset Noise in Diffusion Models

    stat.ML 2024-12 unverdicted novelty 5.0

    A diffusion model variant that adds structured non-zero-mean noise via modified forward/reverse processes, yielding an ELBO loss analogous to offset noise but with time-dependent coefficients, and showing gains on syn...

  29. Balancing User Preferences by Social Networks: A Condition-Guided Social Recommendation Model for Mitigating Popularity Bias

    cs.SI 2024-05 unverdicted novelty 5.0

    CGSoRec denoises social relations and reweights user social preferences to serve as conditions that steer a diffusion recommender away from popularity bias.

  30. On the Redundancy of Timestep Embeddings in Diffusion Models

    cs.LG 2026-06 unverdicted novelty 4.0

    Timestep embeddings are redundant in diffusion models under certain conditions, with time-agnostic variants matching or exceeding conditioned models on FID, precision, and recall for CelebA and CIFAR-10.

  31. Elucidating Representation Degradation Problem in Diffusion Model Training

    cs.LG 2026-05 unverdicted novelty 4.0

    Diffusion models suffer representation degradation at high noise due to recoverability mismatch; ERD mitigates this by dynamic optimization reallocation, accelerating convergence across backbones.

  32. On the Robustness of Distribution Support under Diffusion Guidance

    cs.LG 2026-05 unverdicted novelty 4.0

    Establishes robustness of distribution support for guided diffusion processes under exact score access across DDIM, DDPM, and exponential integrator discretizations.

  33. AI-Generated Image Recognition via Fusion of CNNs and Vision Transformers

    cs.CV 2026-06 unverdicted novelty 2.0

    A fused CNN-ViT model achieves 97.32% accuracy distinguishing AI-generated from real images on the CIFAKE dataset.