Pith. sign in

REVIEW 51 cited by

Understanding Diffusion Models: A Unified Perspective

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.11970 v1 pith:WK42L2EZ submitted 2022-08-25 cs.LG cs.CV

classification cs.LGcs.CV
keywords modelsdiffusionvariationalinputperspectivearbitraryfunctiongenerative
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Diffusion models have shown incredible capabilities as generative models; indeed, they power the current state-of-the-art models on text-conditioned image generation such as Imagen and DALL-E 2. In this work we review, demystify, and unify the understanding of diffusion models across both variational and score-based perspectives. We first derive Variational Diffusion Models (VDM) as a special case of a Markovian Hierarchical Variational Autoencoder, where three key assumptions enable tractable computation and scalable optimization of the ELBO. We then prove that optimizing a VDM boils down to learning a neural network to predict one of three potential objectives: the original source input from any arbitrary noisification of it, the original source noise from any arbitrarily noisified input, or the score function of a noisified input at any arbitrary noise level. We then dive deeper into what it means to learn the score function, and connect the variational perspective of a diffusion model explicitly with the Score-based Generative Modeling perspective through Tweedie's Formula. Lastly, we cover how to learn a conditional distribution using diffusion models via guidance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 51 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 112 citations worldwide. Full citation record

  1. Personalized Federated Learning via Variance-Aware Nonparametric Empirical Bayes

    stat.ML 2026-08 conditional novelty 7.0 of 10

    VANEB generalizes nonparametric empirical Bayes to parameter-dependent noise and uses it to personalize federated models by shrinking local estimates toward a learned population prior.

  2. Bayesian Experimental Design via Score Matching

    stat.ML 2026-07 conditional novelty 7.0 of 10

    SCOREBED isolates EIG double intractability in a policy-independent score-matching stage, then trains design policies with a singly intractable gradient estimator, enabling cheap multi-policy selection.

  3. To Erase, or Not to Erase: Robust Training-Free Concept Erasure with Preservation aware Adaptive Ranked Subspace Expansion

    cs.CV 2026-07 conditional novelty 6.5 of 10

    Diffusion-grounded erase/retain retrieval plus retain-orthogonal value projection and trigger-guided subspace expansion erases concepts more robustly than prior CETs while keeping FID/CLIP near the unedited model.

  4. Tightening the Score Matching Gap for Diffusion Models

    stat.ML 2026-07 conditional novelty 6.0 of 10

    Tighter score-matching gap bounds for diffusion models via entropy flows, LSI and reflection couplings show that low-noise score accuracy dominates sample quality metrics.

  5. CDCP: Conditional Diffusion Model with Contextual Prompts for Multi-task Offline Safe Reinforcement Learning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A conditional diffusion model with contextual prompts and classifier-free cost guidance learns a shared safe multi-task policy from offline data and meets varying cost limits without retraining.

  6. TRE: Training-Free Hallucination Detection for Diffusion Language Models

    cs.AI 2026-06 conditional novelty 6.0 of 10

    TRE weights the entropy of newly revealed tokens by denoising step and detects hallucinations in diffusion LLMs with no training and a single generation.

  7. REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion

    cs.CV 2025-12 conditional novelty 6.0 of 10

    Nonlinear multi-layer compression of frozen VFM patch semantics, jointly denoised with VAE latents, improves ImageNet 256x256 FID (12.9 vs 15.2 for REG at SiT-B/2, 400K) and accelerates convergence over REPA/ReDi/REG.

  8. Joint Model-based Model-free Diffusion for Planning with Constraints

    cs.RO 2025-09 conditional novelty 6.0 of 10

    JM2D samples diffusion plans and safety-filter corrections jointly using a single importance-sampling-guided diffusion process, improving task success and reducing safety-filter interventions.

  9. Dual Orthogonal Guidance for Robust Diffusion-based Handwritten Text Generation

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    A new guidance method for diffusion-based text generation subtracts the orthogonal component of a negative prompt from a positive prompt to reduce artifacts and increase style variation.

  10. Clustering via Self-Supervised Diffusion

    cs.AI 2025-07 conditional novelty 6.0 of 10

    CLUDI trains a student to imitate stochastic diffusion-generated cluster assignments on pre-trained DINO image features and averages multiple assignments to cluster images.

  11. Hyper-modal Imputation Diffusion Embedding with Dual-Distillation for Federated Multimodal Knowledge Graph Completion

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A masked diffusion imputer plus dual distillation (MMFeD3-HidE) improves link prediction on a new federated multimodal knowledge graph benchmark with 50% missing visual/textual modalities.

  12. Exploiting the Exact Denoising Posterior Score in Training-Free Guidance of Diffusion Models

    stat.ML 2025-06 conditional novelty 6.0 of 10

    An exact denoising posterior score is derived and used to compute time-dependent DPS step sizes that transfer to colorization, inpainting, and super-resolution.

  13. Fusion of multi-source precipitation records via coordinate-based generative model

    physics.ao-ph 2025-06 conditional novelty 6.0 of 10

    A coordinate-based diffusion model fuses multi-source precipitation records and corrects biases in unseen operational forecasts.

  14. Diffusion Counterfactual Generation with Semantic Abduction

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Diffusion-based causal image counterfactuals with semantic abduction improve identity preservation at a small cost in intervention effectiveness, demonstrated on Morpho-MNIST, CelebA-HQ, and mammogram artifact removal.

  15. Sparse Autoencoders, Again?

    cs.LG 2025-06 conditional novelty 6.0 of 10

    VAEase gates the VAE decoder input by the encoder's variance, combining sparse-autoencoder adaptive sparsity with a hyperparameter-free loss; a global-minimizer theorem says active latent dimensions recover per-manifo...

  16. Autoregressive regularized score-based diffusion models for multi-scenarios fluid flow prediction

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A regularized autoregressive score-based diffusion model predicts turbulent flows across multiple scenarios, with the variance-preserving SDE formulation performing best.

  17. Unsupervised Learning for Class Distribution Mismatch

    cs.CV 2025-05 conditional novelty 6.0 of 10

    UCDM builds positive and negative image pairs with a text-to-image diffusion model and trains an open-set classifier with no instance labels, outperforming semi-supervised baselines on three benchmarks.

  18. InstaRevive: One-Step Image Enhancement via Dynamic Score Matching

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A one-step image enhancer that combines dynamically controlled score-based distillation with caption prompts, matching multi-step diffusion quality on face and super-resolution benchmarks.

  19. CDM: Contact Diffusion Model for Multi-Contact Point Localization

    cs.RO 2025-02 conditional novelty 6.0 of 10

    A Contact Diffusion Model conditioned on force/torque readings and a signed distance field localizes single and dual robot-arm contacts with 0.44 cm and 1.24 cm real-world error.

  20. Learning to Learn Weight Generation via Local Consistency Diffusion

    cs.LG 2025-02 reject novelty 6.0 of 10

    Training a diffusion weight generator with meta-learning and a local-consistency loss lets it hit intermediate optimizer checkpoints on schedule and end at the global optimum.

  21. Exploring Preference-Guided Diffusion Model for Cross-Domain Recommendation

    cs.IR 2025-01 conditional novelty 6.0 of 10

    DMCDR conditions a diffusion recommender on a source-domain preference representation and reports large MAE/RMSE gains over prior cold-start cross-domain recommenders on three Amazon scenarios.

  22. Generative Physical AI in Vision: A Survey

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A structured review that categorizes physics-aware generative models in vision into explicit-simulation and implicit-learning families and proposes six integration paradigms.

  23. Fast and Robust Visuomotor Riemannian Flow Matching Policy

    cs.RO 2024-12 conditional novelty 6.0 of 10

    A stable Riemannian flow matching policy (SRFMP) that converges to the target action distribution on manifolds, evaluated on ten robotic tasks against diffusion and consistency baselines.

  24. AppGen: Mobility-aware App Usage Behavior Generation for Mobile Users

    cs.HC 2024-12 conditional novelty 6.0 of 10

    An autoregressive diffusion model generates realistic mobile app usage sequences from users' spatio-temporal trajectories, beating state-of-the-art baselines on distributional fidelity metrics on two telecom datasets.

  25. Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    By recasting text-to-image generation as multi-frame video generation and adding a differential camera encoder, the method achieves camera intrinsic control with scene consistency, outperforming current text-to-image ...

  26. RAW-Diffusion: RGB-Guided Diffusion Models for High-Fidelity RAW Image Generation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    RAW-Diffusion generates high-fidelity RAW images from RGB inputs with state-of-the-art PSNR/SSIM on four DSLR datasets, and needs as few as 25 training images.

  27. Controlling Diversity at Inference: Guiding Diffusion Recommender Models with Targeted Category Preferences

    cs.IR 2024-11 conditional novelty 6.0 of 10

    D3Rec controls recommendation diversity at inference by conditioning a diffusion model on a target category distribution, and it outperforms comparable baselines on three datasets.

  28. Free-Lunch Augmentation by Revisiting Diffusion-Based Data Generation for Cross-Domain Few-Shot Object Detection

    cs.CV 2026-08 conditional novelty 5.0 of 10

    SITN improves cross-domain few-shot detection and segmentation by using weakened-noise diffusion and background inpainting to synthesize helpful training images, outperforming prior methods on all reported benchmarks.

  29. On the Redundancy of Timestep Embeddings in Diffusion Models

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    Under high-dimensional concentration conditions, the diffusion denoising objective admits the same global minimizer without timestep embeddings, and time-agnostic U-Nets/DiTs empirically match or improve FID on CelebA...

  30. Agents of Diffusion: Enhancing Diffusion Language Models with Multi-Agent Reinforcement Learning for Structured Data Generation (Extended Version)

    cs.MA 2026-01 reject novelty 5.0 of 10

    AoD pairs a frozen diffusion language model with two LLM agents that iteratively rewrite prompts from natural-language feedback, reporting better JSON diversity and validity, though the claimed RL mechanism and theore...

  31. TFCDiff: Robust ECG Denoising via Time-Frequency Complementary Diffusion

    eess.SP 2025-11 conditional novelty 5.0 of 10

    TFCDiff, a diffusion model trained on truncated DCT coefficients of 10-second ECG segments with time-frequency feature fusion, outperforms eight benchmark denoisers, including on the unseen real SimEMG noise dataset.

  32. Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance

    cs.RO 2025-09 conditional novelty 5.0 of 10

    A plug-and-play fine-tuning method using two VAEs and a latent-distance guidance loss improves cross-embodiment and cross-task success rates of diffusion- and flow-based VLA policies.

  33. Bayesian Radio Map Estimation: Fundamentals and Implementation via Diffusion Models

    eess.SP 2025-08 unverdicted novelty 5.0 of 10

    A diffusion-based Bayesian radio map estimator recovers the posterior of the map, enabling MMSE estimates of arbitrary map functionals while training only for signal power.

  34. Constrained Diffusers for Safe Planning and Control

    eess.SY 2025-06 conditional novelty 5.0 of 10

    Constrained Diffusers enforces trajectory constraints on pre-trained diffusion models without retraining by replacing the reverse process with constrained Langevin sampling.

  35. Frugal Incremental Generative Modeling using Variational Autoencoders

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A single replay-free conditional VAE with fixed-point-separated Gaussian priors and null-space gradient projection achieves competitive continual classification with drastically reduced memory.

  36. Perfect diffusion is $\mathsf{TC}^0$ -- Bad diffusion is Turing-complete

    cs.CC 2025-04 conditional novelty 5.0 of 10

    Perfect diffusion models with TC0 score networks are computationally limited to TC0, while unconstrained diffusion-like SDEs can be Turing-complete.

  37. Generalized Visual Relation Detection with Diffusion Models

    cs.CV 2025-04 conditional novelty 5.0 of 10

    Diff-VRD generates visual relation phrases with a diffusion model conditioned on CLIP features, aiming to detect interactions beyond dataset labels and scoring them with text-to-image retrieval and SPICE.

  38. Variational Rectified Flow Matching

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Variational rectified flow matching uses a latent variable to capture multi-modal velocity fields, improving generation quality and enabling controllable sampling.

  39. HistoSmith: Single-Stage Histology Image-Label Generation via Conditional Latent Diffusion for Enhanced Cell Segmentation and Classification

    cs.CV 2025-02 conditional novelty 5.0 of 10

    A conditional latent diffusion model jointly generates histology images, distance maps, and cell-type masks, and adding its outputs to real training data improves cell segmentation and classification by about 2-3% on ...

  40. Collaborative Diffusion Model for Recommender System

    cs.IR 2025-01 conditional novelty 5.0 of 10

    CDiff4Rec improves diffusion recommenders by injecting item-content pseudo-users and real-user neighbor predictions into the denoising objective, beating DiffRec and other baselines on Yelp, Amazon-Game, and Citeulike-t.

  41. Exploring Structured Semantic Priors Underlying Diffusion Score for Test-time Adaptation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    DUSA adapts classifiers and segmenters at test time by matching their predictions to conditional noise estimates from a pre-trained diffusion model, using a single timestep and active class selection.

  42. VideoDirector: Precise Video Editing via Text-to-Video Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A video editing pipeline that extends null-text inversion and attention control to text-to-video diffusion models, using spatial-temporal decoupled guidance and multi-frame null embeddings to achieve more temporally c...

  43. Machine-Learning-Assisted Photonic Device Development: A Multiscale Approach from Theory to Characterization

    physics.optics 2025-06 accept novelty 4.0 of 10

    This review organizes machine-learning-assisted photonic device development into a five-step Bayesian framework spanning theory, simulation, design, fabrication, and characterization.

  44. Automated Learning of Semantic Embedding Representations for Diffusion Models

    cs.LG 2025-05 conditional novelty 4.0 of 10

    A diffusion model with a timestep-conditioned encoder learns embeddings that reach competitive linear probe accuracy on four of six datasets, with the optimal timestep varying by dataset.

  45. Sch\"odinger Bridge Type Diffusion Models as an Extension of Variational Autoencoders

    cs.LG 2024-12 conditional novelty 4.0 of 10

    A VAE-style derivation shows the Schrödinger bridge diffusion objective decomposes into a prior loss and a drift-matching term.

  46. Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A survey maps the field of MLLM explainability and interpretability into data, model, and training and inference perspectives.

  47. Application of Generative Adversarial Network (GAN) for Synthetic Training Data Creation to improve performance of ANN Classifier for extracting Built-Up pixels from Landsat Satellite Imagery

    cs.CV 2025-01 conditional novelty 3.0 of 10

    A GAN generates synthetic built-up pixels that, when added to a tiny training set, raise an ANN classifier's accuracy and kappa on Landsat7 imagery.

  48. PXGen: A Post-hoc Explainable Method for Generative Models

    cs.LG 2025-01 reject novelty 3.0 of 10

    PXGen is a post-hoc, training-free explanation framework that scores anchor samples with intrinsic and extrinsic criteria, groups them by thresholds, and selects representative examples via k-dispersion or k-center.

  49. Diffusion Models for Hyperspectral Image Analysis: A Comprehensive Review

    eess.IV 2025-05 conditional novelty 2.0 of 10

    A literature review that organizes diffusion-model work for hyperspectral imaging into eight task categories and compiles comparative performance tables from prior papers.

  50. Generative Models: Principles, Architectures, and Applications

    cs.AI 2026-08 conditional novelty 1.0 of 10

    A comprehensive textbook-style review of generative modeling, from ELBO and EM through diffusion models, flow matching, and modern sampling architectures, with no new research findings.

  51. Fundamentals of Data-Driven Approaches to Acoustic Signal Detection, Filtering, and Transformation

    eess.AS 2025-08 unverdicted novelty 1.0 of 10

    A systematic survey of deep-learning acoustic signal processing, organized as detection, filtering, and transformation tasks, with network modules, loss construction, and five application areas.

Pith tools