REVIEW 51 cited by
Understanding Diffusion Models: A Unified Perspective
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Diffusion models have shown incredible capabilities as generative models; indeed, they power the current state-of-the-art models on text-conditioned image generation such as Imagen and DALL-E 2. In this work we review, demystify, and unify the understanding of diffusion models across both variational and score-based perspectives. We first derive Variational Diffusion Models (VDM) as a special case of a Markovian Hierarchical Variational Autoencoder, where three key assumptions enable tractable computation and scalable optimization of the ELBO. We then prove that optimizing a VDM boils down to learning a neural network to predict one of three potential objectives: the original source input from any arbitrary noisification of it, the original source noise from any arbitrarily noisified input, or the score function of a noisified input at any arbitrary noise level. We then dive deeper into what it means to learn the score function, and connect the variational perspective of a diffusion model explicitly with the Score-based Generative Modeling perspective through Tweedie's Formula. Lastly, we cover how to learn a conditional distribution using diffusion models via guidance.
Forward citations
Cited by 51 Pith papers
-
Personalized Federated Learning via Variance-Aware Nonparametric Empirical Bayes
VANEB generalizes nonparametric empirical Bayes to parameter-dependent noise and uses it to personalize federated models by shrinking local estimates toward a learned population prior.
-
Bayesian Experimental Design via Score Matching
SCOREBED isolates EIG double intractability in a policy-independent score-matching stage, then trains design policies with a singly intractable gradient estimator, enabling cheap multi-policy selection.
-
To Erase, or Not to Erase: Robust Training-Free Concept Erasure with Preservation aware Adaptive Ranked Subspace Expansion
Diffusion-grounded erase/retain retrieval plus retain-orthogonal value projection and trigger-guided subspace expansion erases concepts more robustly than prior CETs while keeping FID/CLIP near the unedited model.
-
Tightening the Score Matching Gap for Diffusion Models
Tighter score-matching gap bounds for diffusion models via entropy flows, LSI and reflection couplings show that low-noise score accuracy dominates sample quality metrics.
-
CDCP: Conditional Diffusion Model with Contextual Prompts for Multi-task Offline Safe Reinforcement Learning
A conditional diffusion model with contextual prompts and classifier-free cost guidance learns a shared safe multi-task policy from offline data and meets varying cost limits without retraining.
-
TRE: Training-Free Hallucination Detection for Diffusion Language Models
TRE weights the entropy of newly revealed tokens by denoising step and detects hallucinations in diffusion LLMs with no training and a single generation.
-
REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion
Nonlinear multi-layer compression of frozen VFM patch semantics, jointly denoised with VAE latents, improves ImageNet 256x256 FID (12.9 vs 15.2 for REG at SiT-B/2, 400K) and accelerates convergence over REPA/ReDi/REG.
-
Joint Model-based Model-free Diffusion for Planning with Constraints
JM2D samples diffusion plans and safety-filter corrections jointly using a single importance-sampling-guided diffusion process, improving task success and reducing safety-filter interventions.
-
Dual Orthogonal Guidance for Robust Diffusion-based Handwritten Text Generation
A new guidance method for diffusion-based text generation subtracts the orthogonal component of a negative prompt from a positive prompt to reduce artifacts and increase style variation.
-
Clustering via Self-Supervised Diffusion
CLUDI trains a student to imitate stochastic diffusion-generated cluster assignments on pre-trained DINO image features and averages multiple assignments to cluster images.
-
Hyper-modal Imputation Diffusion Embedding with Dual-Distillation for Federated Multimodal Knowledge Graph Completion
A masked diffusion imputer plus dual distillation (MMFeD3-HidE) improves link prediction on a new federated multimodal knowledge graph benchmark with 50% missing visual/textual modalities.
-
Exploiting the Exact Denoising Posterior Score in Training-Free Guidance of Diffusion Models
An exact denoising posterior score is derived and used to compute time-dependent DPS step sizes that transfer to colorization, inpainting, and super-resolution.
-
Fusion of multi-source precipitation records via coordinate-based generative model
A coordinate-based diffusion model fuses multi-source precipitation records and corrects biases in unseen operational forecasts.
-
Diffusion Counterfactual Generation with Semantic Abduction
Diffusion-based causal image counterfactuals with semantic abduction improve identity preservation at a small cost in intervention effectiveness, demonstrated on Morpho-MNIST, CelebA-HQ, and mammogram artifact removal.
-
Sparse Autoencoders, Again?
VAEase gates the VAE decoder input by the encoder's variance, combining sparse-autoencoder adaptive sparsity with a hyperparameter-free loss; a global-minimizer theorem says active latent dimensions recover per-manifo...
-
Autoregressive regularized score-based diffusion models for multi-scenarios fluid flow prediction
A regularized autoregressive score-based diffusion model predicts turbulent flows across multiple scenarios, with the variance-preserving SDE formulation performing best.
-
Unsupervised Learning for Class Distribution Mismatch
UCDM builds positive and negative image pairs with a text-to-image diffusion model and trains an open-set classifier with no instance labels, outperforming semi-supervised baselines on three benchmarks.
-
InstaRevive: One-Step Image Enhancement via Dynamic Score Matching
A one-step image enhancer that combines dynamically controlled score-based distillation with caption prompts, matching multi-step diffusion quality on face and super-resolution benchmarks.
-
CDM: Contact Diffusion Model for Multi-Contact Point Localization
A Contact Diffusion Model conditioned on force/torque readings and a signed distance field localizes single and dual robot-arm contacts with 0.44 cm and 1.24 cm real-world error.
-
Learning to Learn Weight Generation via Local Consistency Diffusion
Training a diffusion weight generator with meta-learning and a local-consistency loss lets it hit intermediate optimizer checkpoints on schedule and end at the global optimum.
-
Exploring Preference-Guided Diffusion Model for Cross-Domain Recommendation
DMCDR conditions a diffusion recommender on a source-domain preference representation and reports large MAE/RMSE gains over prior cold-start cross-domain recommenders on three Amazon scenarios.
-
Generative Physical AI in Vision: A Survey
A structured review that categorizes physics-aware generative models in vision into explicit-simulation and implicit-learning families and proposes six integration paradigms.
-
Fast and Robust Visuomotor Riemannian Flow Matching Policy
A stable Riemannian flow matching policy (SRFMP) that converges to the target action distribution on manifolds, evaluated on ten robotic tasks against diffusion and consistency baselines.
-
AppGen: Mobility-aware App Usage Behavior Generation for Mobile Users
An autoregressive diffusion model generates realistic mobile app usage sequences from users' spatio-temporal trajectories, beating state-of-the-art baselines on distributional fidelity metrics on two telecom datasets.
-
Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis
By recasting text-to-image generation as multi-frame video generation and adding a differential camera encoder, the method achieves camera intrinsic control with scene consistency, outperforming current text-to-image ...
-
RAW-Diffusion: RGB-Guided Diffusion Models for High-Fidelity RAW Image Generation
RAW-Diffusion generates high-fidelity RAW images from RGB inputs with state-of-the-art PSNR/SSIM on four DSLR datasets, and needs as few as 25 training images.
-
Controlling Diversity at Inference: Guiding Diffusion Recommender Models with Targeted Category Preferences
D3Rec controls recommendation diversity at inference by conditioning a diffusion model on a target category distribution, and it outperforms comparable baselines on three datasets.
-
Free-Lunch Augmentation by Revisiting Diffusion-Based Data Generation for Cross-Domain Few-Shot Object Detection
SITN improves cross-domain few-shot detection and segmentation by using weakened-noise diffusion and background inpainting to synthesize helpful training images, outperforming prior methods on all reported benchmarks.
-
On the Redundancy of Timestep Embeddings in Diffusion Models
Under high-dimensional concentration conditions, the diffusion denoising objective admits the same global minimizer without timestep embeddings, and time-agnostic U-Nets/DiTs empirically match or improve FID on CelebA...
-
Agents of Diffusion: Enhancing Diffusion Language Models with Multi-Agent Reinforcement Learning for Structured Data Generation (Extended Version)
AoD pairs a frozen diffusion language model with two LLM agents that iteratively rewrite prompts from natural-language feedback, reporting better JSON diversity and validity, though the claimed RL mechanism and theore...
-
TFCDiff: Robust ECG Denoising via Time-Frequency Complementary Diffusion
TFCDiff, a diffusion model trained on truncated DCT coefficients of 10-second ECG segments with time-frequency feature fusion, outperforms eight benchmark denoisers, including on the unseen real SimEMG noise dataset.
-
Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance
A plug-and-play fine-tuning method using two VAEs and a latent-distance guidance loss improves cross-embodiment and cross-task success rates of diffusion- and flow-based VLA policies.
-
Bayesian Radio Map Estimation: Fundamentals and Implementation via Diffusion Models
A diffusion-based Bayesian radio map estimator recovers the posterior of the map, enabling MMSE estimates of arbitrary map functionals while training only for signal power.
-
Constrained Diffusers for Safe Planning and Control
Constrained Diffusers enforces trajectory constraints on pre-trained diffusion models without retraining by replacing the reverse process with constrained Langevin sampling.
-
Frugal Incremental Generative Modeling using Variational Autoencoders
A single replay-free conditional VAE with fixed-point-separated Gaussian priors and null-space gradient projection achieves competitive continual classification with drastically reduced memory.
-
Perfect diffusion is $\mathsf{TC}^0$ -- Bad diffusion is Turing-complete
Perfect diffusion models with TC0 score networks are computationally limited to TC0, while unconstrained diffusion-like SDEs can be Turing-complete.
-
Generalized Visual Relation Detection with Diffusion Models
Diff-VRD generates visual relation phrases with a diffusion model conditioned on CLIP features, aiming to detect interactions beyond dataset labels and scoring them with text-to-image retrieval and SPICE.
-
Variational Rectified Flow Matching
Variational rectified flow matching uses a latent variable to capture multi-modal velocity fields, improving generation quality and enabling controllable sampling.
-
HistoSmith: Single-Stage Histology Image-Label Generation via Conditional Latent Diffusion for Enhanced Cell Segmentation and Classification
A conditional latent diffusion model jointly generates histology images, distance maps, and cell-type masks, and adding its outputs to real training data improves cell segmentation and classification by about 2-3% on ...
-
Collaborative Diffusion Model for Recommender System
CDiff4Rec improves diffusion recommenders by injecting item-content pseudo-users and real-user neighbor predictions into the denoising objective, beating DiffRec and other baselines on Yelp, Amazon-Game, and Citeulike-t.
-
Exploring Structured Semantic Priors Underlying Diffusion Score for Test-time Adaptation
DUSA adapts classifiers and segmenters at test time by matching their predictions to conditional noise estimates from a pre-trained diffusion model, using a single timestep and active class selection.
-
VideoDirector: Precise Video Editing via Text-to-Video Models
A video editing pipeline that extends null-text inversion and attention control to text-to-video diffusion models, using spatial-temporal decoupled guidance and multi-frame null embeddings to achieve more temporally c...
-
Machine-Learning-Assisted Photonic Device Development: A Multiscale Approach from Theory to Characterization
This review organizes machine-learning-assisted photonic device development into a five-step Bayesian framework spanning theory, simulation, design, fabrication, and characterization.
-
Automated Learning of Semantic Embedding Representations for Diffusion Models
A diffusion model with a timestep-conditioned encoder learns embeddings that reach competitive linear probe accuracy on four of six datasets, with the optimal timestep varying by dataset.
-
Sch\"odinger Bridge Type Diffusion Models as an Extension of Variational Autoencoders
A VAE-style derivation shows the Schrödinger bridge diffusion objective decomposes into a prior loss and a drift-matching term.
-
Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey
A survey maps the field of MLLM explainability and interpretability into data, model, and training and inference perspectives.
-
Application of Generative Adversarial Network (GAN) for Synthetic Training Data Creation to improve performance of ANN Classifier for extracting Built-Up pixels from Landsat Satellite Imagery
A GAN generates synthetic built-up pixels that, when added to a tiny training set, raise an ANN classifier's accuracy and kappa on Landsat7 imagery.
-
PXGen: A Post-hoc Explainable Method for Generative Models
PXGen is a post-hoc, training-free explanation framework that scores anchor samples with intrinsic and extrinsic criteria, groups them by thresholds, and selects representative examples via k-dispersion or k-center.
-
Diffusion Models for Hyperspectral Image Analysis: A Comprehensive Review
A literature review that organizes diffusion-model work for hyperspectral imaging into eight task categories and compiles comparative performance tables from prior papers.
-
Generative Models: Principles, Architectures, and Applications
A comprehensive textbook-style review of generative modeling, from ELBO and EM through diffusion models, flow matching, and modern sampling architectures, with no new research findings.
-
Fundamentals of Data-Driven Approaches to Acoustic Signal Detection, Filtering, and Transformation
A systematic survey of deep-learning acoustic signal processing, organized as detection, filtering, and transformation tasks, with network modules, loss construction, and five application areas.
Discussion (0). Continue with ORCID to comment.