Pith. sign in

REVIEW 41 cited by

Tackling the Generative Learning Trilemma with Denoising Diffusion GANs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.07804 v2 pith:NOITT5AN submitted 2021-12-15 cs.LG stat.ML

classification cs.LGstat.ML
keywords denoisingdiffusionmodelsgansgenerativemodelsamplesampling
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

A wide variety of deep generative models has been developed in the past decade. Yet, these models often struggle with simultaneously addressing three key requirements including: high sample quality, mode coverage, and fast sampling. We call the challenge imposed by these requirements the generative learning trilemma, as the existing models often trade some of them for others. Particularly, denoising diffusion models have shown impressive sample quality and diversity, but their expensive sampling does not yet allow them to be applied in many real-world applications. In this paper, we argue that slow sampling in these models is fundamentally attributed to the Gaussian assumption in the denoising step which is justified only for small step sizes. To enable denoising with large steps, and hence, to reduce the total number of denoising steps, we propose to model the denoising distribution using a complex multimodal distribution. We introduce denoising diffusion generative adversarial networks (denoising diffusion GANs) that model each denoising step using a multimodal conditional GAN. Through extensive evaluations, we show that denoising diffusion GANs obtain sample quality and diversity competitive with original diffusion models while being 2000$\times$ faster on the CIFAR-10 dataset. Compared to traditional GANs, our model exhibits better mode coverage and sample diversity. To the best of our knowledge, denoising diffusion GAN is the first model that reduces sampling cost in diffusion models to an extent that allows them to be applied to real-world applications inexpensively. Project page and code can be found at https://nvlabs.github.io/denoising-diffusion-gan

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 41 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The GAN is dead; long live the GAN! A Modern GAN Baseline

    cs.LG 2025-01 conditional novelty 7.0 of 10

    A minimalist GAN with a regularized relativistic loss and modern backbone matches or beats StyleGAN2 and several diffusion models on standard FID benchmarks.

  2. ROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative Modelling

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A single-step IMLE generator with per-stage supervision and a robust loss reports FID 2.56 on ImageNet-256 by filtering ~5% of samples at test time.

  3. MiAD: Mirage Atom Diffusion for De Novo Crystal Generation

    cs.LG 2025-11 conditional novelty 6.0 of 10

    Mirage infusion lets crystal diffusion models vary atom counts during generation and raises the S.U.N. rate on MP-20 to 8.2%.

  4. Friend or Foe

    q-bio.QM 2025-08 conditional novelty 6.0 of 10

    Friend or Foe is a 64-dataset compendium of 26M+ simulated bacterial interaction environments, with benchmarks showing deep tabular models classify interaction type with mean MCC of about 0.64.

  5. Quantum latent distributions in deep generative models

    cs.LG 2025-08 conditional novelty 6.0 of 10

    Quantum latent distributions from boson samplers are shown in theory to expand the output distribution class of invertible Lipschitz generators, and in GAN benchmarks on QM9 to beat Gaussian, Bernoulli, and distinguis...

  6. D2Diff : A Dual Domain Diffusion Model for Accurate Multi-Contrast MRI Synthesis

    eess.IV 2025-06 conditional novelty 6.0 of 10

    A dual-domain diffusion model combining spatial and frequency guidance with a shared critic and uncertainty mask loss improves multi-contrast MRI synthesis over existing baselines.

  7. Dual-Expert Consistency Model for Efficient and High-Quality Video Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    By training a semantic expert and a LoRA-based detail expert, DCM reaches nearly teacher-level VBench scores with 4-step video sampling on HunyuanVideo and CogVideoX.

  8. GuideSR: Rethinking Guidance for One-Step High-Fidelity Diffusion-Based Super-Resolution

    eess.IV 2025-05 conditional novelty 6.0 of 10

    GuideSR pairs a full-resolution guidance branch with a one-step latent diffusion branch and improves PSNR, SSIM, LPIPS, DISTS, and FID on DIV2K-Val, RealSR, and DRealSR.

  9. Flow Along the K-Amplitude for Generative Modeling

    cs.LG 2025-04 reject novelty 6.0 of 10

    K-Flow trains flow-matching models with frequency scale as time, enabling competitive image and molecule generation plus scale-level control of outputs.

  10. ScanEdit: Hierarchically-Guided Functional 3D Scan Editing

    cs.CV 2025-04 conditional novelty 6.0 of 10

    ScanEdit uses hierarchical scene graphs and LLM-based planning, placement, and optimization to rearrange objects in real-world 3D scans from text instructions.

  11. Improved Training Technique for Latent Consistency Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Latent consistency models can be trained from scratch for one- or two-step generation when Huber loss is replaced by Cauchy loss and combined with early-timestep diffusion loss, OT coupling, an adaptive scaling schedu...

  12. SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SnapGen is a 379M-parameter UNet with cross-architecture distillation and a 1.38M-parameter decoder that generates 1024x1024 images on a phone in about 1.4 seconds, with GenEval 0.66 and ImageNet FID 2.06.

  13. Adversarial Diffusion Compression for Real-World Image Super-Resolution

    eess.IV 2024-11 conditional novelty 6.0 of 10

    AdcSR distills OSEDiff into a pruned diffusion-GAN that cuts inference time 3.7x and parameters 74% while achieving comparable super-resolution quality.

  14. Layer-Parallel Inference Reduces Encrypted Nonlinear Depth in Transformers

    cs.LG 2026-07 conditional novelty 5.5 of 10

    SNLP reduces symbolic FHE bootstraps from 53 to 20 on a 0.5B model with +1.2% PPL degradation and lower polynomial-error amplification than sequential inference.

  15. Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Training on the best of K generated candidates improves image, video, and language generative models, with the reported gains growing with scale and enabling single-pass end-to-end generation.

  16. FAIL: Flow Matching Adversarial Imitation Learning for Image Generation

    cs.CV 2026-02 conditional novelty 5.0 of 10

    Post-training of flow matching can be framed as adversarial imitation learning, and the proposed FAIL methods improve FLUX's generation quality using 13K expert images without preference pairs.

  17. VarDiU: A Variational Diffusive Upper Bound for One-Step Diffusion Distillation

    cs.LG 2025-08 conditional novelty 5.0 of 10

    VarDiU optimizes a variational upper bound on the diffusive KL divergence with an unbiased gradient estimator, improving one-step generation on a 2D 40-Gaussian toy benchmark compared with Diff-Instruct.

  18. Inference Time Debiasing Concepts in Diffusion Models

    cs.GR 2025-08 reject novelty 5.0 of 10

    DeCoDi subtracts a biased-concept guidance term during diffusion inference to shift generated images away from targeted stereotypes, with evaluation on gender, ethnicity, and age.

  19. Turbulent Injection assisted by Diffusion Models for Scale Resolving Simulations

    physics.flu-dyn 2025-08 conditional novelty 5.0 of 10

    A Reynolds-conditioned diffusion model can generate DHIT turbulence boxes for LES/DNS inflow that match energy spectra and development length, though integral length scale and anisotropy are imperfect.

  20. fastWDM3D: Fast and Accurate 3D Healthy Tissue Inpainting

    eess.IV 2025-07 conditional novelty 5.0 of 10

    fastWDM3D, a wavelet diffusion model with a variance-preserving noise schedule and reconstruction losses, achieves high-quality 3D brain inpainting in two steps and about 1.81 seconds per image.

  21. Reversing Flow for Image Restoration

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ResFlow models HQ-to-LQ degradation as a deterministic augmented flow and inverts it via velocity matching, reporting state-of-the-art restoration in fewer than four sampling steps.

  22. A Robust Local Fr\'echet Regression Using Unbalanced Neural Optimal Transport with Applications to Dynamic Single-cell Genomics Data

    stat.AP 2025-06 reject novelty 5.0 of 10

    A neural-network local Fréchet regression using unbalanced optimal transport is introduced to interpolate single-cell distributions over time, with applications to three differentiation datasets.

  23. Few-Step Diffusion via Score identity Distillation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A few-step, largely data-free extension of Score identity Distillation reaches state-of-the-art FID and CLIP scores on SDXL at 1024x1024 with one or four generation steps.

  24. Ambient Denoising Diffusion Generative Adversarial Networks for Establishing Stochastic Object Models from Noisy Image Data

    cs.CV 2025-01 conditional novelty 5.0 of 10

    ADDGAN learns clean image distributions from noisy CT and DBT measurements by feeding generated objects through the known imaging operator, and it beats AmbientGAN baselines on FID and observer-task metrics.

  25. GMem: A Modular Approach for Ultra-Efficient Generative Models

    cs.CV 2024-12 conditional novelty 5.0 of 10

    GMem conditions diffusion models on a fixed bank of DINOv2 features and reports much lower FID at far fewer epochs than SiT and REPA baselines on ImageNet.

  26. Unpaired Modality Translation for Pseudo Labeling of Histology Images

    eess.IV 2024-12 conditional novelty 5.0 of 10

    Unpaired image translation between labeled and unlabeled microscopy domains can produce pseudo labels useful for segmentation, achieving 0.736 mean Dice for axons on SEM via the tutoring path.

  27. DogLayout: Denoising Diffusion GAN for Discrete and Continuous Layout Generation

    cs.CV 2024-11 conditional novelty 5.0 of 10

    DogLayout uses a denoising diffusion GAN with 4 to 12 timesteps to generate layout boxes and discrete labels, sampling up to 175 times faster than LayoutDM, with mixed FID results across tasks.

  28. Multi-scale Generative Modeling for Fast Sampling

    cs.AI 2024-11 conditional novelty 5.0 of 10

    WMGM generates 128x128 images by diffusing only low-frequency wavelet coefficients and using a shared multi-scale GAN to fill in high-frequency details, improving FID and cutting sampling time and parameters versus SG...

  29. Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution

    cs.CV 2026-07 reject novelty 4.0 of 10

    DDR-SR routes each real-world low-resolution image to one of two diffusion experts based on a high-frequency-loss difficulty score, using a low-compression VAE for hard images and a high-compression VAE for easy image...

  30. WaFusion: A Wavelet-Enhanced Diffusion Framework for Face Morph Generation

    cs.GR 2025-07 conditional novelty 4.0 of 10

    A hybrid wavelet-diffusion framework that morphs only the low-frequency wavelet sub-band to create efficient, high-quality face morphs.

  31. SupResDiffGAN a new approach for the Super-Resolution task

    eess.IV 2025-04 conditional novelty 4.0 of 10

    SupResDiffGAN combines latent-space diffusion with adversarial training and adaptive input noise, achieving faster super-resolution inference than SR3 and I2SB at comparable LPIPS quality.

  32. Conditional diffusion model with spatial attention and latent embedding for medical image segmentation

    eess.IV 2025-02 conditional novelty 4.0 of 10

    A conditional diffusion model with a per-timestep discriminator, spatial attention, and latent embedding reports state-of-the-art accuracy on three medical segmentation datasets using only 2 to 4 diffusion steps.

  33. Nested Annealed Training Scheme for Generative Adversarial Networks

    cs.CV 2025-01 reject novelty 4.0 of 10

    NATS, a nested annealed training scheme for GANs, is claimed to improve FID/IS on CIFAR10, LSUN, CelebA, and ImageNet64, but its key theoretical justification is not provided in the preprint.

  34. A New Formulation of Lipschitz Constrained With Functional Gradient Learning for GANs

    cs.CV 2025-01 reject novelty 4.0 of 10

    Li-CFG adds an ε-centered gradient penalty to the CFG GAN method and claims this enlarges the discriminator gradient norm, shrinking the latent neighborhood size and thereby increasing image diversity.

  35. E2ED^2:Direct Mapping from Noise to Data for Enhanced Diffusion Models

    cs.CV 2024-12 reject novelty 4.0 of 10

    E2ED2 fine-tunes a pretrained diffusion model end-to-end from pure noise to the target latent, improving few-step FID and CLIP on COCO30K and HW30K over the PixArt-delta baseline.

  36. Diffusion-Based Approaches in Medical Image Generation and Analysis

    eess.IV 2024-12 reject novelty 4.0 of 10

    CNNs trained only on diffusion-generated synthetic medical images achieved 72-91% accuracy on real test images across three domains, but no comparison to models trained on real data was performed.

  37. From Text to Pose to Image: Improving Diffusion Model Control and Quality

    cs.CV 2024-11 conditional novelty 4.0 of 10

    A text-to-pose transformer and a face-and-hand-aware pose adapter form a text-to-pose-to-image pipeline for diffusion models, beating the prior adapter baseline on 70 to 78 percent of test cases.

  38. Reinforcement Learning: From Algorithms To Foundation Models

    cs.AI 2026-07 conditional novelty 3.0 of 10

    A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.

  39. Computationally Efficient Diffusion Models in Medical Imaging: A Comprehensive Review

    eess.IV 2025-05 reject novelty 3.0 of 10

    A survey of DDPM, LDM, and WDM diffusion models for medical imaging, organized around training and inference efficiency.

  40. Text to Image Generation and Editing: A Survey

    cs.CV 2025-05 conditional novelty 3.0 of 10

    A broad survey of text-to-image generation and editing research from 2021 to 2024, organized by architecture and comparison tables.

  41. PQD: Post-training Quantization for Efficient Diffusion Models

    cs.CV 2024-12 reject novelty 3.0 of 10

    PQD calibrates diffusion-model quantization on time steps drawn from a tuned normal distribution, reporting competitive 8-bit FID on 64x64 ImageNet but much worse 4-bit FID and no quantitative text-to-image results.

Pith tools