Pith. sign in

REVIEW 27 cited by

ILVR: Conditioning Method for Denoising Diffusion Probabilistic Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.02938 v2 pith:OFUO5GV7 submitted 2021-08-06 cs.CV

classification cs.CV
keywords ddpmimagegenerationimagesmethodgenerativeilvrprocess
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Denoising diffusion probabilistic models (DDPM) have shown remarkable performance in unconditional image generation. However, due to the stochasticity of the generative process in DDPM, it is challenging to generate images with the desired semantics. In this work, we propose Iterative Latent Variable Refinement (ILVR), a method to guide the generative process in DDPM to generate high-quality images based on a given reference image. Here, the refinement of the generative process in DDPM enables a single DDPM to sample images from various sets directed by the reference image. The proposed ILVR method generates high-quality images while controlling the generation. The controllability of our method allows adaptation of a single DDPM without any additional learning in various image generation tasks, such as generation from various downsampling factors, multi-domain image translation, paint-to-image, and editing with scribbles.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 27 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RAG-Audio: Retrieval-Augmented Generation for Faithful Brain-to-Audio Reconstruction

    cs.SD 2026-08 conditional novelty 6.0 of 10

    RAG-Audio starts frozen audio generators from a retrieved exemplar of the fMRI-decoded CLAP embedding, raising 10-way stimulus identification from 0.14-0.18 to 0.40-0.43 on Brain2Music and cutting FAD by about 10x.

  2. MENO: MeanFlow-Enhanced Neural Operators for Dynamical Systems

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    MENO restores multi-scale structure in neural-operator PDE surrogates via one-step improved MeanFlow, claiming up to 2× better power-spectrum accuracy and up to 14× faster inference than DDIM enhancement.

  3. RubricRL: Simple Generalizable Rewards for Text-to-Image Generation

    cs.CV 2025-11 conditional novelty 6.0 of 10

    Using an LLM to generate prompt-specific visual rubrics and grade each criterion independently gives a more interpretable reward that improves text-to-image model alignment beyond composite and learned scalar rewards.

  4. Score-based Diffusion Model for Unpaired Virtual Histology Staining

    eess.IV 2025-06 conditional novelty 6.0 of 10

    An unpaired, mutual-information-guided diffusion model translates H&E histology images into IHC images with improved structural and staining fidelity.

  5. Accelerating Diffusion-based Super-Resolution with Dynamic Time-Spatial Sampling

    cs.CV 2025-05 conditional novelty 6.0 of 10

    TSS accelerates diffusion super-resolution by concentrating denoising steps in early and late iterations and adapting the schedule per image region, improving perceptual scores with fewer steps.

  6. Dissecting Bit-Level Scaling Laws in Quantizing Vision Generative Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Token-based language-style vision models (VAR, LlamaGen) tolerate quantization better than diffusion models, and a custom TopKLD distillation loss pushes their low-bit scaling roughly one precision level higher.

  7. InpDiffusion: Image Inpainting Localization via Conditional Diffusion Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    InpDiffusion treats image inpainting localization as a conditional mask-generation task with a diffusion model, using edge supervision to refine boundaries, and reports state-of-the-art AUC on Inpaint32K, DID, AutoSpl...

  8. SyncDiff: Synchronized Motion Diffusion for Multi-Body Human-Object Interaction Synthesis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SyncDiff synthesizes multi-body human-object interaction motions with one diffusion model plus explicit synchronization and frequency decomposition, improving contact and action-quality metrics over prior methods on f...

  9. Oscillation Inversion: Understand the structure of Large Flow Model through the Lens of Inversion Method

    cs.CV 2024-11 reject novelty 6.0 of 10

    Inverting Flux images with fixed-point iteration oscillates between semantically coherent latent clusters, and this oscillation is repurposed into a distribution-transfer editing and enhancement method.

  10. FlowForm: Synergizing Fluid Physics with Topological Consistency for Satellite Flood Synthesis

    cs.CV 2026-08 conditional novelty 5.0 of 10

    FlowForm is a diffusion model that adds shallow-water-equation penalties and terrain-conditioned adapters to synthesize flood satellite images, and reports top scores on a new 10,000-pair dataset.

  11. Provable diffusion-based posterior sampling for linear inverse problems via DDIM

    cs.LG 2026-07 reject novelty 5.0 of 10

    A SVD-based, coordinate-wise DDIM sampler is claimed to asymptotically sample from the posterior for noisy linear inverse problems, but the proof's posterior identification step does not follow from the stated updates.

  12. Translationese as a Rational Response to Translation Task Difficulty

    cs.CL 2026-03 unverdicted novelty 5.0 of 10

    Translationese is partly predictable from quantifiable translation-task difficulty, especially cross-lingual transfer load, more so for English-to-German than the reverse.

  13. Reconstructing Multi-Scale Physical Fields from Extremely Sparse Measurements with an Autoencoder-Diffusion Cascade

    cs.LG 2025-12 conditional novelty 5.0 of 10

    A cascade of a functional autoencoder (coarse structure) and a residual conditional diffusion model (fine details), with mask-cascade training and manifold-constrained gradients, reconstructs sparse-sensed physical fields.

  14. FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution

    cs.CV 2025-09 conditional novelty 5.0 of 10

    FS-Diff is a diffusion model that jointly fuses and super-resolves low-resolution multimodal image pairs using clarity-aware CLIP semantics and a bidirectional Mamba feature extractor.

  15. Time-variant Image Inpainting via Interactive Distribution Transition Estimation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    The authors introduce time-variant image inpainting (TAMP), a benchmark (TAMP-Street), and InDiTE-Diff, a diffusion-based method with a semantic complementation module that outperforms prior reference-guided inpaintin...

  16. Solving Inverse Problems via Diffusion-Based Priors: An Approximation-Free Ensemble Sampling Approach

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A weighted-particle sampler evolves the posterior through the diffusion model's reverse dynamics, with theoretical error bounds and improved image reconstructions.

  17. Restoring Real-World Images with an Internal Detail Enhancement Diffusion Model

    cs.CV 2025-05 conditional novelty 5.0 of 10

    This paper fine-tunes ControlNet on a frozen Stable Diffusion prior using a self-regularization objective that conditions on both the degraded input and a DDIM-estimated clean version, improving perceptual quality and...

  18. Semantic-Guided Diffusion Model for Single-Step Image Super-Resolution

    cs.CV 2025-05 conditional novelty 5.0 of 10

    SAMSR combines SAM segmentation masks with noise shaping and pixel-wise sampling to improve the perceptual quality of single-step diffusion super-resolution, with modest gains over SinSR.

  19. CountDiffusion: Text-to-Image Synthesis with Training-Free Counting-Guidance Diffusion

    cs.CV 2025-05 reject novelty 5.0 of 10

    CountDiffusion improves object-count accuracy in text-to-image diffusion by detecting objects in a one-step predicted image and applying attention-map guidance to add or remove instances.

  20. Noise Controlled CT Super-Resolution with Conditional Diffusion Model

    cs.CV 2025-02 conditional novelty 5.0 of 10

    A conditional diffusion model trained on hybrid noise-matched simulation and real segmented bone data improves CT spatial resolution without the noise amplification seen in simulation-only training.

  21. Efficient Weighted Sampling via Score-based Generative Models

    cs.LG 2025-02 conditional novelty 5.0 of 10

    A score-based generative model can be reused without retraining to sample from arbitrary weighted versions of its base distribution by adding an approximated guidance term during reverse diffusion.

  22. Exploring the latent space of diffusion models directly through singular value decomposition

    cs.CV 2025-02 reject novelty 5.0 of 10

    The authors report that singular value decomposition of diffusion latent codes reveals stable, order-mobile attribute directions and propose Attribute Vector Integration, a per-pair MLP-based editor that transfers tex...

  23. Exploring Structured Semantic Priors Underlying Diffusion Score for Test-time Adaptation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    DUSA adapts classifiers and segmenters at test time by matching their predictions to conditional noise estimates from a pre-trained diffusion model, using a single timestep and active class selection.

  24. Photovoltaic Defect Image Generator with Boundary Alignment Smoothing Constraint for Domain Shift Mitigation

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A Stable Diffusion generator with text and image adapters plus box-constrained attention produces photovoltaic defect images that improve cross-line defect detection by up to 6.3 mAP.

  25. Causal Disentanglement for Robust Long-tail Medical Image Generation

    cs.CV 2025-04 conditional novelty 4.0 of 10

    A VAE and latent diffusion pipeline disentangles identity from pathology in chest X-rays and generates counterfactual images with text-guided pathology synthesis and inference-time noise optimization for long-tail categories.

  26. Frequency-Aware Guidance for Blind Image Restoration via Diffusion Models

    cs.CV 2024-11 conditional novelty 4.0 of 10

    A plug-and-play wavelet frequency guidance loss improves blind image restoration in diffusion models, giving up to 3.72 dB PSNR gain on motion deblurring.

  27. Projection-Based Correction for Enhancing Deep Inverse Networks

    cs.LG 2025-05 conditional novelty 2.0 of 10

    A projection step that forces a deep network's reconstruction to satisfy y = Ax gives small PSNR gains in low-noise imaging tests, but the supporting theory is a restatement of the definition of a well-trained network.

Pith tools