Pith. sign in

REVIEW 27 cited by

ILVR: Conditioning Method for Denoising Diffusion Probabilistic Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.02938 v2 pith:OFUO5GV7 submitted 2021-08-06 cs.CV

classification cs.CV
keywords ddpmimagegenerationimagesmethodgenerativeilvrprocess
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Denoising diffusion probabilistic models (DDPM) have shown remarkable performance in unconditional image generation. However, due to the stochasticity of the generative process in DDPM, it is challenging to generate images with the desired semantics. In this work, we propose Iterative Latent Variable Refinement (ILVR), a method to guide the generative process in DDPM to generate high-quality images based on a given reference image. Here, the refinement of the generative process in DDPM enables a single DDPM to sample images from various sets directed by the reference image. The proposed ILVR method generates high-quality images while controlling the generation. The controllability of our method allows adaptation of a single DDPM without any additional learning in various image generation tasks, such as generation from various downsampling factors, multi-domain image translation, paint-to-image, and editing with scribbles.

Discussion (0). Sign in to comment.

Forward citations

Cited by 27 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

    cs.LG 2022-09 unverdicted novelty 8.0 of 10

    Rectified flow learns straight-path neural ODEs for distribution transport, yielding efficient generative models and domain transfers that work well even with a single simulation step.

  2. An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

    cs.CV 2022-08 unverdicted novelty 8.0 of 10

    Textual Inversion learns a single embedding vector from a few images to represent personal concepts inside the text embedding space of a frozen text-to-image model, enabling their composition in natural language prompts.

  3. Spectral Guidance for Flexible and Efficient Control of Diffusion Models

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    Spectral Guidance learns singular functions via self-supervised objective to project guidance signals onto diffusion sampling trajectories, enabling stable control without retraining or backpropagation and improving C...

  4. Latent Fourier Transform

    cs.SD 2026-04 unverdicted novelty 7.0 of 10

    LatentFT uses latent-space Fourier transforms and frequency masking in diffusion autoencoders to enable timescale-specific manipulation of musical structure in generative models.

  5. Conflated Inverse Modeling to Generate Diverse and Temperature-Change Inducing Urban Vegetation Patterns

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    A diffusion generative inverse model conditioned on temperature targets produces diverse, physically plausible urban vegetation patterns that achieve specified regional temperature shifts.

  6. LPNSR: Optimal Noise-Guided Diffusion Image Super-Resolution Via Learnable Noise Prediction

    cs.CV 2026-03 conditional novelty 7.0 of 10

    LPNSR derives optimal intermediate noise for diffusion SR via MLE and implements it with an LR-guided noise predictor, reaching SOTA perceptual quality in 4 steps without text priors.

  7. LooseRoPE: Content-aware Attention Manipulation for Semantic Harmonization

    cs.GR 2026-01 unverdicted novelty 7.0 of 10

    LooseRoPE modulates RoPE in diffusion attention maps to continuously trade off between preserving a pasted object's identity and harmonizing it with its new surroundings.

  8. SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

    cs.CV 2021-08 conditional novelty 7.0 of 10

    SDEdit performs guided image synthesis and editing by adding noise to inputs and refining them via denoising with a diffusion model's SDE prior, outperforming GAN methods in human studies without task-specific training.

  9. Conditional Diffusion Under Linear Constraints: Langevin Mixing and Information-Theoretic Guarantees

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Error in approximating the tangent conditional score by the unconditional score in diffusion models is bounded by dimension-free conditional mutual information, with a projected-Langevin method outperforming baselines...

  10. Cross-Modal Generation: From Commodity WiFi to High-Fidelity mmWave and RFID Sensing

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    RF-CMG synthesizes high-quality mmWave and RFID signals from WiFi using a diffusion model with Modality-Guided Embedding for high-frequency details and Low-Frequency Modality Consistency to preserve physical structure.

  11. StructDiff: A Structure-Preserving and Spatially Controllable Diffusion Model for Single-Image Generation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    StructDiff adds adaptive receptive fields and 3D positional encoding to a single-scale diffusion model to preserve structure and enable spatial control in single-image generation.

  12. MENO: MeanFlow-Enhanced Neural Operators for Dynamical Systems

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    MENO enhances neural operators with MeanFlow to restore multi-scale accuracy in dynamical system predictions while keeping inference costs low, achieving up to 2x better power spectrum accuracy and 12x faster inferenc...

  13. Analyzing and Guiding Zero-Shot Posterior Sampling in Diffusion Models

    cs.LG 2026-02 unverdicted novelty 6.0 of 10

    Under a Gaussian prior assumption, zero-shot diffusion posterior samplers for inverse problems admit closed-form spectral representations that enable a new parameter-selection framework balancing perceptual quality an...

  14. RubricRL: Simple Generalizable Rewards for Text-to-Image Generation

    cs.CV 2025-11 conditional novelty 6.0 of 10

    Using an LLM to generate prompt-specific visual rubrics and grade each criterion independently gives a more interpretable reward that improves text-to-image model alignment beyond composite and learned scalar rewards.

  15. Regional climate risk assessment from climate models using probabilistic machine learning

    cs.LG 2024-12 unverdicted novelty 6.0 of 10

    GenFocal uses probabilistic ML to downscale coarse climate projections to fine-scale weather events without paired training data and samples rare high-impact events more accurately than prior methods.

  16. DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory

    cs.CV 2023-08 unverdicted novelty 6.0 of 10

    DragNUWA integrates text, image, and trajectory controls into a diffusion video model using a Trajectory Sampler, Multiscale Fusion, and Adaptive Training to enable fine-grained open-domain video generation.

  17. FlowForm: Synergizing Fluid Physics with Topological Consistency for Satellite Flood Synthesis

    cs.CV 2026-08 conditional novelty 5.0 of 10

    FlowForm is a diffusion model that adds shallow-water-equation penalties and terrain-conditioned adapters to synthesize flood satellite images, and reports top scores on a new 10,000-pair dataset.

  18. Provable diffusion-based posterior sampling for linear inverse problems via DDIM

    cs.LG 2026-07 reject novelty 5.0 of 10

    A SVD-based, coordinate-wise DDIM sampler is claimed to asymptotically sample from the posterior for noisy linear inverse problems, but the proof's posterior identification step does not follow from the stated updates.

  19. TopoStyle: Supporting Iterative Design with Generative AI for 2.5D Topology Optimization

    cs.HC 2026-04 unverdicted novelty 5.0 of 10

    TopoStyle provides an interactive system using 2D diffusion models for 2.5D topology optimization that supports hand-drawn and point-based edits plus masking to enable iterative customization balancing performance and...

  20. SocialMirror: Reconstructing 3D Human Interaction Behaviors from Monocular Videos with Semantic and Geometric Guidance

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    SocialMirror reconstructs 3D meshes of closely interacting humans from monocular videos using semantic guidance from vision-language models and geometric constraints in a diffusion model to handle occlusions and maint...

  21. MENO: MeanFlow-Enhanced Neural Operators for Dynamical Systems

    cs.LG 2026-04 conditional novelty 5.0 of 10

    MENO restores multi-scale structure in neural-operator PDE surrogates via one-step improved MeanFlow, claiming up to 2× better power-spectrum accuracy and up to 14× faster inference than DDIM enhancement.

  22. Translationese as a Rational Response to Translation Task Difficulty

    cs.CL 2026-03 unverdicted novelty 5.0 of 10

    Translationese is partly predictable from quantifiable translation-task difficulty, especially cross-lingual transfer load, more so for English-to-German than the reverse.

  23. Reconstructing Multi-Scale Physical Fields from Extremely Sparse Measurements with an Autoencoder-Diffusion Cascade

    cs.LG 2025-12 conditional novelty 5.0 of 10

    A cascade of a functional autoencoder (coarse structure) and a residual conditional diffusion model (fine details), with mask-cascade training and manifold-constrained gradients, reconstructs sparse-sensed physical fields.

  24. FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution

    cs.CV 2025-09 conditional novelty 5.0 of 10

    FS-Diff is a diffusion model that jointly fuses and super-resolves low-resolution multimodal image pairs using clarity-aware CLIP semantics and a bidirectional Mamba feature extractor.

  25. Dual Ascent Diffusion for Inverse Problems

    cs.CV 2025-05 unverdicted novelty 5.0 of 10

    A dual ascent optimization framework is introduced for MAP estimation with diffusion priors, claimed to outperform prior methods on image restoration in quality, noise robustness, speed, and data fidelity.

  26. SOWing Information: Cultivating Contextual Coherence with MLLMs in Image Generation

    cs.CV 2024-11 unverdicted novelty 5.0 of 10

    SOW uses MLLMs and attention to selectively control unidirectional diffusion for pixel-level fidelity and contextual coherence in text-vision-to-image tasks.

  27. Breaking the Lock-in: Diversifying Text-to-Image Generation via Representation Modulation

    cs.CV 2026-06 unverdicted novelty 4.0 of 10

    Early DC component convergence in text-to-image Transformer features causes output homogeneity; selective early attenuation via DAVE improves diversity without retraining or extra cost.

Pith tools