REVIEW 19 cited by
Not All Noises Are Created Equally:Diffusion Noise Selection and Optimization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Diffusion models that can generate high-quality data from randomly sampled Gaussian noises have become the mainstream generative method in both academia and industry. Are randomly sampled Gaussian noises equally good for diffusion models? While a large body of works tried to understand and improve diffusion models, previous works overlooked the possibility to select or optimize the sampled noise the possibility of selecting or optimizing sampled noises for improving diffusion models. In this paper, we mainly made three contributions. First, we report that not all noises are created equally for diffusion models. We are the first to hypothesize and empirically observe that the generation quality of diffusion models significantly depend on the noise inversion stability. This naturally provides us a noise selection method according to the inversion stability. Second, we further propose a novel noise optimization method that actively enhances the inversion stability of arbitrary given noises. Our method is the first one that works on noise space to generally improve generated results without fine-tuning diffusion models. Third, our extensive experiments demonstrate that the proposed noise selection and noise optimization methods both significantly improve representative diffusion models, such as SDXL and SDXL-turbo, in terms of human preference and other objective evaluation metrics. For example, the human preference winning rates of noise selection and noise optimization over the baselines can be up to 57% and 72.5%, respectively, on DrawBench.
Forward citations
Cited by 19 Pith papers
-
UniNDM: A Unified Noise-driven Detection and Mitigation Framework Against Sexual Content in Text-to-Image Generation
UniNDM detects sexual intent from early-stage diffusion noise and mitigates it via LLM-generated negative prompts and initial-noise optimization, across U-Net and DiT models.
-
HyperNet-Adaptation for Diffusion-Based Test Case Generation
HyNeA adapts a diffusion model's hypernetwork per test case to generate realistic, failure-inducing inputs for deep learning systems without curated failure data.
-
LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion
LatSearch improves video diffusion quality and efficiency by scoring intermediate latents with a trained reward model and performing reward-guided resampling plus final pruning.
-
ReGuidance: A Simple Diffusion Wrapper for Boosting Sample Quality on Hard Inverse Problems
A two-step wrapper (invert candidate to latent, then run DPS from that latent) improves hard inpainting results, with mixed or negative superresolution results and toy-model theory.
-
Test-Time Scaling of Diffusion Models via Noise Trajectory Search
An epsilon-greedy search over per-step noise trajectories improves proxy rewards in diffusion image generation without retraining.
-
A Minimalist Method for Fine-tuning Text-to-Image Diffusion Models
A one-step RL method learns a prompt-conditioned initial noise distribution for a frozen diffusion model, improving scores on the training reward models, with the largest gains at low inference steps.
-
Improving Compositional Generation with Diffusion Models Using Lift Scores
CompLift accepts or rejects generated samples by computing lift scores from conditional and unconditional denoising errors, improving compositional alignment without retraining.
-
Not All Parameters Matter: Masking Diffusion Models for Enhancing Generation Ability
MaskUNet masks U-Net weights with a timestep- and sample-dependent binary mask, improving zero-shot FID on COCO by about 1.1 to 1.5 points while leaving pre-trained weights frozen.
-
From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
ReflectionFlow iteratively refines FLUX.1-dev images with verifier-generated feedback, raising GenEval accuracy from 0.85 to 0.91 at 32 samples per prompt.
-
Zigzag Diffusion Sampling: Diffusion Models Can Self-Improve via Self-Reflection
Z-Sampling alternates high-guidance denoising and low-guidance inversion at each step to improve prompt alignment in pretrained text-to-image diffusion models.
-
A Noise is Worth Diffusion Guidance
A one-step learned noise refinement replaces classifier-free guidance at inference on Stable Diffusion 2.1, giving comparable image quality at about 1.7x lower cost.
-
Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps
Aligning cross-attention map similarity to text self-attention maps at test time improves semantic alignment in Stable Diffusion for prompts with multiple objects and attributes.
-
Bag of Design Choices for Inference of High-Resolution Masked Generative Transformer
A set of inference-time design choices improves image quality, memory, and speed of masked generative Transformers, with combined tricks winning about 70% of human-preference comparisons against vanilla sampling.
-
Anchoring and Steering Diffusion: Enhancing the Faithfulness of Text-to-Image Generation at Inference Time
AnchorSteer improves text-to-image faithfulness by anchoring initial noise with CLIP/DAS-derived semantics (LP-SDS) and correcting errors during denoising with a VLM-driven Think-Erase-Retouch loop.
-
Inference-Time Scaling of Diffusion Models via Progressive Seed Pruning
Progressive Seed Pruning—start many noise seeds, score early partially-denoised images, prune aggressively—improves prompt-aligned image generation at fixed denoising compute over best-of-N, resampling, and tree-searc...
-
Navigating the Exploration-Exploitation Tradeoff in Inference-Time Scaling of Diffusion Models
A funnel-shaped particle schedule and an adaptive temperature schedule improve SMC-based inference-time scaling for text-to-image diffusion models at fixed compute.
-
SimDiffRec: Semantic Similarity-Guided Diffusion for Contrastive Sequential Recommendation
SimDiffRec augments user sequences by replacing items at high-confidence diffusion positions with the model's top prediction and using averaged similar-item embeddings as noise, reporting consistent but unverified gai...
-
A Simple and Efficient Baseline for Zero-Shot Generative Classification
GDC classifies images by fitting one Gaussian per class to DINOv2 embeddings of diffusion-generated reference images, reaching 71.4% on ImageNet at 0.03 seconds per image.
-
Efficient Semantic Splatting for Remote Sensing Multi-view Segmentation
A Gaussian splatting framework with per-point semantic features, SAM2 boundary pseudo-labels, and two aggregation losses gives fast, view-consistent multi-view segmentation for remote sensing under sparse labels.
Discussion (0). Continue with ORCID to comment.