REVIEW 7 cited by
Smoothed Energy Guidance: Guiding Diffusion Models with Reduced Energy Curvature of Attention
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Conditional diffusion models have shown remarkable success in visual content generation, producing high-quality samples across various domains, largely due to classifier-free guidance (CFG). Recent attempts to extend guidance to unconditional models have relied on heuristic techniques, resulting in suboptimal generation quality and unintended effects. In this work, we propose Smoothed Energy Guidance (SEG), a novel training- and condition-free approach that leverages the energy-based perspective of the self-attention mechanism to enhance image generation. By defining the energy of self-attention, we introduce a method to reduce the curvature of the energy landscape of attention and use the output as the unconditional prediction. Practically, we control the curvature of the energy landscape by adjusting the Gaussian kernel parameter while keeping the guidance scale parameter fixed. Additionally, we present a query blurring method that is equivalent to blurring the entire attention weights without incurring quadratic complexity in the number of tokens. In our experiments, SEG achieves a Pareto improvement in both quality and the reduction of side effects. The code is available at https://github.com/SusungHong/SEG-SDXL.
Forward citations
Cited by 7 Pith papers
-
Zigzag Diffusion Sampling: Diffusion Models Can Self-Improve via Self-Reflection
Z-Sampling alternates high-guidance denoising and low-guidance inversion at each step to improve prompt alignment in pretrained text-to-image diffusion models.
-
Perturb-and-Revise: Flexible 3D Editing with Generative Trajectories
Perturb-and-Revise edits 3D scenes by mixing a NeRF's trained parameters with random ones, running multi-view score distillation toward the edit prompt, and refining with identity-preserving gradients.
-
A Noise is Worth Diffusion Guidance
A one-step learned noise refinement replaces classifier-free guidance at inference on Stable Diffusion 2.1, giving comparable image quality at about 1.7x lower cost.
-
Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling
STG boosts video diffusion sample quality by guiding away from a self-produced weak model obtained by skipping spatiotemporal layers, with no extra training.
-
Gradient-Free Classifier Guidance for Diffusion Model Sampling
A gradient-free diffusion sampler that selects a reference class from a pretrained classifier's predictions and adapts guidance strength improves class-conditional fidelity, but its Precision gains are measured with t...
-
Guiding a diffusion model using sliding windows
Masked sliding window guidance improves diffusion sample quality by guiding the model with its own crop-based predictions, without training.
-
Improving Diffusion-Based Image Editing Faithfulness via Guidance and Scheduling
FGS improves faithfulness in diffusion-based image editing by adding a perturbed-feature guidance term and a logarithmic schedule over denoising timesteps.
Discussion (0). Continue with ORCID to comment.