Pith. sign in

REVIEW 2 cited by

Dreamguider: Improved Training free Diffusion-based Conditional Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.02549 v1 pith:OYH2WMTZ submitted 2024-06-04 cs.CV

classification cs.CV
keywords guidancediffusiondreamguiderinference-timebackpropagationcompute-heavyconditionalfurther
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models have emerged as a formidable tool for training-free conditional generation.However, a key hurdle in inference-time guidance techniques is the need for compute-heavy backpropagation through the diffusion network for estimating the guidance direction. Moreover, these techniques often require handcrafted parameter tuning on a case-by-case basis. Although some recent works have introduced minimal compute methods for linear inverse problems, a generic lightweight guidance solution to both linear and non-linear guidance problems is still missing. To this end, we propose Dreamguider, a method that enables inference-time guidance without compute-heavy backpropagation through the diffusion network. The key idea is to regulate the gradient flow through a time-varying factor. Moreover, we propose an empirical guidance scale that works for a wide variety of tasks, hence removing the need for handcrafted parameter tuning. We further introduce an effective lightweight augmentation strategy that significantly boosts the performance during inference-time guidance. We present experiments using Dreamguider on multiple tasks across multiple datasets and models to show the effectiveness of the proposed modules. To facilitate further research, we will make the code public after the review process.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Typography inserted into input images can manipulate CLIP-guided image generation models to produce harmful or biased content, and existing text-focused defenses do not catch it.

  2. DyMO: Training-Free Diffusion Model Alignment with Dynamic Multi-Objective Scheduling

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DyMO improves text-to-image outputs at inference time by dynamically scheduling an LLM-built semantic attention objective with a human-preference reward, without retraining the diffusion model.

Pith tools