Pith. sign in

REVIEW 7 cited by

Guiding a Diffusion Model with a Bad Version of Itself

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.02507 v3 pith:IW7ZHA54 submitted 2024-06-04 cs.CV cs.AIcs.LGcs.NEstat.ML

classification cs.CVcs.AIcs.LGcs.NEstat.ML
keywords modeldiffusionqualityunconditionalvariationamountcontrolgeneration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The primary axes of interest in image-generating diffusion models are image quality, the amount of variation in the results, and how well the results align with a given condition, e.g., a class label or a text prompt. The popular classifier-free guidance approach uses an unconditional model to guide a conditional model, leading to simultaneously better prompt alignment and higher-quality images at the cost of reduced variation. These effects seem inherently entangled, and thus hard to control. We make the surprising observation that it is possible to obtain disentangled control over image quality without compromising the amount of variation by guiding generation using a smaller, less-trained version of the model itself rather than an unconditional model. This leads to significant improvements in ImageNet generation, setting record FIDs of 1.01 for 64x64 and 1.25 for 512x512, using publicly available networks. Furthermore, the method is also applicable to unconditional diffusion models, drastically improving their quality.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Decomposable Probe for Few-Step Diffusion Models: Prompt, Latent, and Score Selectivity across Backbone Families and Distillation Paradigms

    cs.CV 2026-07 conditional novelty 6.5 of 10

    A three-layer perturbation probe shows latent selectivity is a near-binary rectified-flow fingerprint that survives ADD distillation, while score selectivity tracks distillation objective across 23 T2I models.

  2. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  3. Diffuse Everything: Multimodal Diffusion Models on Arbitrary State Spaces

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A unified diffusion framework with per-modality noise clocks lets one model generate images, text, and tabular data jointly or conditionally in their native spaces.

  4. Diffusion Counterfactual Generation with Semantic Abduction

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Diffusion-based causal image counterfactuals with semantic abduction improve identity preservation at a small cost in intervention effectiveness, demonstrated on Morpho-MNIST, CelebA-HQ, and mammogram artifact removal.

  5. Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO

    cs.CV 2026-02 conditional novelty 5.0 of 10

    TP-GRPO replaces terminal rewards with per-step incremental rewards and amplifies rewards at sign-flip 'turning points' to capture delayed denoising effects, improving Flow-GRPO on three benchmarks.

  6. Rethinking Visual Autoregressive Sampling with Information-Grounding Guidance

    cs.CV 2025-09 conditional novelty 5.0 of 10

    IGG, an attention-based reweighting of classifier-free guidance, concentrates guidance on important tokens and modestly improves FID/IS in scale-wise autoregressive image generation.

  7. Improving Diffusion-Based Image Editing Faithfulness via Guidance and Scheduling

    cs.CV 2025-06 conditional novelty 5.0 of 10

    FGS improves faithfulness in diffusion-based image editing by adding a perturbed-feature guidance term and a logarithmic schedule over denoising timesteps.

Pith tools