Pith. sign in

REVIEW 4 cited by

Multistep Distillation of Diffusion Models via Moment Matching

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.04103 v1 pith:LJQKASFU submitted 2024-06-06 cs.LG cs.AIcs.CVcs.NE

classification cs.LGcs.AIcs.CVcs.NE
keywords modelsdiffusionmatchingdatamany-stepmethodmomentone-step
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present a new method for making diffusion models faster to sample. The method distills many-step diffusion models into few-step models by matching conditional expectations of the clean data given noisy data along the sampling trajectory. Our approach extends recently proposed one-step methods to the multi-step case, and provides a new perspective by interpreting these approaches in terms of moment matching. By using up to 8 sampling steps, we obtain distilled models that outperform not only their one-step versions but also their original many-step teacher models, obtaining new state-of-the-art results on the Imagenet dataset. We also show promising results on a large text-to-image model where we achieve fast generation of high resolution images directly in image space, without needing autoencoders or upsamplers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fast Video Generation with Sliding Tile Attention

    cs.CV 2025-02 conditional novelty 7.0 of 10

    Sliding tile attention (STA) replaces full 3D attention in video diffusion transformers with dense tile-local windows, achieving 1.89x training-free and up to 3.53x fine-tuned end-to-end speedups on HunyuanVideo with ...

  2. Diffusion Autoencoders are Scalable Image Tokenizers

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A single diffusion L2 loss can train scalable image tokenizers that match or outperform GAN-LPIPS tokenizers for reconstruction and downstream generation.

  3. VarDiU: A Variational Diffusive Upper Bound for One-Step Diffusion Distillation

    cs.LG 2025-08 conditional novelty 5.0 of 10

    VarDiU optimizes a variational upper bound on the diffusive KL divergence with an unbiased gradient estimator, improving one-step generation on a 2D 40-Gaussian toy benchmark compared with Diff-Instruct.

  4. Reinforcement Learning: From Algorithms To Foundation Models

    cs.AI 2026-07 conditional novelty 3.0 of 10

    A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.

Pith tools