Pith. sign in

REVIEW 45 cited by

Adversarial Diffusion Distillation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.17042 v1 pith:V4NPFM3B submitted 2023-11-28 cs.CV

classification cs.CV
keywords diffusionimagemodelsadversarialdistillationstepshighhttps
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Adversarial Diffusion Distillation (ADD), a novel training approach that efficiently samples large-scale foundational image diffusion models in just 1-4 steps while maintaining high image quality. We use score distillation to leverage large-scale off-the-shelf image diffusion models as a teacher signal in combination with an adversarial loss to ensure high image fidelity even in the low-step regime of one or two sampling steps. Our analyses show that our model clearly outperforms existing few-step methods (GANs, Latent Consistency Models) in a single step and reaches the performance of state-of-the-art diffusion models (SDXL) in only four steps. ADD is the first method to unlock single-step, real-time image synthesis with foundation models. Code and weights available under https://github.com/Stability-AI/generative-models and https://huggingface.co/stabilityai/ .

Discussion (0). Sign in to comment.

Forward citations

Cited by 45 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. NEvo: Neural-Guided Evolutionary Video Synthesis for Dynamic Visual Selectivity

    cs.CV 2026-07 unverdicted novelty 7.0 of 10

    NEvo performs evolutionary search guided by a dynamic voxel-level encoding model to synthesize videos that maximize predicted activity in target brain ROIs, recovering known selectivities and revealing temporal dynami...

  2. Language-Assisted Super-Resolution from Real-World Low-Resolution Patches

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    LA-SR redefines unpaired super-resolution in language space by projecting images into a semantically rich representation and applying vision-language model guided losses to handle real-world degradations extracted fro...

  3. Oracle Noise: Faster Semantic Spherical Alignment for Interpretable Latent Optimization

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    Oracle Noise optimizes diffusion model noise on a Riemannian hypersphere guided by key prompt words to preserve the Gaussian prior, eliminate norm inflation, and achieve faster semantic alignment than Euclidean methods.

  4. $Z^2$-Sampling: Zero-Cost Zigzag Trajectories for Semantic Alignment in Diffusion Models

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    Z²-Sampling implicitly realizes zero-cost zigzag trajectories for curvature-aware semantic alignment in diffusion models by reducing multi-step paths via operator dualities and temporal caching while synthesizing a di...

  5. PromptEvolver: Prompt Inversion through Evolutionary Optimization in Natural-Language Space

    cs.LG 2026-04 unverdicted novelty 7.0 of 10

    PromptEvolver recovers high-fidelity natural language prompts for given images by evolving them via genetic algorithm guided by a vision-language model, outperforming prior methods on benchmarks.

  6. It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models

    cs.CV 2025-12 unverdicted novelty 7.0 of 10

    Noise optimization during sampling recovers diversity in mode-collapsed diffusion models while preserving output fidelity.

  7. One Step Diffusion via Shortcut Models

    cs.LG 2024-10 conditional novelty 7.0 of 10

    Shortcut models enable high-quality single or few-step sampling in diffusion models with one network and training phase by conditioning on desired step size.

  8. A Decomposable Probe for Few-Step Diffusion Models: Prompt, Latent, and Score Selectivity across Backbone Families and Distillation Paradigms

    cs.CV 2026-07 conditional novelty 6.5 of 10

    A three-layer perturbation probe shows latent selectivity is a near-binary rectified-flow fingerprint that survives ADD distillation, while score selectivity tracks distillation objective across 23 T2I models.

  9. Diagnosing Under-Development of Irreversible Processes in Video Generation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Text-to-video generators under-develop irreversible processes: near-zero progress and 92-100% stasis versus real footage (rho=+0.40, 35% static), confirmed by human raters.

  10. Parallel Decoding Distillation for Fast Image and Video Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A trajectory-based distillation method trains a student to predict multiple mean velocities per network evaluation, enabling 4-8 step generation with competitive quality and improved diversity.

  11. DiTango: Cost-Effective Parallel Diffusion Generation with Selective Attention State Reuse

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Selective reuse of composed attention states across denoising steps lets DiTango skip both remote KV communication and attention compute for low-contribution sequence partitions, cutting multi-GPU diffusion latency by...

  12. DiffusionBench: On Holistic Evaluation of Diffusion Transformers

    cs.CV 2026-06 conditional novelty 6.0 of 10

    NanoGen unifies DiT training on ImageNet and T2I, reveals negative Pearson correlations (-0.377 to -0.580) in method rankings across metrics from 21 models, and motivates DiffusionBench for holistic evaluation.

  13. Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    CACFM applies RL to adaptively select critical regions in probability flow ODE trajectories for consistency distillation, yielding SOTA few-step results on FLUX and SDXL.

  14. Addressing Detail Bottlenecks in Latent Diffusion for RGB-to-SWIR Image Translation

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Introduces SCAE with skip connections and LGE to fix detail loss in LDMs for RGB-to-SWIR translation, yielding up to 2x mAP gains and 3.4x on small objects while reaching SOTA FID.

  15. Variance Reduction for Expectations with Diffusion Teachers

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    CARV amortizes upstream diffusion teacher costs over noise resamples with timestep importance sampling and stratified-inverse-CDF sampling, delivering 2-3x effective compute gains in text-to-3D experiments and order-o...

  16. Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    CLVR couples verified logical planning with pixel diffusion, uses proxy reinforcement learning on distilled histories, and merges weights to cut inference to 4 NFEs while outperforming open-source T2I models on comple...

  17. Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    CLVR framework adds closed-loop visual verification, proxy prompt reinforcement learning, and delta-space weight merge to improve complex text-to-image generation over single-step or unverified multi-step baselines.

  18. MetaSR: Content-Adaptive Metadata Orchestration for Generative Super-Resolution

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    MetaSR adaptively orchestrates metadata in a DiT-based generative SR model to deliver up to 1 dB PSNR gains and 50% bitrate savings across diverse content and degradations.

  19. BiasIG: Benchmarking Multi-dimensional Social Biases in Text-to-Image Models

    cs.CY 2026-04 conditional novelty 6.0 of 10

    BiasIG is a multi-dimensional benchmark for social biases in T2I models that shows debiasing interventions frequently cause confounding discrimination effects.

  20. Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion Models

    cs.CV 2026-03 unverdicted novelty 6.0 of 10

    Implicit generative choices in diffusion models for ambiguous prompts are localized principally in self-attention layers, enabling a targeted ICM steering method that outperforms prior debiasing approaches.

  21. Na\"ive PAINE: Lightweight Text-to-Image Generation Improvement with Prompt Evaluation

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A lightweight predictor ranks initial noises by expected human-preference score for a prompt, selecting the best few for diffusion generation and reporting prompt difficulty.

  22. Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Category-level retrieval improves when text queries are converted into multiple generated images, aggregated with a learned attention module, and fused with CLIP text similarity.

  23. FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

    cs.GR 2025-06 unverdicted novelty 6.0 of 10

    FLUX.1 Kontext unifies image generation and editing via flow matching and sequence concatenation, delivering improved multi-turn consistency and speed on the new KontextBench benchmark.

  24. Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

    cs.CV 2024-10 unverdicted novelty 6.0 of 10

    Aligning noisy hidden states in diffusion transformers to clean features from pretrained visual encoders speeds up training over 17x and reaches FID 1.42.

  25. Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

    cs.CV 2024-05 conditional novelty 6.0 of 10

    Hunyuan-DiT is a new multi-resolution diffusion transformer that achieves state-of-the-art Chinese text-to-image generation through custom architecture, data pipelines, and multimodal caption refinement.

  26. Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

    cs.CV 2024-03 conditional novelty 6.0 of 10

    Biased noise sampling for rectified flows combined with a bidirectional text-image transformer architecture yields state-of-the-art high-resolution text-to-image results that scale predictably with model size.

  27. DiFaReli++: Diffusion Face Relighting with Consistent Cast Shadows

    cs.CV 2023-04 unverdicted novelty 6.0 of 10

    DiFaReli++ conditions a DDIM on shading references and inferred shadow maps to relight single-view faces with consistent shadows, trained only on 2D images and claiming SOTA on Multi-PIE.

  28. Examining the Efficacy of Graph Neural Network Message-Passing in Regression Contexts

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Across four NAS/DNN-predictor regression benchmarks, GEN (deep graph convolution) achieves the best average rank over 11 GNN message-passing layers, though attention GATv2 wins on the largest graphs.

  29. MicroZoom: Structure-Preserving Detail Synthesis at Extreme Scale

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A cascaded, segmentation-conditioned, per-instance diffusion method synthesizes globally coherent gigapixel microscopic detail from a phone photo and sparse microscope references at up to 350×.

  30. Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A compact 4B image generation/editing system with a fast one-step VAE, native-resolution packing, RL alignment, and 4-step distillation reports competitive benchmarks against 6B–80B open models.

  31. Language-Assisted Super-Resolution from Real-World Low-Resolution Patches

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    LA-SR extracts real LR patches from depth-varying regions in single images and uses vision-language models with linguistic content and quality losses for unpaired super-resolution.

  32. Scenario-conditioned flow matching for probabilistic generation of three-component ground-motion waveforms

    physics.geo-ph 2026-06 unverdicted novelty 5.0 of 10

    WaveFlowGMM generates scenario-conditioned three-component ground-motion waveforms by using symbolic learning for PGA amplitude and AlphaFlow for normalized wavelet-packet waveforms that are later rescaled.

  33. Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning

    cs.CV 2026-06 conditional novelty 5.0 of 10

    A reinforcement-learning agent that picks high-error trajectory segments during consistency distillation improves few-step text-to-image generation on FLUX and SDXL.

  34. On the Redundancy of Timestep Embeddings in Diffusion Models

    cs.LG 2026-06 conditional novelty 5.0 of 10

    Under high-dimensional concentration conditions, the diffusion denoising objective admits the same global minimizer without timestep embeddings, and time-agnostic U-Nets/DiTs empirically match or improve FID on CelebA...

  35. EPIG: Emotion-Based Prompting for Personalised Image Generation

    cs.AI 2026-06 unverdicted novelty 5.0 of 10

    EPIG is a training-free prompt enrichment technique using valence-arousal representations that reduces mean arousal error by 14% versus naive insertion and 12% versus LLM expansion on a 10-prompt benchmark while prese...

  36. LUCID: Learning Unified Control for Image Deflaring and Exposure Mastery in Nighttime Photography

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    LUCID introduces a unified controllable framework for nighttime image restoration that disentangles flares and uses diffusion priors with four-mode training for selective exposure and artifact control.

  37. Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    Lens is a 3.8B-parameter text-to-image model that reaches competitive or superior performance to >6B-parameter systems using 19.3% of the training compute of Z-Image through a densely captioned 800M dataset, multi-res...

  38. Variance Reduction for Expectations with Diffusion Teachers

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    CARV introduces a hierarchical Monte Carlo estimator with amortized reuse, importance sampling, and stratification that yields 2-3x effective compute gains on diffusion-teacher pipelines while cutting gradient varianc...

  39. MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    MONET is an open 104.9M image-text pair dataset created via safety filtering, deduplication, and multi-VLM recaptioning from 2.9B raw pairs, validated by training a competitive 4B-parameter latent diffusion model.

  40. Brain-to-Image Retrieval and Reconstruction via Multimodal EEG Alignment

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    A multimodal alignment pipeline decodes EEG signals recorded during natural image viewing into image retrieval (86.3% Top-1) and reconstruction (CLIP 0.903) tasks.

  41. Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion Models

    cs.CV 2026-03 unverdicted novelty 5.0 of 10

    Implicit generative choices in diffusion models concentrate in self-attention layers; targeted ICM interventions there outperform broader debiasing methods with fewer artifacts.

  42. Systematic Optimization of Real-Time Diffusion Model Inference on Apple M3 Ultra

    cs.LG 2026-02 conditional novelty 5.0 of 10

    Systematic benchmarking of diffusion model optimizations on Apple M3 Ultra produces 22.7 FPS real-time img2img at 512x512 and demonstrates that CUDA-derived techniques do not transfer directly to Apple Silicon.

  43. On the Redundancy of Timestep Embeddings in Diffusion Models

    cs.LG 2026-06 unverdicted novelty 4.0 of 10

    Timestep embeddings are redundant in diffusion models under certain conditions, with time-agnostic variants matching or exceeding conditioned models on FID, precision, and recall for CelebA and CIFAR-10.

  44. Learning from Limited and Imperfect Data

    cs.LG 2025-07 unverdicted novelty 3.0 of 10

    A doctoral thesis compiling nine peer-reviewed papers on long-tailed image generation, long-tailed recognition, semi-supervised learning, and domain adaptation.

  45. Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

    cs.CV 2024-02 unverdicted novelty 2.0 of 10

    The paper reviews the background, technology, applications, limitations, and future directions of OpenAI's Sora text-to-video generative model based on public information.

Pith tools