Pith. sign in

REVIEW 6 cited by

Unifying Diffusion Models' Latent Space, with Applications to CycleDiffusion and Guidance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.05559 v2 pith:5TYCXHRV submitted 2022-10-11 cs.CV cs.GRcs.LG

Unifying Diffusion Models' Latent Space, with Applications to CycleDiffusion and Guidance

classification cs.CV cs.GRcs.LG
keywords modelsdiffusionlatentspaceformulationcyclediffusionganscode
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Diffusion models have achieved unprecedented performance in generative modeling. The commonly-adopted formulation of the latent code of diffusion models is a sequence of gradually denoised samples, as opposed to the simpler (e.g., Gaussian) latent space of GANs, VAEs, and normalizing flows. This paper provides an alternative, Gaussian formulation of the latent space of various diffusion models, as well as an invertible DPM-Encoder that maps images into the latent space. While our formulation is purely based on the definition of diffusion models, we demonstrate several intriguing consequences. (1) Empirically, we observe that a common latent space emerges from two diffusion models trained independently on related domains. In light of this finding, we propose CycleDiffusion, which uses DPM-Encoder for unpaired image-to-image translation. Furthermore, applying CycleDiffusion to text-to-image diffusion models, we show that large-scale text-to-image diffusion models can be used as zero-shot image-to-image editors. (2) One can guide pre-trained diffusion models and GANs by controlling the latent codes in a unified, plug-and-play formulation based on energy-based models. Using the CLIP model and a face recognition model as guidance, we demonstrate that diffusion models have better coverage of low-density sub-populations and individuals than GANs. The code is publicly available at https://github.com/ChenWu98/cycle-diffusion.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Semantic Browsing: Controllable Diversity for Image Generation

    cs.CV 2026-06 unverdicted novelty 7.0

    A technique for controllable diversity in text-to-image generation by inducing structured semantic variations at the prompt level via VLM and agentic workflow.

  2. IMPLICITSTAINER: Resolution Agnostic Data-Efficient Virtual Staining Using Neural Implicit Functions

    eess.IV 2025-05 unverdicted novelty 7.0

    Neural implicit functions enable resolution-agnostic, deterministic virtual staining from H&E to IHC images with SOTA results and better low-data performance than patch-based GAN or diffusion methods.

  3. Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing

    cs.CV 2025-09 conditional novelty 6.0

    VARIN uses a Location-aware Argmax Inversion pseudo-inverse of Gumbel-max sampling to extract editable discrete noises, enabling training-free prompt-guided editing for visual autoregressive models.

  4. Regional climate risk assessment from climate models using probabilistic machine learning

    cs.LG 2024-12 unverdicted novelty 6.0

    GenFocal uses probabilistic ML to downscale coarse climate projections to fine-scale weather events without paired training data and samples rare high-impact events more accurately than prior methods.

  5. FlowForm: Synergizing Fluid Physics with Topological Consistency for Satellite Flood Synthesis

    cs.CV 2026-08 conditional novelty 5.0

    FlowForm is a diffusion model that adds shallow-water-equation penalties and terrain-conditioned adapters to synthesize flood satellite images, and reports top scores on a new 10,000-pair dataset.

  6. Style-Based Neural Architectures for Real-Time Weather Classification

    cs.CV 2026-04 unverdicted novelty 5.0

    Three style-based neural architectures are proposed for real-time weather classification from images, with two truncated ResNet variants claimed to outperform prior methods and generalize across public datasets.