Pith. sign in

REVIEW 1 cited by

Unsupervised Discovery of Interpretable Directions in h-space of Pre-trained Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.09912 v3 pith:LV7LNI3U submitted 2023-10-15 cs.CV cs.LG

classification cs.CVcs.LG
keywords diffusiondirectionsmodelsmethodgradienth-spaceinterpretablepre-trained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose the first unsupervised and learning-based method to identify interpretable directions in h-space of pre-trained diffusion models. Our method is derived from an existing technique that operates on the GAN latent space. Specifically, we employ a shift control module that works on h-space of pre-trained diffusion models to manipulate a sample into a shifted version of itself, followed by a reconstructor to reproduce both the type and the strength of the manipulation. By jointly optimizing them, the model will spontaneously discover disentangled and interpretable directions. To prevent the discovery of meaningless and destructive directions, we employ a discriminator to maintain the fidelity of shifted sample. Due to the iterative generative process of diffusion models, our training requires a substantial amount of GPU VRAM to store numerous intermediate tensors for back-propagating gradient. To address this issue, we propose a general VRAM-efficient training algorithm based on gradient checkpointing technique to back-propagate any gradient through the whole generative process, with acceptable occupancy of VRAM and sacrifice of training efficiency. Compared with existing related works on diffusion models, our method inherently identifies global and scalable directions, without necessitating any other complicated procedures. Extensive experiments on various datasets demonstrate the effectiveness of our method.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On Designing Diffusion Autoencoders for Efficient Generation and Representation Learning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Small binary latents conditioned via cross-attention let a diffusion autoencoder generate from a uniform Bernoulli prior with fewer steps while keeping representation quality.

Pith tools