Pith. sign in

REVIEW 20 cited by

Diffusion Models already have a Semantic Latent Space

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.10960 v2 pith:ZITBY3TX submitted 2022-10-20 cs.CV

Diffusion Models already have a Semantic Latent Space

classification cs.CV
keywords semanticlatentspacediffusiongenerativemodelsprocessasyrp
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Diffusion models achieve outstanding generative performance in various domains. Despite their great success, they lack semantic latent space which is essential for controlling the generative process. To address the problem, we propose asymmetric reverse process (Asyrp) which discovers the semantic latent space in frozen pretrained diffusion models. Our semantic latent space, named h-space, has nice properties for accommodating semantic image manipulation: homogeneity, linearity, robustness, and consistency across timesteps. In addition, we introduce a principled design of the generative process for versatile editing and quality boost ing by quantifiable measures: editing strength of an interval and quality deficiency at a timestep. Our method is applicable to various architectures (DDPM++, iD- DPM, and ADM) and datasets (CelebA-HQ, AFHQ-dog, LSUN-church, LSUN- bedroom, and METFACES). Project page: https://kwonminki.github.io/Asyrp/

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AutoAnchor: Stable Diffusion Unlearning Using Cross-Attention as a Manifold Surrogate

    cs.LG 2026-07 conditional novelty 7.0

    Cross-attention maps serve as a tractable surrogate for manifold proximity, enabling automatic synthesis of anchors that suppress normal-space drift in diffusion unlearning.

  2. Towards More General Control of Diffusion Models Using Jeffrey Guidance

    cs.LG 2026-06 unverdicted novelty 7.0

    Jeffrey guidance applies Jeffrey's rule of conditioning to diffusion models to target prescribed marginal distributions while preserving conditional structure, demonstrated via embedding matching and fairness enforcement.

  3. Filtering Memorization from Parameter-Space in Diffusion Models

    cs.CV 2026-05 conditional novelty 7.0

    Base-Anchored Filtering suppresses weakly backbone-aligned LoRA spectral channels to cut memorization while preserving or improving generation quality, without data or re-training.

  4. The Interplay of Data Structure and Imbalance in the Learning Dynamics of Diffusion Models

    stat.ML 2026-05 unverdicted novelty 7.0

    Higher-variance classes are learned first in diffusion models; strong class imbalance reverses the order and imposes distinct delayed learning times on minority classes.

  5. Grokking of Diffusion Models: Case Study on Modular Addition

    cs.LG 2026-04 unverdicted novelty 7.0

    Diffusion models show grokking on modular addition by composing periodic operand representations in simple data regimes or by separating arithmetic computation from visual denoising across timesteps in varied regimes.

  6. Unboxing Diffusion Models for the Arts: Interactive Model Bending and Practice-Based Explainability

    cs.HC 2026-07 conditional novelty 6.0

    Bending different layers of a diffusion model's UNet produces distinct and fairly consistent visual effects across seeds and prompts, and an interactive ComfyUI tool lets artists explore these effects hands-on.

  7. Don't Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance

    cs.CV 2026-06 unverdicted novelty 6.0

    Feature self-guidance disperses internal features of flow models during batch generation and applies manifold regularization to increase output diversity while preserving condition alignment.

  8. AnchorSteer: Self-Discovered Concept Injection for Structure-Preserving Music Editing

    cs.SD 2026-05 unverdicted novelty 6.0

    AnchorSteer couples self-discovered semantic concept vectors with structural anchoring in diffusion models to achieve controllable music editing with preserved structure.

  9. Are Watermarked Images Editable? SafeMark for Watermark-Preserving Text-Guided Image Editing

    cs.CV 2026-05 unverdicted novelty 6.0

    SafeMark integrates a thresholded watermark-decoding loss into diffusion editors to enable text-guided edits that preserve embedded watermarks with high bit accuracy.

  10. Whispers in the Noise: Surrogate-Guided Concept Awakening via a Multi-Agent Framework

    cs.AI 2026-05 unverdicted novelty 6.0

    ConceptAgent is a black-box multi-agent system that awakens erased concepts in diffusion models by initializing denoising trajectories from surrogate-guided noisy states.

  11. Filtering Memorization from Parameter-Space in Diffusion Models

    cs.CV 2026-05 unverdicted novelty 6.0

    BAF reduces memorization in diffusion LoRAs by filtering spectral channels of the adaptation weights that show weak alignment with the base model's principal subspace.

  12. Evaluation without Generation: Non-Generative Assessment of Harmful Model Specialization with Applications to CSAM

    cs.LG 2026-04 unverdicted novelty 6.0

    Gaussian probing infers harmful model specialization from parameter perturbations and internal representation responses to Gaussian latent ensembles rather than from generated outputs.

  13. Geometric Decoupling: Diagnosing the Structural Instability of Latent

    cs.CV 2026-04 unverdicted novelty 6.0

    Latent diffusion models exhibit geometric decoupling where curvature in out-of-distribution generation is misallocated to unstable semantic boundaries instead of image details, identifying geometric hotspots as the st...

  14. Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing

    cs.CV 2025-10 conditional novelty 6.0

    Kontinuous Kontext adds continuous edit-strength control to instruction-based image editing by projecting a scalar strength and text embedding into the modulation space of a Flux Kontext diffusion editor.

  15. SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse Autoencoders

    cs.CV 2025-09 conditional novelty 6.0

    A supervised sparse autoencoder binds each concept to a single neuron, letting Stable Diffusion erase a concept by steering one latent.

  16. Text Slider: Efficient and Plug-and-Play Continuous Concept Control for Image/Video Synthesis via LoRA Adapters

    cs.GR 2025-09 unverdicted novelty 6.0

    Text Slider uses LoRA adapters on pre-trained text encoders to identify low-rank directions for efficient, plug-and-play continuous concept control in diffusion-based image and video synthesis.

  17. Unified Backbone Refinement for Diffusion Models via Internal-Latent Analysis

    cs.CV 2026-07 conditional novelty 5.0

    Early abrupt deviations in deep diffusion latents track artifacts; EMA detection plus backbone-specific suppression (DUNE) reduces them without retraining.

  18. AI for Cultural Heritage Textiles: Fine-Tuned Latent Diffusion for Novel Ulos Motif Synthesis

    cs.CV 2026-07 conditional novelty 4.0

    Fine-tuned Protogen v3.4 generates novel Ulos motifs with ~10.5 imes lower FID and 2× higher IS than Stable Diffusion v1.4, with guidance scale 5–9 balancing fidelity and diversity.

  19. Towards Controllable Image Generation through Representation-Conditioned Diffusion Models

    cs.CV 2026-05 unverdicted novelty 4.0

    Self-supervised representation conditioning on diffusion models boosts unconditional generation quality and yields a controllable space with smooth, disentangled variations.

  20. Pixel-Space Diffusion Transformers

    cs.CV 2026-07 conditional novelty 3.0

    A systematic review of pixel-space diffusion transformers, categorizing architectures and challenges for end-to-end image generation without latent compression.