Pith. sign in

REVIEW 31 cited by

Diffusion Models already have a Semantic Latent Space

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.10960 v2 pith:ZITBY3TX submitted 2022-10-20 cs.CV

classification cs.CV
keywords semanticlatentspacediffusiongenerativemodelsprocessasyrp
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Diffusion models achieve outstanding generative performance in various domains. Despite their great success, they lack semantic latent space which is essential for controlling the generative process. To address the problem, we propose asymmetric reverse process (Asyrp) which discovers the semantic latent space in frozen pretrained diffusion models. Our semantic latent space, named h-space, has nice properties for accommodating semantic image manipulation: homogeneity, linearity, robustness, and consistency across timesteps. In addition, we introduce a principled design of the generative process for versatile editing and quality boost ing by quantifiable measures: editing strength of an interval and quality deficiency at a timestep. Our method is applicable to various architectures (DDPM++, iD- DPM, and ADM) and datasets (CelebA-HQ, AFHQ-dog, LSUN-church, LSUN- bedroom, and METFACES). Project page: https://kwonminki.github.io/Asyrp/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 31 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AutoAnchor: Stable Diffusion Unlearning Using Cross-Attention as a Manifold Surrogate

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Cross-attention maps serve as a tractable surrogate for manifold proximity, enabling automatic synthesis of anchors that suppress normal-space drift in diffusion unlearning.

  2. Filtering Memorization from Parameter-Space in Diffusion Models

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    Base-Anchored Filtering suppresses weakly backbone-aligned LoRA spectral channels to cut memorization while preserving or improving generation quality, without data or re-training.

  3. ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation

    cs.CV 2025-07 conditional novelty 7.0 of 10

    ReFlex edits real images with FLUX by extracting attention and residual features from a mid-step latent and adapting them during generation, improving text alignment while preserving structure.

  4. Instruction-based Image Manipulation by Watching How Things Move

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A diffusion editing model, InstructMove, is trained on video frame pairs annotated by MLLMs using spatial conditioning, enabling non-rigid edits and viewpoint changes.

  5. Unboxing Diffusion Models for the Arts: Interactive Model Bending and Practice-Based Explainability

    cs.HC 2026-07 conditional novelty 6.0 of 10

    Bending different layers of a diffusion model's UNet produces distinct and fairly consistent visual effects across seeds and prompts, and an interactive ComfyUI tool lets artists explore these effects hands-on.

  6. Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing

    cs.CV 2025-10 conditional novelty 6.0 of 10

    Kontinuous Kontext adds continuous edit-strength control to instruction-based image editing by projecting a scalar strength and text embedding into the modulation space of a Flux Kontext diffusion editor.

  7. SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse Autoencoders

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A supervised sparse autoencoder binds each concept to a single neuron, letting Stable Diffusion erase a concept by steering one latent.

  8. Generative Latent Diffusion for Efficient Spatiotemporal Data Reduction

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A latent diffusion model conditioned on keyframe latents reconstructs non-key frames, giving higher compression ratios than prior scientific data compressors.

  9. DiffEx: Explaining a Classifier with Diffusion Models to Identify Microscopic Cellular Variations

    cs.CV 2025-02 conditional novelty 6.0 of 10

    DiffEx builds a classifier-aware latent space with a diffusion model, finds contrastive directions in it, and ranks them to produce visual explanations that reveal cellular phenotype changes.

  10. Training-Free Style and Content Transfer by Leveraging U-Net Skip Connections in Stable Diffusion

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Injecting the fourth and fifth U-Net skip connections from one Stable Diffusion image into another transfers content or style without any training.

  11. CA-Edit: Causality-Aware Condition Adapter for High-Fidelity Local Facial Attribute Editing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    CA-Edit uses a causality-aware condition adapter and low-frequency sampling guidance to make local facial attribute edits follow text prompts while preserving skin detail fidelity.

  12. FluxSpace: Disentangled Semantic Editing in Rectified Flow Transformers

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FluxSpace performs training-free, disentangled semantic editing in rectified flow transformers by combining attention outputs with prompt-derived linear directions.

  13. MyTimeMachine: Personalized Facial Age Transformation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A personalized facial age transformation method that uses an adapter network on top of the SAM global aging model, trained with 10 to 50 photos of one person, to produce re-aged images that resemble that person's actu...

  14. Generating Compositional Scenes via Text-to-image RGBA Instance Generation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A multi-stage text-to-image approach that generates individual objects as RGBA images and composes them scene-by-scene via noise blending, enabling fine-grained layout and attribute control.

  15. Latent Space Disentanglement in Diffusion Transformers Enables Precise Zero-shot Semantic Editing

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Diffusion transformer latent spaces are shown to be semantically disentangled, and prompt-difference directions plus a score-distillation step enable zero-shot fine-grained image editing.

  16. Unified Backbone Refinement for Diffusion Models via Internal-Latent Analysis

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Early abrupt deviations in deep diffusion latents track artifacts; EMA detection plus backbone-specific suppression (DUNE) reduces them without retraining.

  17. Towards Efficient Exemplar Based Image Editing with Multimodal VLMs

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ReEdit transfers exemplar-based edits to new images by conditioning Stable Diffusion on a LLaVA-written caption plus a CLIP edit-direction vector, with no per-example optimization.

  18. ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A released 6.4 million pair dataset and 613 sample benchmark for instruction-guided image editing of non-rigid motions, plus a Flux.1-dev based baseline that outperforms open-source methods on the new benchmark.

  19. Exploring the latent space of diffusion models directly through singular value decomposition

    cs.CV 2025-02 reject novelty 5.0 of 10

    The authors report that singular value decomposition of diffusion latent codes reveals stable, order-mobile attribute directions and propose Attribute Vector Integration, a per-pair MLP-based editor that transfers tex...

  20. StyleStudio: Text-Driven Style Transfer with Selective Control of Style Elements

    cs.CV 2024-12 conditional novelty 5.0 of 10

    StyleStudio improves text-driven style transfer with cross-modal AdaIN, a negative-style-image classifier-free guidance, and teacher-model layout stabilization.

  21. $\textit{Revelio}$: Interpreting and leveraging semantic information in diffusion models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    Diffusion model features, when decoded with k-sparse autoencoders, reveal interpretable visual concepts, and a lightweight classifier on the best layer (up_ft1 at t=25) beats prior diffusion-based classifiers on fine-...

  22. FastGrasp: Efficient Grasp Synthesis with Diffusion

    cs.RO 2024-11 conditional novelty 5.0 of 10

    A one-stage latent diffusion model with an adaptation module generates MANO hand grasping poses from object point clouds faster and with lower penetration than two-stage optimization baselines.

  23. Mediffusion: Joint Diffusion for Self-Explainable Semi-Supervised Classification and Medical Image Generation

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A joint latent diffusion model and classifier, sharing one UNet, improves semi-supervised chest X-ray and skin lesion classification and produces counterfactual explanations evaluated by an external classifier and sev...

  24. AI for Cultural Heritage Textiles: Fine-Tuned Latent Diffusion for Novel Ulos Motif Synthesis

    cs.CV 2026-07 conditional novelty 4.0 of 10

    Fine-tuned Protogen v3.4 generates novel Ulos motifs with ~10.5 imes lower FID and 2× higher IS than Stable Diffusion v1.4, with guidance scale 5–9 balancing fidelity and diversity.

  25. Diffusion Once and Done: Degradation-Aware LoRA for Efficient All-in-One Image Restoration

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    Proposes DOD, a one-step Stable Diffusion model for all-in-one image restoration, but the submitted manuscript text is an unrelated software engineering review, leaving the claim unverifiable.

  26. A Plug-and-Play Multi-Criteria Guidance for Diverse In-Betweening Human Motion Generation

    cs.GR 2025-08 reject novelty 4.0 of 10

    MCG-IMM uses evolutionary multi-criteria optimization during sampling to increase the intra-batch diversity of in-betweening human motions generated by any pretrained model.

  27. SecureT2I: No More Unauthorized Manipulation on AI Generated Images from Prompts

    cs.CR 2025-07 reject novelty 4.0 of 10

    A diffusion editing model is fine-tuned with a blur target for forbidden images and the original output for permitted images, claiming selective suppression of unauthorized edits.

  28. DCI: Dual-Conditional Inversion for Boosting Diffusion-Based Image Editing

    cs.CV 2025-06 reject novelty 4.0 of 10

    DCI combines reference-guided noise correction with fixed-point latent refinement and reports state-of-the-art reconstruction and editing metrics on PIE-Bench.

  29. Unsupervised Region-Based Image Editing of Denoising Diffusion Models

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A masking and Jacobian projection technique discovers unsupervised semantic directions in diffusion model latent space, enabling region-local editing without fine-tuning.

  30. Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A survey maps the field of MLLM explainability and interpretability into data, model, and training and inference perspectives.

  31. Pixel-Space Diffusion Transformers

    cs.CV 2026-07 conditional novelty 3.0 of 10

    A systematic review of pixel-space diffusion transformers, categorizing architectures and challenges for end-to-end image generation without latent compression.

Pith tools