Pith. sign in

REVIEW 3 cited by

EDICT: Exact Diffusion Inversion via Coupled Transformations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.12446 v2 pith:M2XCHJOP submitted 2022-11-22 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords imageedictdiffusioninversionrealimagesnoisecoupled
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Finding an initial noise vector that produces an input image when fed into the diffusion process (known as inversion) is an important problem in denoising diffusion models (DDMs), with applications for real image editing. The state-of-the-art approach for real image editing with inversion uses denoising diffusion implicit models (DDIMs) to deterministically noise the image to the intermediate state along the path that the denoising would follow given the original conditioning. However, DDIM inversion for real images is unstable as it relies on local linearization assumptions, which result in the propagation of errors, leading to incorrect image reconstruction and loss of content. To alleviate these problems, we propose Exact Diffusion Inversion via Coupled Transformations (EDICT), an inversion method that draws inspiration from affine coupling layers. EDICT enables mathematically exact inversion of real and model-generated images by maintaining two coupled noise vectors which are used to invert each other in an alternating fashion. Using Stable Diffusion, a state-of-the-art latent diffusion model, we demonstrate that EDICT successfully reconstructs real images with high fidelity. On complex image datasets like MS-COCO, EDICT reconstruction significantly outperforms DDIM, improving the mean square error of reconstruction by a factor of two. Using noise vectors inverted from real images, EDICT enables a wide range of image edits--from local and global semantic edits to image stylization--while maintaining fidelity to the original image structure. EDICT requires no model training/finetuning, prompt tuning, or extra data and can be combined with any pretrained DDM. Code is available at https://github.com/salesforce/EDICT.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs

    cs.CV 2025-07 conditional novelty 7.0 of 10

    A large human-annotated benchmark of AI-edited images (EBench-18K) plus a fine-tuned LMM metric (LMM4Edit) that predicts human preference scores across three dimensions and answers editing-specific questions.

  2. TITAN-Guide: Taming Inference-Time AligNment for Guided Text-to-Video Diffusion Models

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A training-free guidance method that uses forward gradients instead of backpropagation to steer text-to-video diffusion latents with lower GPU memory.

  3. Defensive Adversarial CAPTCHA: A Semantics-Driven Framework for Natural Adversarial Example Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    DAC and BP-DAC generate 'unsourced' adversarial CAPTCHAs from semantic prompts and report transfer attack success rates above 95% on ImageNet classifiers in black-box settings.

Pith tools