REVIEW 14 cited by
Null-text Inversion for Editing Real Images using Guided Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent text-guided diffusion models provide powerful image generation capabilities. Currently, a massive effort is given to enable the modification of these images using text only as means to offer intuitive and versatile editing. To edit a real image using these state-of-the-art tools, one must first invert the image with a meaningful text prompt into the pretrained model's domain. In this paper, we introduce an accurate inversion technique and thus facilitate an intuitive text-based modification of the image. Our proposed inversion consists of two novel key components: (i) Pivotal inversion for diffusion models. While current methods aim at mapping random noise samples to a single input image, we use a single pivotal noise vector for each timestamp and optimize around it. We demonstrate that a direct inversion is inadequate on its own, but does provide a good anchor for our optimization. (ii) NULL-text optimization, where we only modify the unconditional textual embedding that is used for classifier-free guidance, rather than the input text embedding. This allows for keeping both the model weights and the conditional embedding intact and hence enables applying prompt-based editing while avoiding the cumbersome tuning of the model's weights. Our Null-text inversion, based on the publicly available Stable Diffusion model, is extensively evaluated on a variety of images and prompt editing, showing high-fidelity editing of real images.
Forward citations
Cited by 14 Pith papers
-
D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples
Mask-guided self-attention fusion creates well-aligned target images that stay visually close to poorly-aligned base images, with full denoising trajectories, and DPO on these pairs improves alignment.
-
Beyond Objects: Contextual Synthetic Data Generation for Fine-Grained Classification
Conditioning a fine-tuned text-to-image model on per-image background/pose captions and then randomly recombining those contexts across classes improves few-shot fine-grained classifier accuracy.
-
DGMO: Training-Free Audio Source Separation through Diffusion-Guided Mask Optimization
Diffusion-Guided Mask Optimization shows a frozen text-to-audio diffusion model can perform zero-shot language-queried source separation by fitting a spectrogram mask to the model's generated reference.
-
Diffusion Instruction Tuning
Lavender fine-tunes vision-language models by aligning their attention maps with Stable Diffusion's attention targets, improving accuracy on 20 benchmarks with as few as 0.13 million training examples.
-
Unpaired Multi-Domain Histopathology Virtual Staining using Dual Path Prompted Inversion
A dual-path prompt inversion method performs unpaired virtual staining by matching a structural inversion trajectory and a style reference trajectory in a pre-trained diffusion model.
-
FluxSpace: Disentangled Semantic Editing in Rectified Flow Transformers
FluxSpace performs training-free, disentangled semantic editing in rectified flow transformers by combining attention outputs with prompt-derived linear directions.
-
MyTimeMachine: Personalized Facial Age Transformation
A personalized facial age transformation method that uses an adapter network on top of the SAM global aging model, trained with 10 to 50 photos of one person, to produce re-aged images that resemble that person's actu...
-
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
A compact 4B image generation/editing system with a fast one-step VAE, native-resolution packing, RL alignment, and 4-step distillation reports competitive benchmarks against 6B–80B open models.
-
RORem: Training a Robust Object Remover with Human-in-the-Loop
RORem trains an SDXL-based object remover on a 200K-pair dataset grown by iterative human feedback and a learned discriminator, surpassing prior methods by roughly 18 points in human-judged success rate.
-
Z-STAR+: A Zero-shot Style Transfer Method via Adjusting Style Distribution
A training-free style transfer method that fuses content and style latent features in Stable Diffusion via cross-attention reweighting and a scaled adaptive instance normalization.
-
VideoDirector: Precise Video Editing via Text-to-Video Models
A video editing pipeline that extends null-text inversion and attention control to text-to-video diffusion models, using spatial-temporal decoupled guidance and multi-frame null embeddings to achieve more temporally c...
-
Instructing Text-to-Image Diffusion Models via Classifier-Guided Semantic Optimization
CASO learns per-attribute continuous embeddings via classifier gradients, steering Stable Diffusion for text-free, disentangled image editing.
-
DiffuEraser: A Diffusion Model for Video Inpainting
DiffuEraser is a stable-diffusion video inpainting model that injects ProPainter priors via DDIM inversion and expands temporal receptive fields for long-sequence consistency.
-
Watermarking across Modalities for Content Tracing and Generative AI
A thesis showing that invisible watermarks can be embedded across images, audio, text, and model weights, with statistical tests for tracing AI-generated content.
Discussion (0). Continue with ORCID to comment.