REVIEW 13 cited by
Diffusion Model-Based Image Editing: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Denoising diffusion models have emerged as a powerful tool for various image generation and editing tasks, facilitating the synthesis of visual content in an unconditional or input-conditional manner. The core idea behind them is learning to reverse the process of gradually adding noise to images, allowing them to generate high-quality samples from a complex distribution. In this survey, we provide an exhaustive overview of existing methods using diffusion models for image editing, covering both theoretical and practical aspects in the field. We delve into a thorough analysis and categorization of these works from multiple perspectives, including learning strategies, user-input conditions, and the array of specific editing tasks that can be accomplished. In addition, we pay special attention to image inpainting and outpainting, and explore both earlier traditional context-driven and current multimodal conditional methods, offering a comprehensive analysis of their methodologies. To further evaluate the performance of text-guided image editing algorithms, we propose a systematic benchmark, EditEval, featuring an innovative metric, LMM Score. Finally, we address current limitations and envision some potential directions for future research. The accompanying repository is released at https://github.com/SiatMMLab/Awesome-Diffusion-Model-Based-Image-Editing-Methods.
Forward citations
Cited by 13 Pith papers
-
Understanding Generative AI Capabilities in Everyday Image Editing Tasks
On real Reddit photo-editing requests, human judges prefer human edits over AI edits 66% of the time, and AI editors can satisfactorily handle about 33% of requests.
-
Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
VARIN uses a Location-aware Argmax Inversion pseudo-inverse of Gumbel-max sampling to extract editable discrete noises, enabling training-free prompt-guided editing for visual autoregressive models.
-
Make Your MoVe: Make Your 3D Contents by Adapting Multi-View Diffusion Models to External Editing
A tuning-free dual-pipeline that injects original normal latents into an edited multi-view diffusion stream, preserving geometry during 2D-to-3D appearance editing.
-
OutDreamer: Video Outpainting with a Diffusion Transformer
OutDreamer couples a diffusion transformer with mask-driven self-attention and a latent alignment loss to outpaint videos in a zero-shot manner, exceeding prior zero-shot baselines on standard benchmarks.
-
KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models
A new benchmark, KRIS-Bench, evaluates image editing models on knowledge-grounded reasoning across factual, conceptual, and procedural tasks, and finds large performance gaps in current models.
-
Provable diffusion-based posterior sampling for linear inverse problems via DDIM
A SVD-based, coordinate-wise DDIM sampler is claimed to asymptotically sample from the posterior for noisy linear inverse problems, but the proof's posterior identification step does not follow from the stated updates.
-
PairEdit: Learning Semantic Variations for Exemplar-based Image Editing
PairEdit trains two LoRA adapters on a pretrained diffusion model to capture the semantic direction between paired source-target images, enabling text-free, controllable image editing from as few as one pair.
-
Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning
A new micro-edit dataset and fine-tuning recipe appear to help multimodal LLMs notice small visual changes, but the central 'feature consistency loss' claim is not present in the method.
-
TaxaDiffusion: Progressively Trained Diffusion Model for Fine-Grained Species Generation
A diffusion model trained progressively from Kingdom to Species generates more accurate fine-grained animal images, including rare species with as few as one training sample.
-
Stationary Power-Law Solutions of Kinetic-Alfv\'{e}nic Turbulence
The submission cannot be assessed because the supplied full text is a different paper than the abstract and metadata describe.
-
SecureT2I: No More Unauthorized Manipulation on AI Generated Images from Prompts
A diffusion editing model is fine-tuned with a blur target for forbidden images and the original output for permitted images, claiming selective suppression of unauthorized edits.
-
Component Adaptive Clustering for Generalized Category Discovery
AdaGCD applies adaptive slot attention to decompose DINO image features into semantic components and pools them with global features, reporting SOTA accuracy on six GCD benchmarks.
-
DLSF: Dual-Layer Synergistic Fusion for High-Fidelity Image Syn-thesis
A softmax and spatial attention fusion of SDXL base and refiner latents yields an ImageNet FID drop of about 1 point, but the result is not statistically supported.
Discussion (0). Sign in to comment.