REVIEW 6 cited by
Paint by Example: Exemplar-based Image Editing with Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Language-guided image editing has achieved great success recently. In this paper, for the first time, we investigate exemplar-guided image editing for more precise control. We achieve this goal by leveraging self-supervised training to disentangle and re-organize the source image and the exemplar. However, the naive approach will cause obvious fusing artifacts. We carefully analyze it and propose an information bottleneck and strong augmentations to avoid the trivial solution of directly copying and pasting the exemplar image. Meanwhile, to ensure the controllability of the editing process, we design an arbitrary shape mask for the exemplar image and leverage the classifier-free guidance to increase the similarity to the exemplar image. The whole framework involves a single forward of the diffusion model without any iterative optimization. We demonstrate that our method achieves an impressive performance and enables controllable editing on in-the-wild images with high fidelity.
Forward citations
Cited by 6 Pith papers
-
VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control
A zero-shot diffusion framework that inserts a reference object into a video with high-fidelity appearance preservation and precise key-point trajectory motion control.
-
PixelMan: Consistent Object Editing with Diffusion Models via Pixel Manipulation and Generation
An inversion-free, training-free diffusion editing method that anchors output latents to a pixel-manipulated copy of the image achieves consistent object repositioning, resizing, and pasting in 16 steps.
-
Dynamic Try-On: Taming Video Virtual Try-on with Dynamic Attention Mechanism
A DiT-based video try-on framework that reuses the backbone as garment encoder and uses limb-aware dynamic attention to improve temporal consistency.
-
Sharp-It: A Multi-view to Multi-view Diffusion Model for 3D Synthesis and Manipulation
Sharp-It fine-tunes a multi-view diffusion model to enhance low-quality Shap-E renderings into high-quality multi-view sets that can be reconstructed into detailed 3D assets.
-
TryOffAnyone: Tiled Cloth Generation from a Dressed Person
A mask-conditioned, Stable Diffusion-based model generates tiled garment images from dressed-person photos and reports best-seed metrics that improve on prior work but with a flawed evaluation protocol.
-
DynamicAvatars: Accurate Dynamic Facial Avatars Reconstruction and Precise Editing with Diffusion Models
DynamicAvatars reconstructs dynamic 3D head avatars from video and enables prompt-based editing via dual Gaussian tracking, semantic masks, and LLM-guided diffusion editing.
Discussion (0). Continue with ORCID to comment.