Pith. sign in

REVIEW 3 cited by

DiffUHaul: A Training-Free Method for Object Dragging in Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.01594 v2 pith:OEPTFYUA submitted 2024-06-03 cs.CV cs.GRcs.LG

classification cs.CVcs.GRcs.LG
keywords imagesobjectlocalizedmodeldenoisingdiffuhauleditingmethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-image diffusion models have proven effective for solving many image editing tasks. However, the seemingly straightforward task of seamlessly relocating objects within a scene remains surprisingly challenging. Existing methods addressing this problem often struggle to function reliably in real-world scenarios due to lacking spatial reasoning. In this work, we propose a training-free method, dubbed DiffUHaul, that harnesses the spatial understanding of a localized text-to-image model, for the object dragging task. Blindly manipulating layout inputs of the localized model tends to cause low editing performance due to the intrinsic entanglement of object representation in the model. To this end, we first apply attention masking in each denoising step to make the generation more disentangled across different objects and adopt the self-attention sharing mechanism to preserve the high-level object appearance. Furthermore, we propose a new diffusion anchoring technique: in the early denoising steps, we interpolate the attention features between source and target images to smoothly fuse new layouts with the original appearance; in the later denoising steps, we pass the localized features from the source images to the interpolated images to retain fine-grained object details. To adapt DiffUHaul to real-image editing, we apply a DDPM self-attention bucketing that can better reconstruct real images with the localized model. Finally, we introduce an automated evaluation pipeline for this task and showcase the efficacy of our method. Our results are reinforced through a user preference study.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Motion by Queries: Identity-Motion Trade-offs in Text-to-Video Generation

    cs.CV 2024-12 conditional novelty 7.0 of 10

    Query features in video diffusion models encode both motion and identity, enabling efficient zero-shot motion transfer and training-free multi-shot character consistency.

  2. Motion Modes: What Could Happen Next?

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Guiding a pre-trained image-to-video model with hand-designed energy functions discovers multiple distinct, object-focused motions from a single image, without training.

  3. Stable Flow: Vital Layers for Training-Free Image Editing

    cs.CV 2024-11 conditional novelty 6.0 of 10

    An automatic vital-layer selection for FLUX enables training-free, stable text-driven image editing via selective attention injection.

Pith tools