Pith. sign in

REVIEW 2 cited by

Cross-Image Attention for Zero-Shot Appearance Transfer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.03335 v1 pith:DDDYLTNU submitted 2023-11-06 cs.CV cs.GR

classification cs.CVcs.GR
keywords appearanceimageimagessemanticattentioncross-imagestructureacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in text-to-image generative models have demonstrated a remarkable ability to capture a deep semantic understanding of images. In this work, we leverage this semantic knowledge to transfer the visual appearance between objects that share similar semantics but may differ significantly in shape. To achieve this, we build upon the self-attention layers of these generative models and introduce a cross-image attention mechanism that implicitly establishes semantic correspondences across images. Specifically, given a pair of images -- one depicting the target structure and the other specifying the desired appearance -- our cross-image attention combines the queries corresponding to the structure image with the keys and values of the appearance image. This operation, when applied during the denoising process, leverages the established semantic correspondences to generate an image combining the desired structure and appearance. In addition, to improve the output image quality, we harness three mechanisms that either manipulate the noisy latent codes or the model's internal representations throughout the denoising process. Importantly, our approach is zero-shot, requiring no optimization or training. Experiments show that our method is effective across a wide range of object categories and is robust to variations in shape, size, and viewpoint between the two input images.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PixelMan: Consistent Object Editing with Diffusion Models via Pixel Manipulation and Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    An inversion-free, training-free diffusion editing method that anchors output latents to a pixel-manipulated copy of the image achieves consistent object repositioning, resizing, and pasting in 16 steps.

  2. MixSA: Training-free Reference-based Sketch Extraction via Mixture-of-Self-Attention

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A training-free method transfers a reference sketch's stroke style onto a photo's edge map by replacing self-attention keys and values in Stable Diffusion, with two sliders for texture and style.

Pith tools