Pith. sign in

REVIEW 3 cited by

Object-aware Inversion and Reassembly for Image Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.12149 v2 pith:VBEHCQLQ submitted 2023-10-18 cs.CV

classification cs.CV
keywords editingimageinversionpairsstepstargetnumberoptimal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

By comparing the original and target prompts, we can obtain numerous editing pairs, each comprising an object and its corresponding editing target. To allow editability while maintaining fidelity to the input image, existing editing methods typically involve a fixed number of inversion steps that project the whole input image to its noisier latent representation, followed by a denoising process guided by the target prompt. However, we find that the optimal number of inversion steps for achieving ideal editing results varies significantly among different editing pairs, owing to varying editing difficulties. Therefore, the current literature, which relies on a fixed number of inversion steps, produces sub-optimal generation quality, especially when handling multiple editing pairs in a natural image. To this end, we propose a new image editing paradigm, dubbed Object-aware Inversion and Reassembly (OIR), to enable object-level fine-grained editing. Specifically, we design a new search metric, which determines the optimal inversion steps for each editing pair, by jointly considering the editability of the target and the fidelity of the non-editing region. We use our search metric to find the optimal inversion step for each editing pair when editing an image. We then edit these editing pairs separately to avoid concept mismatch. Subsequently, we propose an additional reassembly step to seamlessly integrate the respective editing results and the non-editing region to obtain the final edited image. To systematically evaluate the effectiveness of our method, we collect two datasets called OIRBench for benchmarking single- and multi-object editing, respectively. Experiments demonstrate that our method achieves superior performance in editing object shapes, colors, materials, categories, etc., especially in multi-object editing scenarios.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality

    cs.CV 2025-12 unverdicted novelty 7.0 of 10

    LivingSwap is the first video reference-guided face swapping model that uses keyframe conditioning and temporal stitching to preserve source video realism with high fidelity across long sequences.

  2. BindEdit: Taming Attention Leakage for Precise Multi-Object Image Editing

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    BindEdit suppresses two forms of attention leakage in diffusion-based editing by binding target tokens to regions, rebalancing cross-attention, and adding a region fidelity term, plus a new multi-object benchmark.

  3. RectifiedHR: Enable Efficient High-Resolution Synthesis via Energy Rectification

    cs.CV 2025-03 unverdicted novelty 5.0 of 10

    RectifiedHR is a training-free method that uses noise refresh and latent energy analysis to enable efficient high-resolution synthesis in diffusion models.

Pith tools