Pith. sign in

REVIEW 5 cited by

DiT4Edit: Diffusion Transformer for Image Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.03286 v2 pith:HIOHGIYO submitted 2024-11-05 cs.CV

classification cs.CV
keywords editingimagediffusiondit4editimagesalgorithmcompareddemonstrate
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Despite recent advances in UNet-based image editing, methods for shape-aware object editing in high-resolution images are still lacking. Compared to UNet, Diffusion Transformers (DiT) demonstrate superior capabilities to effectively capture the long-range dependencies among patches, leading to higher-quality image generation. In this paper, we propose DiT4Edit, the first Diffusion Transformer-based image editing framework. Specifically, DiT4Edit uses the DPM-Solver inversion algorithm to obtain the inverted latents, reducing the number of steps compared to the DDIM inversion algorithm commonly used in UNet-based frameworks. Additionally, we design unified attention control and patches merging, tailored for transformer computation streams. This integration allows our framework to generate higher-quality edited images faster. Our design leverages the advantages of DiT, enabling it to surpass UNet structures in image editing, especially in high-resolution and arbitrary-size images. Extensive experiments demonstrate the strong performance of DiT4Edit across various editing scenarios, highlighting the potential of Diffusion Transformers in supporting image editing.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent

    cs.CV 2025-08 conditional novelty 5.0 of 10

    DescriptiveEdit turns semantic editing into reference-conditioned text-to-image generation, reporting state-of-the-art scores on the Emu Edit benchmark with a frozen backbone and about 75M trainable parameters.

  2. AE-NeRF: Augmenting Event-Based Neural Radiance Fields for Non-ideal Conditions and Larger Scene

    cs.CV 2025-01 conditional novelty 5.0 of 10

    AE-NeRF jointly optimizes camera poses and an event-based NeRF with a proposal network and four event-specific losses, improving novel view synthesis under noisy poses and non-uniform motion.

  3. MagicQuill: An Intelligent Interactive Image Editing System

    cs.CV 2024-11 conditional novelty 5.0 of 10

    MagicQuill combines brush-based edge and color control with an MLLM that guesses user intent, enabling fast interactive image edits without typing prompts.

  4. DFVEdit: Conditional Delta Flow Vector for Zero-shot Video Editing

    cs.CV 2025-06 conditional novelty 4.0 of 10

    DFVEdit edits videos by iteratively subtracting a conditional delta flow vector, the difference between the model's predictions under the target and source prompts, from the latent representation of the source video.

  5. Enhancing Image Generation Fidelity via Progressive Prompts

    cs.CV 2025-01 reject novelty 3.0 of 10

    DiTPipe injects LLM-written high-level and low-level prompts into masked regional cross-attention layers of a diffusion transformer to improve local prompt following.

Pith tools