Pith. sign in

REVIEW 11 cited by

UltraEdit: Instruction-based Fine-Grained Image Editing at Scale

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.05282 v2 pith:KQMOFJAW submitted 2024-07-07 cs.CV

classification cs.CV
keywords editingimageultraeditmodelsautomaticallydatadatasetdatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents UltraEdit, a large-scale (approximately 4 million editing samples), automatically generated dataset for instruction-based image editing. Our key idea is to address the drawbacks in existing image editing datasets like InstructPix2Pix and MagicBrush, and provide a systematic approach to producing massive and high-quality image editing samples. UltraEdit offers several distinct advantages: 1) It features a broader range of editing instructions by leveraging the creativity of large language models (LLMs) alongside in-context editing examples from human raters; 2) Its data sources are based on real images, including photographs and artworks, which provide greater diversity and reduced bias compared to datasets solely generated by text-to-image models; 3) It also supports region-based editing, enhanced by high-quality, automatically produced region annotations. Our experiments show that canonical diffusion-based editing baselines trained on UltraEdit set new records on MagicBrush and Emu-Edit benchmarks. Our analysis further confirms the crucial role of real image anchors and region-based editing data. The dataset, code, and models can be found in https://ultra-editing.github.io.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. REALEDIT: Reddit Edits As a Large-scale Empirical Dataset for Image Transformations

    cs.CV 2025-02 conditional novelty 7.0 of 10

    REALEDIT provides a large-scale, real-world image editing dataset from Reddit and demonstrates that training on it improves performance on authentic user requests.

  2. Instruction-based Image Manipulation by Watching How Things Move

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A diffusion editing model, InstructMove, is trained on video frame pairs annotated by MLLMs using spatial conditioning, enabling non-rigid edits and viewpoint changes.

  3. ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    An automatically generated training dataset and a fine-tuned LLaVA-NeXT model produce an image editing evaluation scorer that aligns with human preference and serves as a reward model for improving editing models.

  4. Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing

    cs.CV 2025-07 conditional novelty 6.0 of 10

    X-Planner, an MLLM-based planner, decomposes complex image-editing instructions into localized sub-edits with masks and boxes, improving editing quality on standard and new complex benchmarks.

  5. EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new human-labeled benchmark shows leading vision-language models are unreliable at judging image edits, and the authors' methods improve artifact detection and difference captioning.

  6. OmniStyle: Filtering High Quality Style Transfer Data at Scale

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new million-triplet dataset and a diffusion transformer model that performs text-guided and image-guided style transfer, with a filtering pipeline used to curate high-quality training examples.

  7. Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning

    cs.CV 2026-03 conditional novelty 5.5 of 10

    Goal-driven selection of 1× multimodal instruction subsets reaches a 512k Uni-10x baseline after ~27–35k samples and improves accuracy by up to +3.08 pp under a fixed Qwen3-VL recipe.

  8. Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A compact 4B image generation/editing system with a fast one-step VAE, native-resolution packing, RL alignment, and 4-step distillation reports competitive benchmarks against 6B–80B open models.

  9. ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A released 6.4 million pair dataset and 613 sample benchmark for instruction-guided image editing of non-rigid motions, plus a Flux.1-dev based baseline that outperforms open-source methods on the new benchmark.

  10. DrivingGaussian++: Towards Realistic Reconstruction and Editable Simulation for Surrounding Dynamic Driving Scenes

    cs.CV 2025-08 conditional novelty 4.0 of 10

    DrivingGaussian++ reconstructs dynamic surround-view driving scenes and performs training-free multi-task editing (weather, texture, object manipulation) using Gaussians, diffusion models, and LLM-generated trajectories.

  11. TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model

    cs.CV 2025-07 conditional novelty 4.0 of 10

    TalkFashion, a text-driven virtual try-on assistant, reports better semantic consistency and visual quality than four baselines on VITON-HD by combining an LLM router, catalog matching, and automatic mask generation.

Pith tools