Pith. sign in

REVIEW 6 cited by

UltraEdit: Instruction-based Fine-Grained Image Editing at Scale

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.05282 v2 pith:KQMOFJAW submitted 2024-07-07 cs.CV

classification cs.CV
keywords editingimageultraeditmodelsautomaticallydatadatasetdatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents UltraEdit, a large-scale (approximately 4 million editing samples), automatically generated dataset for instruction-based image editing. Our key idea is to address the drawbacks in existing image editing datasets like InstructPix2Pix and MagicBrush, and provide a systematic approach to producing massive and high-quality image editing samples. UltraEdit offers several distinct advantages: 1) It features a broader range of editing instructions by leveraging the creativity of large language models (LLMs) alongside in-context editing examples from human raters; 2) Its data sources are based on real images, including photographs and artworks, which provide greater diversity and reduced bias compared to datasets solely generated by text-to-image models; 3) It also supports region-based editing, enhanced by high-quality, automatically produced region annotations. Our experiments show that canonical diffusion-based editing baselines trained on UltraEdit set new records on MagicBrush and Emu-Edit benchmarks. Our analysis further confirms the crucial role of real image anchors and region-based editing data. The dataset, code, and models can be found in https://ultra-editing.github.io.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    An automatically generated training dataset and a fine-tuned LLaVA-NeXT model produce an image editing evaluation scorer that aligns with human preference and serves as a reward model for improving editing models.

  2. Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing

    cs.CV 2025-07 conditional novelty 6.0 of 10

    X-Planner, an MLLM-based planner, decomposes complex image-editing instructions into localized sub-edits with masks and boxes, improving editing quality on standard and new complex benchmarks.

  3. Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning

    cs.CV 2026-03 conditional novelty 5.5 of 10

    Goal-driven selection of 1× multimodal instruction subsets reaches a 512k Uni-10x baseline after ~27–35k samples and improves accuracy by up to +3.08 pp under a fixed Qwen3-VL recipe.

  4. Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A compact 4B image generation/editing system with a fast one-step VAE, native-resolution packing, RL alignment, and 4-step distillation reports competitive benchmarks against 6B–80B open models.

  5. DrivingGaussian++: Towards Realistic Reconstruction and Editable Simulation for Surrounding Dynamic Driving Scenes

    cs.CV 2025-08 conditional novelty 4.0 of 10

    DrivingGaussian++ reconstructs dynamic surround-view driving scenes and performs training-free multi-task editing (weather, texture, object manipulation) using Gaussians, diffusion models, and LLM-generated trajectories.

  6. TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model

    cs.CV 2025-07 conditional novelty 4.0 of 10

    TalkFashion, a text-driven virtual try-on assistant, reports better semantic consistency and visual quality than four baselines on VITON-HD by combining an LLM router, catalog matching, and automatic mask generation.

Pith tools