Pith. sign in

REVIEW 29 cited by

HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.09990 v1 pith:RBCTCVX5 submitted 2024-04-15 cs.CV cs.AI

HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing

classification cs.CV cs.AI
keywords editingimagehigh-qualityhq-editmodelsalignmentdatadataset
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

This study introduces HQ-Edit, a high-quality instruction-based image editing dataset with around 200,000 edits. Unlike prior approaches relying on attribute guidance or human feedback on building datasets, we devise a scalable data collection pipeline leveraging advanced foundation models, namely GPT-4V and DALL-E 3. To ensure its high quality, diverse examples are first collected online, expanded, and then used to create high-quality diptychs featuring input and output images with detailed text prompts, followed by precise alignment ensured through post-processing. In addition, we propose two evaluation metrics, Alignment and Coherence, to quantitatively assess the quality of image edit pairs using GPT-4V. HQ-Edits high-resolution images, rich in detail and accompanied by comprehensive editing prompts, substantially enhance the capabilities of existing image editing models. For example, an HQ-Edit finetuned InstructPix2Pix can attain state-of-the-art image editing performance, even surpassing those models fine-tuned with human-annotated data. The project page is https://thefllood.github.io/HQEdit_web.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 29 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. C3-Bench: A Context-Aware Change Captioning Benchmark

    cs.CV 2026-06 unverdicted novelty 7.0

    C3-Bench supplies a multi-domain dataset and LLM-based evaluation protocol that exposes systematic failures in existing change captioning models outside their training regimes.

  2. RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation

    cs.CV 2026-06 unverdicted novelty 7.0

    RS-Gen proposes a plug-and-play agentic framework with a closed-loop reasoning mechanism that augments base image models to achieve SOTA results on WISE Verified and RISEBench.

  3. CV-Arena: An Open Benchmark for Instructional Computer Vision Problem Solving with Human-AI Collaborative Preferences

    cs.CV 2026-05 unverdicted novelty 7.0

    CV-Arena is a new 12K-pair benchmark for instruction-guided real-image editing with 16 task types, CogRetriever curation, and Active Elo mixed human-AI evaluation that finds gaps in 21 models and presents CV-Agent.

  4. VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset

    cs.CV 2026-05 unverdicted novelty 7.0

    VINS-120K supplies the first large-scale set of instruction-image-edited-image triplets at ultra-high resolution together with an adaptation strategy that improves detail synthesis.

  5. RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition

    cs.CV 2026-05 unverdicted novelty 7.0

    RevealLayer decomposes natural images into multiple RGBA layers using diffusion models with region-aware attention, occlusion-guided adaptation, and a composite loss, outperforming prior methods on a new benchmark dataset.

  6. EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement

    cs.CV 2026-05 unverdicted novelty 7.0

    EditRefiner uses a perception-reasoning-action-evaluation agent loop and the EditFHF-15K human feedback dataset to refine text-guided image edits more accurately than prior methods.

  7. Edit Where You Mean: Region-Aware Adapter Injection for Mask-Free Local Image Editing

    cs.CV 2026-04 unverdicted novelty 7.0

    A co-trained adapter framework enables mask-free local editing in DiTs by factorizing edit semantics from spatial location and jointly learning a mask predictor.

  8. A Sanity Check on Composed Image Retrieval

    cs.CV 2026-04 unverdicted novelty 7.0

    The paper creates FISD, a controlled benchmark for composed image retrieval that removes query ambiguity via generative models, and proposes a multi-round agentic evaluation to assess models in interactive settings.

  9. AIM-Bench: Benchmarking and Improving Affective Image Manipulation via Fine-Grained Hierarchical Control

    cs.CV 2026-04 unverdicted novelty 7.0

    AIM-Bench is the first dedicated benchmark for editing images to evoke specific emotions with fine-grained control, paired with AIM-40k dataset that delivers a 9.15% performance gain by correcting training data imbalances.

  10. Do-Undo Bench: Reversibility for Action Understanding in Image Generation

    cs.CV 2025-12 unverdicted novelty 7.0

    Do-Undo Bench is a new evaluation task and dataset that forces models to simulate forward action effects and then undo them to measure genuine action understanding in image generation.

  11. Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning

    cs.CV 2025-07 unverdicted novelty 7.0

    Presents Reason50K dataset and ReasonBrain framework for hypothetical instruction-based image editing that requires physical, temporal, causal, and story reasoning.

  12. UniEdit-Flow: Unleashing Inversion and Editing in the Era of Flow Models

    cs.CV 2025-04 unverdicted novelty 7.0

    UniEdit-Flow presents tuning-free Uni-Inv and Uni-Edit methods for inversion and editing in flow models that achieve accurate reconstruction and robust region-preserving edits across generative models.

  13. Making Implicit Preservation Intent Explicit in Conversational Image Editing

    cs.CV 2026-07 conditional novelty 6.0

    Conversational image editors fail to restore temporarily occluded content; ReSpec fixes this by explicitly selecting historical visual references and rewriting instructions to guide restoration.

  14. SpatialFlow-GRPO: Where Spatial Credit Drives Image Editing

    cs.CV 2026-06 unverdicted novelty 6.0

    SpatialFlow-GRPO improves image editing quality by converting region-aware rewards into semantic-region-level optimization signals aligned with latent positions during policy updates.

  15. SpatialFlow-GRPO: Where Spatial Credit Drives Image Editing

    cs.CV 2026-06 unverdicted novelty 6.0

    SpatialFlow-GRPO adds region-level reward feedback and spatial alignment to Flow-GRPO-style RL for image editing, reporting gains on GEdit-Bench, ImgEdit-Bench, and a new MultiEditBench.

  16. 4KLSDB: A Large-Scale Dataset for 4K Image Restoration and Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    4KLSDB supplies 129k+ curated 4K images plus validation/test splits to support training of super-resolution and text-to-image diffusion models.

  17. DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning

    cs.CV 2026-05 unverdicted novelty 6.0

    DiffCap-Bench supplies a diverse IDC benchmark with ten categories and LLM judging grounded in human difference lists to evaluate MLLMs more robustly than prior lexical metrics.

  18. EditCaption: Human-Refined SFT and HAE-DPO for Image Editing Instruction Synthesis

    cs.CV 2026-04 unverdicted novelty 6.0

    EditCaption reduces critical errors in automated image editing instructions from 47.75% to 23% via SFT and DPO, yielding fine-tuned models that match or exceed closed-source VLMs on Eval-400 and ByteMorph-Bench.

  19. EditCaption: Human-Refined SFT and HAE-DPO for Image Editing Instruction Synthesis

    cs.CV 2026-04 conditional novelty 6.0

    A 235B VLM trained with human-refined SFT and hardness-adaptive error-aware DPO cuts critical instruction errors from ~48% to ~18% and beats Gemini-3-Pro on three editing-instruction benchmarks.

  20. HorizonWeaver: Generalizable Multi-Level Semantic Editing for Driving Scenes

    cs.CV 2026-04 unverdicted novelty 6.0

    HorizonWeaver enables photorealistic, instruction-driven multi-level editing of complex driving scenes with improved generalization via a new paired dataset, language-guided masks, and joint training losses.

  21. EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning

    cs.CV 2025-09 unverdicted novelty 6.0

    EditVerse unifies image and video editing and generation in one transformer model via unified token sequences and in-context learning, trained jointly on curated video editing data plus image/video corpora and evaluat...

  22. ImgEdit: A Unified Image Editing Dataset and Benchmark

    cs.CV 2025-05 conditional novelty 6.0

    ImgEdit supplies 1.2 million curated edit pairs and a three-part benchmark that let a VLM-based model outperform prior open-source editors on adherence, quality, and detail preservation.

  23. ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations

    cs.CV 2026-06 unverdicted novelty 5.0

    ARM is a 7B autoregressive multimodal model with a unified discrete visual tokenizer and RL that performs image understanding, generation, and editing while showing cross-task synergy from preference optimization.

  24. MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow Matching

    cs.CV 2026-06 unverdicted novelty 5.0

    MT-EditFlow applies flow-matching RL with multi-reward aggregation to improve multi-turn image editing performance on models like FLUX.1-Kontext-dev by 6.85 points at turn-3.

  25. FineEdit: Fine-Grained Image Edit with Bounding Box Guidance

    cs.CV 2026-04 unverdicted novelty 5.0

    FineEdit adds multi-level bounding box injection to diffusion image editing, releases a 1.2M-pair dataset with box annotations, and shows better instruction following and background consistency than prior open models ...

  26. Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing

    cs.CV 2025-09 unverdicted novelty 5.0

    Rebalancing designer-painter roles by assigning design to the understanding module via the new DIM dataset yields SOTA image editing performance with a 4.6B model.

  27. Emerging Properties in Unified Multimodal Pretraining

    cs.CV 2025-05 unverdicted novelty 5.0

    BAGEL is a unified decoder-only model that develops emerging complex multimodal reasoning abilities after pretraining on large-scale interleaved data and outperforms prior open-source unified models.

  28. Step1X-Edit: A Practical Framework for General Image Editing

    cs.CV 2025-04 unverdicted novelty 4.0

    Step1X-Edit integrates a multimodal LLM with a diffusion decoder, trained on a custom high-quality dataset, to deliver image editing performance that surpasses open-source baselines and approaches proprietary models o...

  29. Toward Native Multimodal Modeling: A Roadmap

    cs.CV 2026-05 unverdicted novelty 3.0

    A roadmap that defines architectural nativity for multimodal models and categorizes them into Multi-to-Text, Multi-to-Target, and Multi-to-Multi types while outlining an industrial pipeline toward unified transformer-...