Pith. sign in

REVIEW 32 cited by

HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.09990 v1 pith:RBCTCVX5 submitted 2024-04-15 cs.CV cs.AI

classification cs.CVcs.AI
keywords editingimagehigh-qualityhq-editmodelsalignmentdatadataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study introduces HQ-Edit, a high-quality instruction-based image editing dataset with around 200,000 edits. Unlike prior approaches relying on attribute guidance or human feedback on building datasets, we devise a scalable data collection pipeline leveraging advanced foundation models, namely GPT-4V and DALL-E 3. To ensure its high quality, diverse examples are first collected online, expanded, and then used to create high-quality diptychs featuring input and output images with detailed text prompts, followed by precise alignment ensured through post-processing. In addition, we propose two evaluation metrics, Alignment and Coherence, to quantitatively assess the quality of image edit pairs using GPT-4V. HQ-Edits high-resolution images, rich in detail and accompanied by comprehensive editing prompts, substantially enhance the capabilities of existing image editing models. For example, an HQ-Edit finetuned InstructPix2Pix can attain state-of-the-art image editing performance, even surpassing those models fine-tuned with human-annotated data. The project page is https://thefllood.github.io/HQEdit_web.

Discussion (0). Sign in to comment.

Forward citations

Cited by 32 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. C3-Bench: A Context-Aware Change Captioning Benchmark

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    C3-Bench supplies a multi-domain dataset and LLM-based evaluation protocol that exposes systematic failures in existing change captioning models outside their training regimes.

  2. RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    RS-Gen proposes a plug-and-play agentic framework with a closed-loop reasoning mechanism that augments base image models to achieve SOTA results on WISE Verified and RISEBench.

  3. CV-Arena: An Open Benchmark for Instructional Computer Vision Problem Solving with Human-AI Collaborative Preferences

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    CV-Arena is a new 12K-pair benchmark for instruction-guided real-image editing with 16 task types, CogRetriever curation, and Active Elo mixed human-AI evaluation that finds gaps in 21 models and presents CV-Agent.

  4. VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    VINS-120K supplies the first large-scale set of instruction-image-edited-image triplets at ultra-high resolution together with an adaptation strategy that improves detail synthesis.

  5. RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    RevealLayer decomposes natural images into multiple RGBA layers using diffusion models with region-aware attention, occlusion-guided adaptation, and a composite loss, outperforming prior methods on a new benchmark dataset.

  6. EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    EditRefiner uses a perception-reasoning-action-evaluation agent loop and the EditFHF-15K human feedback dataset to refine text-guided image edits more accurately than prior methods.

  7. Edit Where You Mean: Region-Aware Adapter Injection for Mask-Free Local Image Editing

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    A co-trained adapter framework enables mask-free local editing in DiTs by factorizing edit semantics from spatial location and jointly learning a mask predictor.

  8. A Sanity Check on Composed Image Retrieval

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    The paper creates FISD, a controlled benchmark for composed image retrieval that removes query ambiguity via generative models, and proposes a multi-round agentic evaluation to assess models in interactive settings.

  9. AIM-Bench: Benchmarking and Improving Affective Image Manipulation via Fine-Grained Hierarchical Control

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    AIM-Bench is the first dedicated benchmark for editing images to evoke specific emotions with fine-grained control, paired with AIM-40k dataset that delivers a 9.15% performance gain by correcting training data imbalances.

  10. Do-Undo Bench: Reversibility for Action Understanding in Image Generation

    cs.CV 2025-12 unverdicted novelty 7.0 of 10

    Do-Undo Bench is a new evaluation task and dataset that forces models to simulate forward action effects and then undo them to measure genuine action understanding in image generation.

  11. Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning

    cs.CV 2025-07 unverdicted novelty 7.0 of 10

    Presents Reason50K dataset and ReasonBrain framework for hypothetical instruction-based image editing that requires physical, temporal, causal, and story reasoning.

  12. UniEdit-Flow: Unleashing Inversion and Editing in the Era of Flow Models

    cs.CV 2025-04 unverdicted novelty 7.0 of 10

    UniEdit-Flow presents tuning-free Uni-Inv and Uni-Edit methods for inversion and editing in flow models that achieve accurate reconstruction and robust region-preserving edits across generative models.

  13. Making Implicit Preservation Intent Explicit in Conversational Image Editing

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Conversational image editors fail to restore temporarily occluded content; ReSpec fixes this by explicitly selecting historical visual references and rewriting instructions to guide restoration.

  14. SpatialFlow-GRPO: Where Spatial Credit Drives Image Editing

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    SpatialFlow-GRPO improves image editing quality by converting region-aware rewards into semantic-region-level optimization signals aligned with latent positions during policy updates.

  15. SpatialFlow-GRPO: Where Spatial Credit Drives Image Editing

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    SpatialFlow-GRPO adds region-level reward feedback and spatial alignment to Flow-GRPO-style RL for image editing, reporting gains on GEdit-Bench, ImgEdit-Bench, and a new MultiEditBench.

  16. 4KLSDB: A Large-Scale Dataset for 4K Image Restoration and Generation

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    4KLSDB supplies 129k+ curated 4K images plus validation/test splits to support training of super-resolution and text-to-image diffusion models.

  17. DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    DiffCap-Bench supplies a diverse IDC benchmark with ten categories and LLM judging grounded in human difference lists to evaluate MLLMs more robustly than prior lexical metrics.

  18. EditCaption: Human-Refined SFT and HAE-DPO for Image Editing Instruction Synthesis

    cs.CV 2026-04 conditional novelty 6.0 of 10

    A 235B VLM trained with human-refined SFT and hardness-adaptive error-aware DPO cuts critical instruction errors from ~48% to ~18% and beats Gemini-3-Pro on three editing-instruction benchmarks.

  19. EditCaption: Human-Refined SFT and HAE-DPO for Image Editing Instruction Synthesis

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    EditCaption reduces critical errors in automated image editing instructions from 47.75% to 23% via SFT and DPO, yielding fine-tuned models that match or exceed closed-source VLMs on Eval-400 and ByteMorph-Bench.

  20. HorizonWeaver: Generalizable Multi-Level Semantic Editing for Driving Scenes

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    HorizonWeaver enables photorealistic, instruction-driven multi-level editing of complex driving scenes with improved generalization via a new paired dataset, language-guided masks, and joint training losses.

  21. EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning

    cs.CV 2025-09 unverdicted novelty 6.0 of 10

    EditVerse unifies image and video editing and generation in one transformer model via unified token sequences and in-context learning, trained jointly on curated video editing data plus image/video corpora and evaluat...

  22. Trade-offs in Image Generation: How Do Different Dimensions Interact?

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new benchmark and VLM-as-judge metric map trade-offs among ten image-generation dimensions across 14 models, with a visualization called DTM.

  23. ImgEdit: A Unified Image Editing Dataset and Benchmark

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ImgEdit supplies 1.2 million curated edit pairs and a three-part benchmark that let a VLM-based model outperform prior open-source editors on adherence, quality, and detail preservation.

  24. JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

    cs.CV 2026-07 conditional novelty 5.0 of 10

    JarvisHub open-sources a three-layer canvas-state, protocol-bridge, and agent-runtime harness so multimodal creative agents can inspect and update a shared editable project graph over long workflows.

  25. ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    ARM is a 7B autoregressive multimodal model with a unified discrete visual tokenizer and RL that performs image understanding, generation, and editing while showing cross-task synergy from preference optimization.

  26. MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow Matching

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    MT-EditFlow applies flow-matching RL with multi-reward aggregation to improve multi-turn image editing performance on models like FLUX.1-Kontext-dev by 6.85 points at turn-3.

  27. FineEdit: Fine-Grained Image Edit with Bounding Box Guidance

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    FineEdit adds multi-level bounding box injection to diffusion image editing, releases a 1.2M-pair dataset with box annotations, and shows better instruction following and background consistency than prior open models ...

  28. SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models

    cs.CV 2025-10 conditional novelty 5.0 of 10

    A unified multimodal model can improve its own text-to-image generation by using its understanding module as a rewarder in a global-plus-local reward-weighted training loop.

  29. Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing

    cs.CV 2025-09 unverdicted novelty 5.0 of 10

    Rebalancing designer-painter roles by assigning design to the understanding module via the new DIM dataset yields SOTA image editing performance with a 4.6B model.

  30. Emerging Properties in Unified Multimodal Pretraining

    cs.CV 2025-05 unverdicted novelty 5.0 of 10

    BAGEL is a unified decoder-only model that develops emerging complex multimodal reasoning abilities after pretraining on large-scale interleaved data and outperforms prior open-source unified models.

  31. Step1X-Edit: A Practical Framework for General Image Editing

    cs.CV 2025-04 unverdicted novelty 4.0 of 10

    Step1X-Edit integrates a multimodal LLM with a diffusion decoder, trained on a custom high-quality dataset, to deliver image editing performance that surpasses open-source baselines and approaches proprietary models o...

  32. Toward Native Multimodal Modeling: A Roadmap

    cs.CV 2026-05 unverdicted novelty 3.0 of 10

    A roadmap that defines architectural nativity for multimodal models and categorizes them into Multi-to-Text, Multi-to-Target, and Multi-to-Multi types while outlining an industrial pipeline toward unified transformer-...

Pith tools