{"id":"964bf75f-120c-496d-844e-cae53828c2d4","arxiv_id":"2412.04280","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"HumanEdit provides 5,751 human-annotated, high-resolution image editing pairs with masks and a six-type instruction taxonomy, plus baseline benchmark results.","lead":"HumanEdit is a new human-annotated dataset for instruction-based image editing, with 5,751 image pairs, masks, and six edit types. It aims to fix a gap in prior datasets by adding human feedback and high-resolution real photos.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The high-quality claim rests on unverified alignment between DALL-E 2 targets and instructions; Appendix C's documented Counting/Relation failures make this load-bearing for benchmark validity.","rationale":"The reader's weakest assumption correctly identifies the DALL-E 2 generation plus human filtering as the key unverified link. I agree fully: the paper provides no independent evidence that the retained 5,751 target images are faithful to their instructions, despite the documented failure modes in Appendix C. This concern is concrete and testable, and it directly affects the benchmark's validity: if a substantial fraction of targets are misaligned, then L1/L2/CLIP/DINO scores in Tables 3-5 are measuring agreement with a faulty reference, not editing quality. The paper has real strengths—a released dataset, masks, a six-category taxonomy, and a documented four-stage pipeline—so the appropriate outcome is not rejection but a requirement for external validation before the 'high-quality' label and benchmark numbers are taken at face value. Since the reader already issued CONDITIONAL, my stress-test does not move the verdict; it sharpens the condition and specifies how to satisfy it.","tokens_in":18327,"tokens_out":3404,"duration_ms":36760,"concrete_test":"Stratified external audit: sample 100 pairs from each of Counting, Action, and Relation from HumanEdit-full. Have three external annotators judge, with source image and instruction but no access to the authors' mask or target, whether the target image obeys the instruction (e.g., for Counting, count objects; for Relation, check spatial relation). Compute per-category pass rate and inter-annotator agreement (Cohen's kappa). If pass rate <90% in any category or kappa <0.7, the human-rewarded quality claim is not supported and benchmark tables need a caveat or re-grounding.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that HumanEdit is a high-quality benchmark assumes that each retained DALL-E 2 edited image is a faithful realization of the instruction. This is load-bearing because the edited images serve both as training targets and as ground truth in Tables 3-5. The paper's own Appendix C documents that DALL-E 2 has 'limited editing capabilities for specific types,' including counting and relational edits, with examples where object counts change in the wrong direction ('remove rather than add') and where instructed relations cannot be achieved even after dozens of trials. These limitations are not merely discarded failures: they are the same generation process used for the 5,751 retained pairs. The only filter is a two-tier internal review by administrators (Section 2, Stage 4); no inter-annotator agreement, no external validation, and no quantitative check of instruction-target alignment is reported. In particular, the Counting subset (698 pairs) is especially vulnerable because DALL-E 2 is documented as insensitive to object counts; if retained Counting pairs contain wrong counts, benchmark conclusions about Counting performance are invalid. Similarly, subtle Action/Relation misalignments can survive subjective review. Without independent evidence that the retained targets satisfy their instructions, the dataset's 'high-quality' claim and its benchmark scores are not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces HumanEdit, a dataset of 5,751 instruction-based image-editing pairs constructed on real high-resolution images, with masks for every pair, six instruction categories (Action, Add, Counting, Relation, Remove, Replace), and a 400-pair core subset. The authors describe a four-stage annotation pipeline in which human annotators create instructions and use DALL-E 2 to produce edited targets, followed by administrator review that they term human-rewarded. The paper reports dataset statistics, compares with prior editing datasets, and benchmarks eight methods in mask-free and mask-provided settings. The central claim is that HumanEdit is a high-quality, human-rewarded dataset that supports both masked and mask-free editing and serves as a versatile benchmark.","tokens_in":18683,"tokens_out":3640,"duration_ms":32900,"significance":"If the quality claim can be substantiated, HumanEdit would be a useful community resource: it provides real-image editing pairs, detailed masks, a six-category taxonomy for fine-grained evaluation, and an open benchmark with multiple baselines. The paper's strengths include the detailed four-stage pipeline description, the explicit documentation of failure cases in Appendix C, the release of the dataset, and the breadth of reported statistics and baseline comparisons. The significance, however, rests on the validity of the DALL-E 2 targets as ground truth, and that validity is not yet independently established.","major_comments":[{"comment":"The central 'high-quality' claim depends on the assumption that each retained DALL-E 2 edited image is a faithful realization of its instruction. Appendix C explicitly documents DALL-E 2's limited editing capabilities for counting and relational edits, including cases where the model removes rather than adds objects (Fig. 48) and where instructed relations could not be achieved despite dozens of trials (Figs. 46, 49). Because the Counting subset (698 pairs) and Relation subset (410 pairs) in Table 1 are produced by the same model and filtered only by internal administrator review, the paper does not demonstrate that the retained targets satisfy their instructions. I request inter-annotator agreement on a random sample, an independent instruction-target alignment check, or category-wise retention statistics from the 20,000 annotated images to the final 5,751. This is load-bearing because Tables 3-5 use these targets as ground truth for evaluation.","section":"Section 2, Stage 4 and Appendix C"},{"comment":"The proportion of the dataset that supports mask-free editing is reported inconsistently: the abstract says 'a subset', the text in Section 3 says '46.5% of the data supporting editing without masks', and Figure 6(a) reports 'no need for mask 53.1%' with 'need mask 46.9%'. Since the benchmark includes mask-free settings (Tables 3 and 5), the exact split and the criterion used to determine it must be clarified. Without this, the mask-free benchmark results are not reproducible.","section":"Section 3 and Figure 6(a)"},{"comment":"The benchmark evaluates models by comparing their outputs to DALL-E 2 targets using L1, L2, CLIP-I, DINO, and CLIP-T. If some retained targets contain the artifacts or misalignments documented in Appendix C, then these scores partly measure fidelity to imperfect targets rather than editing quality. I recommend adding a human evaluation on a sample of model outputs, at least on HumanEdit-core, or reporting per-instance instruction-target alignment scores to validate the benchmark conclusions.","section":"Section 4, Tables 3-5"},{"comment":"The 'human-rewarded' mechanism is described only as 'annotators with good performance receive higher rewards, while those with poor performance are removed from the annotator teams.' No details are given on the reward scheme, the scoring rubric used by administrators, or the number of administrators and their agreement. Since this quality-control procedure is the primary evidence for the dataset's quality claim, it should be quantified, for example by reporting the number of administrators per submission and the rate of returned versus discarded submissions.","section":"Section 2, Stage 4"}],"minor_comments":[{"comment":"The header contains a typo: 'Rmove' should be 'Remove'.","section":"Table 1"},{"comment":"The label 'Toturial' is a typo and should be 'Tutorial'.","section":"Figure 2"},{"comment":"The sentence 'MagicBrush has only 46.6 input images above 1000' appears to mean 46.6% of images, not 46.6 images; please correct the typo.","section":"Section 3"},{"comment":"The Introduction mentions 'Vendi Score calculations' but does not define or cite the Vendi Score; please add a definition and reference.","section":"Introduction"},{"comment":"The sentence 'with the resulting images mostly [Podell et al., 2023, Ge et al., 2024b] showing a reduction in the number of objects' contains stray citation markers that interrupt the prose; these should be removed or moved to the end of the sentence.","section":"Appendix C.1"},{"comment":"The column 'Real-world Scenario' is defined as 'whether images edited by users in the real world are included', but HumanEdit is marked with a checkmark even though its instructions are created by annotators rather than collected from real user editing requests. The column definition or the table entry should be clarified to avoid overclaiming.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"This is a solid dataset-contribution paper with a detailed pipeline and useful resources, but the main quality claim needs independent validation. If the authors can provide inter-annotator agreement or an external alignment check, clarify the mask-free split, and quantify the review process, I would support acceptance after revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful thing here is the artifact: 5,751 human-annotated editing pairs, all with masks, about 46% mask-free, a six-way instruction taxonomy, and high-resolution Unsplash sources. That is a real resource for people working on instruction-based image editing, and the paper describes the four-stage pipeline in enough detail to be reproducible. The decision to document failure cases in Appendix C and the annotator guidance in Appendix B is honest and above the usual bar for dataset papers. I also credit them for releasing the data on Hugging Face and for running a reasonable set of baselines with default hyperparameters.\n\nThe soft spot is exactly where the stress-test note lands. The ground-truth edited images come from DALL-E 2, filtered only by internal administrator review. There is no inter-annotator agreement, no external validation, and no quantitative check that the retained targets actually satisfy their instructions. The paper's own Appendix C shows that DALL-E 2 is unreliable for counting and relational edits — including examples where counts change in the wrong direction and where the desired relation could not be produced after dozens of trials. Since the Counting subset is 698 pairs and Relation is 410, a non-trivial fraction of the benchmark's ground truth may be misaligned. If that is true, the per-category scores in Table 5 for Counting and Relation are not trustworthy. The \"human-rewarded\" framing is mostly a quality-control mechanism, which is fine, but it does not by itself establish alignment.\n\nOther issues are more minor: no error bars or significance tests in the benchmark tables, no benchmark code released, and one baseline (Meissonic) comes from the same group. None of these are fatal.\n\nBottom line: this paper deserves peer review, not desk rejection. The dataset is a plausible contribution even if the quality claim is not yet fully established. A serious referee should ask for inter-annotator agreement on a subsample, a quantitative alignment check (e.g., using a VLM to verify instruction-target consistency), or at minimum an honest caveat about the Counting and Relation subsets. I would use the dataset, but I would not rely on the benchmark rankings without those additions.","headline":"A genuinely useful released dataset for instruction-based editing, but the 'high-quality' claim rests on unverified DALL-E 2 alignment — especially shaky for Counting and Relation — so the benchmark scores should be read with caution.","tokens_in":19066,"tokens_out":1408,"would_cite":true,"duration_ms":16653,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"HumanEdit claims that a human-rewarded, four-stage annotation pipeline can produce instruction-editing pairs aligned with human preferences, and delivers a 5,751-pair dataset plus a benchmark to support that claim.","keywords":["instruction-based image editing","human-annotated dataset","human preference alignment","image editing benchmark","diffusion models","image masks","DALL-E 2","six edit categories"],"falsifier":"A count audit of the Counting subset: automatically or manually count the relevant objects in the ground-truth target images and compare against the numbers stated in the instructions; if a substantial share of pairs violates the stated count, the fidelity claim collapses. A complementary test is a human-preference study in which fresh annotators judge whether each target image satisfies its instruction, with the pass rate reported.","tokens_in":18130,"feed_emoji":"🖼️","tokens_out":4939,"duration_ms":44779,"temperature":0.7,"pith_summary":"The paper introduces HumanEdit, a curated dataset for instruction-guided image editing whose pairs were authored by human annotators and then reviewed by administrators, rather than assembled automatically. The claim is that this human-rewarded pipeline yields source-target pairs that better match what real users ask for, addressing a fidelity gap in large auto-generated editing datasets. The dataset spans six instruction categories (Action, Add, Counting, Relation, Remove, Replace), comes with masks for every image, and includes a subset detailed enough for mask-free editing. The paper also benchmarks several existing editing models on the dataset, reporting where they succeed and where they fall short. If the claim holds, HumanEdit gives the field a reliable training and evaluation resource for making instruction-following image editors behave more like human editors.","feed_headline":"5,751 human-verified pairs for instruction-based image editing","feed_subtitle":"The dataset splits edits into six categories and supports both masked and mask-free image editing.","key_machinery":"The four-stage annotation pipeline is the load-bearing mechanism: after a tutorial and quiz select annotators, images are curated from high-resolution sources, annotators create instructions and use DALL-E 2 with masks to generate edited images, and administrators perform a two-tier quality review that returns or discards submissions. Roughly 20,000 annotated images were reduced to 5,751 retained pairs. The six-category taxonomy (Action, Add, Counting, Relation, Remove, Replace) is the organizing device that turns the collection into a benchmark capable of reporting per-task strengths and weaknesses.","core_discovery":"HumanEdit is a 5,751-pair dataset for instruction-guided image editing in which every pair was hand-built: annotators wrote the edit instruction, drew the mask, and used DALL-E 2 to generate the edited image, and administrators then accepted, returned for re-annotation, or discarded each submission. The central claim is that roughly 2,500 hours of human effort across four stages make the dataset better aligned with human preferences than prior large-scale editing datasets built largely from language models and synthesis pipelines. A further contribution is the six-way taxonomy of editing tasks (Action, Add, Counting, Relation, Remove, Replace), which the authors argue supports fine-grained evaluation, and the release of a benchmark with mask-free and mask-provided baselines showing, for example, that most methods perform better on Add than on Remove. The dataset also provides masks for every image while keeping a mask-free subset, and it draws on high-resolution images from diverse sources rather than a single dataset.","pith_inferences":["At 5,751 pairs, the dataset's practical value may be more as an evaluation benchmark than as a large-scale training corpus; scaling the pipeline to millions of pairs would be costly, though the quality-controlled subset could be used to filter or validate larger auto-generated collections.","Because DALL-E 2 is weakest at Counting and Relation edits, the retained pairs in those categories may over-represent easy instances, which would make benchmark scores on those categories optimistic relative to real-world difficulty.","The fact that only 46.5 percent of instructions are detailed enough for mask-free editing suggests natural user instructions are often spatially ambiguous, signaling a need for research on instruction-driven region grounding.","The human-rewarded annotation scheme could transfer to other instruction-following generation tasks, such as video or 3D editing, where alignment with human preference is currently a bottleneck."],"forward_implications":["Models fine-tuned on HumanEdit should produce edits that follow user instructions more faithfully than models trained only on auto-generated editing data, as measured by human preference.","The six-part taxonomy makes per-task reporting possible; the benchmark numbers indicate Relation and Action edits are the hardest for current methods, pointing to where training data and architectures must improve.","The provided masks enable mask-conditioned training, while the mask-free subset allows evaluation of whether purely instruction-driven localization can replace explicit masks.","The HI-EDIT benchmark gives future work a fixed, human-verified test bed, making results across editing models comparable.","The mask-versus-mask-free split lets researchers measure how much spatial supervision is actually needed for reliable instruction editing."],"supporting_citations":[{"why":"Supplies DALL-E 2, the platform used to generate the edited target images for every dataset pair.","marker":"[Ramesh et al., 2022]"},{"why":"Defines MagicBrush, the prior human-annotated editing dataset that HumanEdit improves on and uses as a baseline and comparison point.","marker":"[Zhang et al., 2024a]"},{"why":"Provides InstructPix2Pix, the foundational instruction-editing framework and the fine-tuning/evaluation backbone for several baselines.","marker":"[Brooks et al., 2023]"},{"why":"Offers HQ-Edit, the main example of a large auto-generated editing dataset lacking human feedback that HumanEdit contrasts with.","marker":"[Hui et al., 2024]"},{"why":"Supplies SEED-Data-Edit, a large unmasked dataset whose VLM-driven annotation pipeline the authors argue introduces hallucination and misalignment.","marker":"[Ge et al., 2024a]"},{"why":"Provides CLIP, which supplies the CLIP-I and CLIP-T evaluation metrics used to report benchmark scores.","marker":"[Radford et al., 2021]"}],"fun_headline_variants":["2,500 hours of human effort yield 5,751 precise image-edit pairs","Six-way taxonomy, 5,751 human-verified edits, masks included","Human-rewarded editing dataset: 5,751 pairs, 6 instruction types","High-resolution, mask-provided image edits from real human feedback"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reliability of the dataset rests on the assumption that DALL-E 2's edited images, after human review, actually do what the instruction says; if the model's known failures in counting and spatial-relation edits slip through the reviewers' filter, the dataset's ground truth can be systematically wrong.","fun_headline_variants_meta":{"raw":{"variants":["2,500 hours of human effort yield 5,751 precise image-edit pairs","Six-way taxonomy, 5,751 human-verified edits, masks included","Human-rewarded editing dataset: 5,751 pairs, 6 instruction types","High-resolution, mask-provided image edits from real human feedback"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00035,"raw_usage":{"total_tokens":1922,"prompt_tokens":966,"completion_tokens":956,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":582,"completion_tokens_details":{"reasoning_tokens":872}},"tokens_in":582,"tokens_out":956,"duration_ms":8187,"temperature":1.0,"reasoning_tokens":872,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:32:37.276587+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A count audit of the Counting subset: automatically or manually count the relevant objects in the ground-truth target images and compare against the numbers stated in the instructions; if a substantial share of pairs violates the stated count, the fidelity claim collapses. A complementary test is a human-preference study in which fresh annotators judge whether each target image satisfies its instruction, with the pass rate reported.","supporting_citations":[],"review_version":1}