{"id":"fe8fbb7c-45d6-447c-9883-961ba81bb8dd","arxiv_id":"2505.16181","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"On real Reddit photo-editing requests, human judges prefer human edits over AI edits 66% of the time, and AI editors can satisfactorily handle about 33% of requests.","lead":"Researchers analyzed 83,000 real photo-editing requests from Reddit along with 305,000 human-made edits and tested 49 AI editors on a subset. They found human judges still prefer human edits about two-thirds of the time, and current AI editors can only handle about one third of everyday requests.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 33.35% figure is a vote-level weighted average, not a request-level estimate; the paper's own formula cannot support the claim that one-third of requests are handleable.","rationale":"The reader identified sample representativeness as the weakest assumption. That is a real limitation, but the more fundamental problem is that the 33.35% number does not measure what the abstract claims. Sec. 5.4's formula uses vote-level AI win/tie rates from Tab. A8; because each request produces many pairwise comparisons, the resulting percentage is a comparison-level average, not a request-level fraction. The paper's own sentence 'if it is rated Tie or AI wins ... we consider it satisfactorily handled' is not operationalized at the request level anywhere. This is internal to the paper's method, not an external validity issue, so it is more load-bearing than representativeness. The pooling of all 49 models further widens the gap between the headline and the computation. Both issues are fixable by re-analysis, so the verdict stays conditional rather than reject; the dataset and the VLM-human divergence result remain valuable. My read therefore does not change the reader's CONDITIONAL verdict, but for a different reason.","tokens_in":39079,"tokens_out":8471,"duration_ms":63294,"concrete_test":"Recompute request-level handleability on PSR-328: define a request as handled if at least one AI edit ties or beats the majority of its human edits (or, alternatively, if the best AI edit beats the best human edit); aggregate these request-level outcomes by action, reweight by D_v, and compare to 33.35%. If the estimate shifts by more than a few absolute points, the headline claim should be revised. Also report the same computation restricted to SeedEdit, GPT-4o, and Gemini-2.0-Flash to separate 'best' from 'average' performance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantity in Sec. 5.4 is computed as sum_v D_v * AI_v = 33.35%, where D_v is the request-level proportion of action v in the full dataset and AI_v is the AI win+tie rate from Tab. A8. Tab. A8 is built from pairwise comparisons: each vote pits one AI edit against one human edit. With ~5 human edits and 7 AI edits per request (Sec. 4), the number of votes per request varies, so requests with more edits dominate AI_v. A vote-level win/tie rate is not a request-level handleability rate: a request can have some AI-vs-human pairs win and others lose. The formula mixes a request-level weight with a vote-level rate, so 33.35% estimates the chance that a randomly drawn AI edit ties or beats a randomly drawn human edit, not the fraction of requests for which an AI editor produces a satisfactory result. In addition, pooling all 49 models, including many weak HuggingFace tools, contradicts the abstract's 'best AI editors'; per-action rates for SeedEdit alone (Tab. A9) are substantially higher on most actions. The headline number is therefore neither a request-level percentage nor an upper bound on best-model performance.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces PSR, a large-scale dataset of 83k real-world image-editing requests collected from r/PhotoshopRequest over 2013-2025, together with 305k human-made edits. The authors annotate each request with WordNet-derived subjects, 15 editing-action labels, and three creativity levels. On a stratified subset of 328 requests, they collect 4,359 human preference votes comparing 1,644 human edits against 2,296 AI edits from 49 models, and report that human raters prefer human edits 66% of the time. The paper further estimates that AI editors can satisfactorily handle 33.35% of requests, finds that VLM judges (GPT-4o, o1, Gemini-2.0-Flash-Thinking) agree poorly with human raters and show strong biases (e.g., o1 prefers GPT-4o edits 83.9% of the time), and documents that AI editors tend to make unrequested aesthetic enhancements and struggle to preserve identity.","tokens_in":39292,"tokens_out":5292,"duration_ms":45432,"significance":"If the headline results hold, the dataset and findings would be a valuable contribution: PSR is the largest real-world image-editing request dataset with human edits, the action/creativity taxonomy is more user-intent-oriented than prior tool-based taxonomies, and the human study with 4.3k votes plus VLM-judge comparisons is unusually large for this task. The paper also ships code, qualitative examples, and a controlled identity-drift experiment with DINOv2 distances, all of which strengthen reproducibility. The observation that VLM judges can be severely biased in edit evaluation is practically important for automated benchmarking. However, the central quantitative claim of '33.35% of requests can be handled by AI' is not supported by the computation as reported, because the estimate is built from vote-level win/tie rates rather than request-level outcomes; this must be corrected before the headline result can be accepted.","major_comments":[{"comment":"The 33.35% estimate conflates vote-level win/tie rates with request-level handleability. The formula multiplies D_v (the request-level proportion of each action in the full dataset) by AI_v, which Table A8 defines as the percentage of pairwise votes won or tied by the AI across all requests and all 49 models. Because each request has roughly five human edits and seven AI edits, a single request contributes many votes and can have both AI wins and AI losses across those pairs. The stated definition of 'satisfactorily handled' is per request ('if it is rated Tie or AI wins by human raters'), so AI_v must be computed after aggregating votes to the request level (e.g., a request is handled if at least one AI edit ties or beats all human edits for that request). As written, 33.35% is the expected probability that a randomly selected AI-edit versus human-edit pair is won or tied by the AI, not the fraction of requests for which an AI editor produces a satisfactory result.","section":"Sec. 5.4 and Table A8"},{"comment":"The abstract attributes the 33% figure to 'the best AI editors (including GPT-4o, Gemini-2.0-Flash, SeedEdit)', but the computation in Sec. 5.4 pools all 49 models, including many low-performance Hugging Face tools, as shown in Table A8. This is not equivalent to best-model performance. For example, Table A9 reports SeedEdit's per-action AI win+Tie rates that are substantially higher than the pooled rates for most actions (merge 63.0% vs 39.3%, add 52.0% vs 38.0%, delete 46.5% vs 34.9%). The paper should either report request-level handleability separately for each of the three SOTA models, or explicitly characterize 33.35% as an aggregate over all 49 models, not as the capability of the best AI editors.","section":"Abstract and Sec. 5.4"},{"comment":"The weighted estimate is fragile because per-action vote counts are small for several actions and the sample was not stratified by action. PSR-328 was stratified by creativity level only, and Table A8 shows very low counts for clone (n=25), zoom (n=75), crop (n=129), and specialized operation (n=116). The formula weights these noisy per-action rates by the full-dataset action frequencies, but the paper reports only the point estimate 33.35% with no confidence intervals, bootstrap, or sensitivity analysis. Given the central role of this number, the authors should provide uncertainty quantification and show that the conclusions are robust to excluding low-count actions.","section":"Sec. 5.4 and Table A8"}],"minor_comments":[{"comment":"The headline comparison (66.0% human preference vs. 25.8% AI preference) is reported without confidence intervals or significance tests. Because votes are clustered by request and by rater, cluster-robust intervals would be appropriate for assessing the strength of the claim.","section":"Sec. 5.1"},{"comment":"The human rater pool is a convenience sample of 122 volunteers, one-third of whom are professional image editors. The paper should discuss how this composition may affect the measured human-preference rates and the generalizability of the comparison.","section":"Appendix E"},{"comment":"The figure labels '2k AI edits' and '1.6k human edits', while the text and Table 3 report 2,296 AI edits and 1,644 human edits; the labels should be made consistent with the exact counts.","section":"Fig. 1"},{"comment":"The limitations paragraph acknowledges LLM-based annotation biases and the exclusion of some unavailable models, which is appropriate. It should also mention the vote-level versus request-level aggregation issue in Sec. 5.4 and the small per-action sample sizes as limitations of the 33.35% estimate.","section":"Sec. 6 Limitations"},{"comment":"The column header 'AI Win+Tie' is described as 'indicating the percentage of % requests that can already be handled', which is misleading given that the values are vote-level rates. The header should say 'percentage of votes' unless the quantity is redefined at the request level.","section":"Table A8"},{"comment":"The mathematical notation 'Pv v=1 Dv ×AIv' is garbled; the summation symbol and indices need to be typeset correctly, and the equation should explicitly state that D_v and AI_v are defined at different levels (request-level and vote-level, respectively).","section":"Sec. 5.4"}],"recommendation":"major_revision","confidential_remarks":"The dataset and empirical evaluation are genuinely valuable and appear to be conducted in good faith; there is no sign of circular reasoning or fabricated results. My main concern is that the paper's most public-facing number, the 33.35% 'requests handled by AI' claim, is computed from vote-level aggregates and is also pooled over many weak models, so it does not support the abstract's wording. This is fixable by re-aggregating the existing vote data at the request level and reporting per-model results, which is why I recommend major revision rather than rejection. The paper fits the journal's scope well, and with the corrected analysis it could be a strong contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the useful news: PSR is a genuinely valuable dataset—83k real editing requests with 305k human edits, tagged with WordNet subjects, 15 action verbs, and creativity levels. It's the largest of its kind and gives the community a realistic test set instead of synthetic instructions. The human study is careful: 328 requests stratified by creativity, 4.3k human votes, and a separate VLM-as-judge experiment. The VLM result is the most solid finding: all three VLMs diverge sharply from human raters, with o1 preferring GPT-4o edits 85% of the time and Cohen's kappa around 0.2. That's an important caution for anyone using VLMs to auto-evaluate editing. The qualitative failure analysis—identity drift, unrequested touch-ups—rings true and is backed by the controlled shirt-color DINOv2 experiment.\n\nNow the soft spot, and it's a load-bearing one. The headline 'AI can handle ~33% of requests' (Sec 5.4) doesn't mean what it says. The 33.35% is sum_v D_v * AI_v, where AI_v is the win+tie rate from pairwise votes, not the fraction of requests with a satisfactory AI edit. Each request contributes multiple votes (up to ~35), so a vote-level rate is not a request-level probability. The formula is also averaged over all 49 models, including a long tail of weak HuggingFace tools; the abstract's 'best AI editors' is not what Table A8 reports. SeedEdit alone (Table A9) reaches 46–58% win+tie on the most common actions. The paper's own definition—'if it is rated Tie or AI wins, consider it satisfactorily handled'—implies a request-level judgment, but the computation doesn't implement it. That needs to be fixed, either by aggregating votes per request or by clearly relabeling the number as a vote-level comparison.\n\nThe other concerns are minor-to-moderate: no confidence intervals, a convenience sample of human raters (one-third professional editors), and PSR-328 stratified by creativity rather than action, so per-action rates may not generalize. The reweighting by action frequency helps but can't fix action–creativity interactions.\n\nBottom line: the dataset, the VLM-judge result, and the failure-mode analysis are real contributions. The 33% number is not, as stated. Rewrite it and the paper becomes a solid benchmark contribution. I'd send it to review, with the explicit request that the referee scrutinize the definition of 'handleable'.","headline":"Valuable dataset and a solid VLM-judge result, but the headline 'one-third of requests' is a vote-level statistic, not a request-level estimate.","tokens_in":39863,"tokens_out":4280,"would_cite":true,"duration_ms":36022,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Current AI image editors can handle about one-third of real-world editing requests, according to 4,359 human votes on 328 real requests.","keywords":["image editing","generative AI","human evaluation","VLM as judge","Reddit dataset","request taxonomy","identity preservation","perceptual quality"],"falsifier":"Conduct a new human preference study on a larger random sample of requests drawn from the same PSR corpus, stratified by editing action to match the corpus distribution, and compute the action-weighted AI win-plus-tie rate; if the result falls clearly outside 30–36%, the paper's headline estimate does not generalize.","tokens_in":38893,"feed_emoji":"🖼️","tokens_out":6348,"duration_ms":48589,"temperature":0.7,"pith_summary":"The paper analyzes 83,000 real image-editing requests and 305,000 human-made edits collected from a large Reddit community over 12 years, together with edits produced by 49 AI editors. By running a controlled human preference study on 328 of these requests, it claims that human judges prefer the human-made edits over the AI edits 66% of the time, and that after weighting by how often each editing action appears in the corpus, only 33.35% of real-world requests can be satisfactorily handled by the best current AI editors. The paper further claims that AI editors are weakest on low-creativity, precise requests, that they frequently change the identity of people and animals, and that they make unrequested aesthetic touch-ups. Finally, it claims that vision-language model judges disagree sharply with human judges, so automated VLM ratings are not a reliable proxy for human preference.","feed_headline":"AI handles only 1 in 3 real photo-editing requests","feed_subtitle":"Human raters prefer human edits 66% of the time; identity drift and unsolicited touch-ups are the biggest AI failures.","key_machinery":"The load-bearing object is the PSR dataset itself: 82,976 real-world requests with 305,806 human edits, annotated with 15 user-intent editing actions, WordNet-based subject labels, and three creativity levels. The quantitative estimate is produced by a stratified sample (PSR-328) balanced across creativity levels, pairwise human and VLM judgments between human and AI edits, and a weighting formula: overall handleable percentage equals the sum over actions of the action's frequency in the full corpus times the AI win-plus-tie rate for that action.","core_discovery":"On its own terms, the paper's central discovery is an estimate of the current ceiling of automated image editing on real user requests: human raters prefer human edits over AI edits 66.0% of the time, and the weighted AI win-plus-tie rate across all editing actions is 33.35%. The per-action breakdown shows AI editors succeed most often on open-ended actions such as 'add', 'apply', and 'merge' and least often on spatially precise actions such as 'zoom', 'crop', and 'move'. A separate controlled experiment shows that both GPT-4o and Gemini-2.0-Flash gradually change a person's facial identity and body shape over a sequence of simple shirt-color edits, measured by growing DINOv2 feature distance, confirming identity preservation as a systematic weakness.","pith_inferences":["Beyond the paper: the 33.35% figure is a snapshot of models available in early 2025; rerunning the same protocol on newer editors would give a direct measure of progress in closing the largest gaps.","Beyond the paper: the per-action weighting rests on the PSR-328 subset's action mix, and a study sampled to match the corpus's action frequencies could produce a different overall estimate, so the number should be treated as a point estimate rather than a tight bound.","Beyond the paper: the observed 'polish bias'—AI edits raising aesthetic scores even when not requested—might be part of why VLM judges over-prefer AI edits, since aesthetic quality is easier for a VLM to verify than faithfulness to the request.","Beyond the paper: the taxonomy of 15 intent-level actions could be reused as a template for building future training sets that match real user request distributions."],"forward_implications":["If 33.35% is accurate, current text-to-image editors are not close to replacing human editors on general real-world requests, and benchmarks built from synthetic requests likely overstate real capability.","Identity preservation and avoidance of unrequested changes are the two highest-leverage targets for improving AI editors.","AI editors are comparatively more useful for open-ended, high-creativity requests than for precise, low-creativity ones.","VLM judges should not be used as the primary evaluator for image editing without human calibration, since they can exhibit strong model-specific biases (e.g., preferring one editor's output up to 85% of the time)."],"supporting_citations":[{"why":"Provides the synthetic-request baseline that PSR contrasts with, showing that prior datasets rely on made-up requests; also a reference for the action distribution in AI training sets.","marker":"[4]"},{"why":"The concurrent Reddit-derived dataset RealEdit that PSR compares to, establishing the scale and novelty of PSR's 83k requests and creativity labels.","marker":"[47]"},{"why":"One of the three SOTA AI editors (SeedEdit) whose edits are compared against human edits in the human preference study.","marker":"[45]"},{"why":"One of the three SOTA AI editors (GPT-4o image generation) whose edits are compared against human edits.","marker":"[35]"},{"why":"One of the three SOTA AI editors (Gemini-2.0-Flash) whose edits are compared against human edits.","marker":"[21]"},{"why":"The o1 VLM judge whose preference for one AI editor's edits (85%) is contrasted with human preferences to show VLM bias.","marker":"[33]"},{"why":"The Gemini-2.0-Flash-Thinking VLM judge used in the VLM-as-judge experiment.","marker":"[14]"},{"why":"WordNet supplies the synset taxonomy used to standardize subject labels in the PSR dataset.","marker":"[12]"},{"why":"Prior editing-action taxonomy tied to software functions, which the paper's user-intent taxonomy is designed to improve upon.","marker":"[31]"}],"fun_headline_variants":["AI editors succeed on only 33% of real requests","Human editors beat AI in two-thirds of real image edits","AI photo editing fails on precise, identity-critical requests","Real-world photo requests: AI wins just a third of the time","Identity drift: Why AI editors fail at simple photo fixes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 33.35% figure assumes that the 328 requests selected for human evaluation represent every type of editing request in the full 83k corpus in proportions that allow the per-action success rates to be reweighted into a global estimate.","fun_headline_variants_meta":{"raw":{"variants":["AI editors succeed on only 33% of real requests","Human editors beat AI in two-thirds of real image edits","AI photo editing fails on precise, identity-critical requests","Real-world photo requests: AI wins just a third of the time","Identity drift: Why AI editors fail at simple photo fixes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000848,"raw_usage":{"total_tokens":3710,"prompt_tokens":984,"completion_tokens":2726,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":600,"completion_tokens_details":{"reasoning_tokens":2644}},"tokens_in":600,"tokens_out":2726,"duration_ms":17702,"temperature":1.0,"reasoning_tokens":2644,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T15:05:14.503505+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Conduct a new human preference study on a larger random sample of requests drawn from the same PSR corpus, stratified by editing action to match the corpus distribution, and compute the action-weighted AI win-plus-tie rate; if the result falls clearly outside 30–36%, the paper's headline estimate does not generalize.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the synthetic-request baseline that PSR contrasts with, showing that prior datasets rely on made-up requests; also a reference for the action distribution in AI training sets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The concurrent Reddit-derived dataset RealEdit that PSR compares to, establishing the scale and novelty of PSR's 83k requests and creativity labels."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"One of the three SOTA AI editors (GPT-4o image generation) whose edits are compared against human edits."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"One of the three SOTA AI editors (Gemini-2.0-Flash) whose edits are compared against human edits."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The o1 VLM judge whose preference for one AI editor's edits (85%) is contrasted with human preferences to show VLM bias."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The Gemini-2.0-Flash-Thinking VLM judge used in the VLM-as-judge experiment."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"WordNet supplies the synset taxonomy used to standardize subject labels in the PSR dataset."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior editing-action taxonomy tied to software functions, which the paper's user-intent taxonomy is designed to improve upon."}],"review_version":1}