{"id":"224a02d0-5cea-4509-91bf-3b2a80d00017","arxiv_id":"2504.20340","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Human-driven iterative prompt refinement improves image regeneration similarity in a 20-person study, while image similarity metrics show only moderate agreement with human rankings.","lead":"This paper ran a 20-person study where participants tried to recreate target AI images by repeatedly editing their text prompts, and found that later iterations produced images judged closer to the target. It also found that automated image-similarity scores matched human judgment only moderately, so they should be used with caution as feedback.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Recency/order bias in post-session rankings may explain the strong preference for late iterations; the subjective pillar of the central claim is not secure without a blinded re-ranking test.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the human-ranking evidence could reflect recency/effort bias rather than genuine similarity improvement. This is the most critical issue because the paper's headline claim is supported by 'both subjective evaluations and quantitative measures,' but the subjective measure is the one showing the largest apparent effect (44 of 150 top picks at iteration 10 vs. 15 expected). The objective mixed-effects model, while statistically significant, shows only small effect sizes and depends on the post-hoc exclusion of ImageHash, which removes a quarter of participants. If the subjective ranking is biased, the central claim loses its strongest support and the paper's contribution is substantially weakened. The authors themselves flag the fixed iteration count as a potential source of bias, which reinforces that this is a genuine threat rather than a manufactured concern. A concrete blinded re-ranking test would settle the issue: if the late-iteration preference persists when presentation order is randomized and iteration labels are hidden, the concern does not land; if it disappears, the human-ranking evidence must be discounted. Since the paper already holds a CONDITIONAL verdict based on this and related limitations, my read does not change that verdict, but it confirms that the condition is essential.","tokens_in":14218,"tokens_out":8764,"duration_ms":93042,"concrete_test":"Re-run the ranking task with a new set of participants (or independent raters) who are shown the 10 generated images for each target in a fully randomized order, without labels indicating which iteration produced them, and then asked to pick the most similar image. Compare the distribution of chosen iteration numbers to the uniform expectation. If the preference for iterations 9 and 10 disappears or sharply decreases under randomized/blinded presentation, the recency-bias concern is confirmed and Section 5.2's chi-square result should not be treated as evidence for Hypothesis 2.3. Even re-analyzing the existing data by recording the initial display order and testing whether top picks correlate with order would be a useful first step.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest human-centric evidence for the central claim is the chi-square test on top-ranked images (Section 5.2, Table 5), where 65 of 150 top picks came from iterations 9 and 10. These rankings were made after the participant completed all 10 iterations for a target image, with images likely shown in chronological order. Recency or position bias could inflate the apparent preference for later iterations; the authors acknowledge in Section 6 that the fixed iteration count 'may have introduced potential biases in the later iterations.' The chi-square test also ignores clustering (150 choices from 15 participants after ImageHash exclusion), so even the reported significance is not robustly established. The objective mixed-effects model does not independently rescue the claim: it relies on the same post-hoc exclusion, shows small improvements (iteration 1 coefficient -0.053 on a 0-1 scale), and does not control for the possibility that later iterations simply benefit from more sampling rather than genuine refinement. Thus the conclusion that 'iterative prompt refinement substantially enhances alignment... particularly in the early stages' hinges on whether users' top-ranked images are chosen for similarity or because they saw those images last.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a user study (n=20) of 'image regeneration,' in which participants iteratively edit prompts over 10 iterations per target image to recreate a target visual, with one of four image similarity metrics (ISMs) available and visible for half of the targets. The authors evaluate (RQ1) whether ISMs align with human similarity rankings via ICC, and (RQ2) whether iterative refinement improves similarity via a linear mixed-effects model on an adjusted ISM score and via a chi-square test on the iteration of users' top-ranked images. PS and CLIP variants show moderate ICC; ImageHash is excluded. The mixed model finds a significant iteration effect, with early iterations differing from iteration 10 and a plateau after iteration 6; the top-ranked images concentrate in iterations 9 and 10. The paper concludes that iterative prompt refinement substantially enhances alignment, especially early, and that select ISMs can act as feedback proxies.","tokens_in":14429,"tokens_out":7823,"duration_ms":83902,"significance":"If the findings are robust, this is a useful empirical contribution to an underexplored area: human-driven prompt refinement for image regeneration, with practical implications for novice prompt engineering, art restoration, and educational feedback tools. The paper has notable strengths: a structured within-subject design (10 iterations × 10 targets), a stand-alone human ranking measure, and an explicit limitations section. The central claims, however, rest on three fragile pillars: the post-hoc exclusion of the ImageHash condition, the possibly order-biased and clustered human-ranking test, and small objective effect sizes relative to the word 'substantially.' These are fixable with additional analyses, so the paper's contribution is promising but not yet established.","major_comments":[{"comment":"The decision to exclude ImageHash (ICC = 0.250) was made after observing the same data that feed the main analysis. Because participants were randomly assigned to metrics, dropping the ImageHash condition removes 25% of the sample (5 of 20) and breaks the per-metric balance; if the ImageHash-assigned participants happened to differ in iteration behavior, the mixed model and chi-square results in Section 5.2 are no longer from a randomized comparison. Please report a sensitivity analysis that retains ImageHash or otherwise show that the conclusions are unchanged, and either specify the ICC threshold a priori or clearly frame the main analysis as exploratory.","section":"Section 5.1, Table 1"},{"comment":"The chi-square test treats 150 top-ranked choices as independent, but they come from 15 participants with 10 choices each; ignoring this clustering can inflate significance. In addition, the manuscript does not report whether the ranking display order was randomized; if images were shown chronologically, recency or position bias could inflate the counts at iterations 9 and 10. The authors acknowledge this possibility in Section 6 ('may have introduced potential biases in the later iterations'). Please provide a cluster-robust analysis (e.g., a mixed-effects multinomial or logistic model with participant as a random effect) and, if feasible, a blinded re-ranking in randomized order.","section":"Section 5.2, Table 5"},{"comment":"The direction of the adjusted score is stated inconsistently. Appendix C defines the adjusted ISM score so that 'higher score meaning better similarity,' but the text describes the negative coefficients for early iterations as 'improved' and as 'significantly lower (improved) adjusted scores.' Under the stated definition, a negative coefficient relative to iteration 10 means the earlier iteration has a lower score, which is worse, not better. Please correct either the definition or the interpretation; as written, the objective evidence for Hypothesis 2.1 is internally contradictory and difficult for a reader to verify.","section":"Section 5.2, Table 3 and Appendix C"},{"comment":"Even after correcting the sign, the objective effect is modest: the cumulative difference between iteration 1 and iteration 10 is about 0.053 on a normalized 0–1 scale, or roughly 8.5% of the reference mean (0.620). The conclusion that iterative refinement 'substantially enhances alignment' is stronger than the data support; a more precise statement would describe a statistically significant but modest improvement concentrated in the early iterations.","section":"Section 5.2, Table 3 and Conclusion"}],"minor_comments":[{"comment":"The word 'subject' is used both for human participants and for the target prompt content (e.g., cat, astronaut); Table 2's 'subject' fixed effect should be renamed to 'target subject' to avoid ambiguity.","section":"Throughout"},{"comment":"There are duplicated words that should be corrected: 'may have have introduced' and 'by by trends'.","section":"Section 6"},{"comment":"The phrase 'how such content are inspired and generated' should be 'how such content is inspired and generated' (or 'such contents are').","section":"Abstract"},{"comment":"The manuscript does not state whether data or code will be made available; a data availability statement would improve reproducibility.","section":"Data Availability"},{"comment":"The reference to Mañas et al. contains an unnormalized tilde glyph in the author name; please ensure the LaTeX/PDF rendering is correct.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The main concerns are all addressable in revision: the post-hoc exclusion needs a sensitivity analysis, the human-ranking test needs a cluster-robust and order-controlled version, and the sign/effect-size reporting needs to be made consistent. I would also encourage the editor to ask for a data availability statement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's my take on arXiv:2504.20340. The paper is a small user study (20 participants) on human-driven iterative prompt refinement for image regeneration. It asks whether iterative prompt edits improve similarity to a target image, and whether image similarity metrics (ISMs) agree with human rankings. The results point in the expected direction, but the headline claim that refinement 'substantially enhances alignment' is stronger than the evidence supports.\n\nWhat's genuinely useful: this is one of the few studies that puts the human in the loop for multiple iterations, rather than using an automated prompt optimizer. The authors also test ISM-human alignment empirically with ICC instead of assuming it. Their mixed-effects model is a reasonable way to handle the nested, repeated-measures design, and they are transparent about their choices, including the post-hoc exclusion of ImageHash after it showed poor ICC.\n\nThe soft spots are in three places. First, the exclusion of ImageHash happens after seeing the low ICC, and it removes a quarter of the participants from the main analysis. That's a data-dependent decision that inflates the apparent consistency of the remaining metrics. Second, the objective effect sizes are small: the largest coefficient is -0.053 on a 0-1 scale, which is about five percentage points, and the paper's 'substantially' is doing a lot of work. Third, the human-ranking evidence—users picking later images as best, with 44 of 150 top picks at iteration 10—is vulnerable to order/recency bias because participants ranked all ten images after the session, in what appears to be chronological order. The chi-square test also treats 150 choices as independent when they come from 15 participants and 10 target images each. The authors acknowledge the fixed-iteration bias in their limitations, but they still lean on this result in the conclusion.\n\nThe paper is honest and clearly written, and it makes a modest empirical contribution to prompt engineering and human-AI interaction. It does not release data or code, which is a shame for a study this small. I'd send it to peer review—the question is worthwhile and the methodology is basically sound—but I'd expect the reviewers to ask for a softer conclusion, or a re-analysis that accounts for clustering and order effects.\n\nIf you're working on prompt optimization tools, this is worth knowing about, but I wouldn't treat it as strong evidence on its own.","headline":"A small, honest user study showing iterative human prompt refinement helps in image regeneration, but the 'substantial' claim outruns the small effects and order-biased rankings.","tokens_in":14945,"tokens_out":2651,"would_cite":false,"duration_ms":26259,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Iterative human-driven prompt refinement substantially improves how closely AI-generated images match a target visual, with gains concentrated in the early iterations and moderate agreement between perceptual/CLIP similarity metrics and…","keywords":["iterative prompt refinement","image regeneration","text-to-image generation","image similarity metrics","human-AI collaboration","prompt engineering","user study","perceptual similarity"],"falsifier":"A replication that presents the ten generated images in a randomized order, or ranks them against the target one at a time, should still show users disproportionately selecting later-iteration images if the paper's human-centric claim is right; if the preference for iterations 9 and 10 weakens or disappears, the subjective-improvement evidence would be substantially overstated.","tokens_in":14052,"feed_emoji":"🎨","tokens_out":5638,"duration_ms":50874,"temperature":0.7,"pith_summary":"The paper claims that when a person tries to recreate a specific target image with a text-to-image model, iteratively refining the prompt yields images that are measurably and perceptibly closer to the target, with the largest gains in the first several iterations. It also claims that common image similarity metrics, Perceptual Similarity and two CLIP variants, agree with human similarity judgments at a moderate level, while an image-hash metric does not. If true, ordinary users can close much of the gap to a desired visual through repeated prompt edits, and objective metrics can serve as rough feedback in such workflows.","feed_headline":"Iterative prompt edits measurably close the gap to any target image","feed_subtitle":"Twenty users tried to recreate target images; most score gains came in the first six rounds of prompt edits.","key_machinery":"The central mechanism is the iterative prompt-refinement loop: the user inspects the target image, writes a prompt, generates an image, compares the output to the target, and edits the prompt for the next round. The study measures this loop with three tools: Intraclass Correlation Coefficient (ICC) to quantify agreement between each ISM's ranking and human rankings, a linear mixed-effects model with an AR(1) residual structure to test how iteration and other factors change adjusted ISM scores, and a chi-square goodness-of-fit test on the iteration from which each user's top-ranked image came. The ISMs themselves are defined by their comparison machinery: Perceptual Similarity (LPIPS) compares deep CNN feature maps, CLIP B32 and L14 compare image embeddings, and ImageHash compares Hamming distances between perceptual hashes.","core_discovery":"The study's central discovery is that human-driven iterative prompt refinement improves image-regeneration alignment, and that the improvement appears both in objective similarity scores and in users' own rankings. In a study with 20 participants, 10 target images per participant, and 10 iterations per image (2,000 prompts total), the mixed-effects model found significant score gains for iterations 1 through 6 relative to iteration 10, after which gains were no longer statistically significant. Users' top-ranked images came disproportionately from the last two iterations, with 44 of 150 top choices at iteration 10. Intraclass correlation coefficients placed Perceptual Similarity at 0.686, CLIP B32 at 0.620, and CLIP L14 at 0.527, all moderate, while ImageHash scored 0.250. The authors interpret these results as evidence that iterative refinement works, that gains plateau around the seventh iteration, and that only some ISMs are trustworthy proxies for human perception.","pith_inferences":["Beyond the paper: if the fixed ten-iteration design introduced recency bias, the disproportionate preference for iterations 9 and 10 may overstate true perceptual gains; a replication with randomly ordered or pairwise rankings would separate genuine improvement from a last-seen effect.","Beyond the paper: the plateau after iteration 6 suggests an optimal-stopping rule for image-regeneration tools, where the system could signal users when further edits are unlikely to pay off and save time and compute.","Beyond the paper: the moderate ICC values imply that ISM-guided feedback should be presented as suggestions rather than verdicts, and richer feedback such as localized visual differences may be more helpful than a single aggregate score.","Beyond the paper: because the visibility of the ISM score did not alter improvement, future designs might test adaptive feedback, such as showing which regions of the image are most dissimilar, instead of repeating the same global score each round."],"forward_implications":["Users who iterate on prompts rather than relying on a single attempt move consistently closer to a target image, as measured by Perceptual Similarity and CLIP scores.","Most measurable improvement occurs in iterations 1 through 6; beyond that, additional prompt edits yield no statistically significant score gains, implying a practical plateau.","Users subjectively favor later iterations, with the most-similar image most often coming from iteration 9 or 10, confirming that the perceived benefit of iteration matches the objective trend.","Perceptual Similarity and the two CLIP variants can serve as moderate proxies for human similarity judgment in iterative workflows, but ImageHash should not be used this way.","Providing the numeric ISM score during the task did not change the rate of improvement, so simply showing a metric is not enough to boost user performance."],"supporting_citations":[{"why":"Supplies the Stable Diffusion 3.0 model used to generate both target images and user-generated images in the study.","marker":"[Esser and others, 2024]"},{"why":"Defines the Perceptual Similarity (LPIPS) metric and provides evidence that deep-feature distances track human judgments, making it the study's best-aligned ISM.","marker":"[Zhang et al., 2018]"},{"why":"Defines the CLIP image-embedding model used to compute the B32 and L14 similarity scores.","marker":"[OpenAI, 2021]"},{"why":"Provides the ImageHash implementation used as the fourth image similarity metric, which the study finds poorly aligned with human rankings.","marker":"[Buchner, 2024]"},{"why":"Supplies the ICC interpretation thresholds (poor, moderate, good, excellent) used to classify the metrics' agreement with human raters.","marker":"[Koo and Li, 2016]"},{"why":"Prior single-shot prompt-inference study that this work extends by allowing multiple iterations of prompt refinement.","marker":"[Trinh et al., 2024]"},{"why":"Automatic prompt-optimization framework (OPT2I) that establishes a baseline for how iterative refinement can improve text-to-image alignment.","marker":"[Ma˜nas et al., 2024]"},{"why":"Capability-aware prompt reformulation work that motivates the study's focus on user-driven refinement and its comparison with AI-led prompt adjustment.","marker":"[Zhan et al., 2024]"}],"fun_headline_variants":["Prompt tweaks boost image match, gains plateau by iteration 7","Iterative prompts improve regeneration, first 6 rounds key","Human prompt edits: score gains stop after 6 rounds","Perceptual Similarity best matches human judgment in prompt loops","2,000 prompt edits show regeneration gains plateau early"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The human-ranking evidence assumes that users' top-ranked images reflect genuine similarity improvement rather than a preference for whichever image was generated last, since every session used a fixed ten iterations and users ranked all ten images only after the session ended.","fun_headline_variants_meta":{"raw":{"variants":["Prompt tweaks boost image match, gains plateau by iteration 7","Iterative prompts improve regeneration, first 6 rounds key","Human prompt edits: score gains stop after 6 rounds","Perceptual Similarity best matches human judgment in prompt loops","2,000 prompt edits show regeneration gains plateau early"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000614,"raw_usage":{"total_tokens":2870,"prompt_tokens":979,"completion_tokens":1891,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":595,"completion_tokens_details":{"reasoning_tokens":1809}},"tokens_in":595,"tokens_out":1891,"duration_ms":12758,"temperature":1.0,"reasoning_tokens":1809,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:31:04.591942+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication that presents the ten generated images in a randomized order, or ranks them against the target one at a time, should still show users disproportionately selecting later-iteration images if the paper's human-centric claim is right; if the preference for iterations 9 and 10 weakens or disappears, the subjective-improvement evidence would be substantially overstated.","supporting_citations":[{"cited_title":"Stable diffusion 3: re- search paper–stability ai","cited_arxiv_id":null,"evidence_quote":"Supplies the Stable Diffusion 3.0 model used to generate both target images and user-generated images in the study."},{"cited_title":"The unreasonable effectiveness of deep features as a perceptual metric","cited_arxiv_id":null,"evidence_quote":"Defines the Perceptual Similarity (LPIPS) metric and provides evidence that deep-feature distances track human judgments, making it the study's best-aligned ISM."},{"cited_title":"https: //openai.com/research/clip,","cited_arxiv_id":null,"evidence_quote":"Defines the CLIP image-embedding model used to compute the B32 and L14 similarity scores."},{"cited_title":"Imagehash","cited_arxiv_id":null,"evidence_quote":"Provides the ImageHash implementation used as the fourth image similarity metric, which the study finds poorly aligned with human rankings."},{"cited_title":"A guide- line of selecting and reporting intraclass correlation co- efficients for reliability research","cited_arxiv_id":null,"evidence_quote":"Supplies the ICC interpretation thresholds (poor, moderate, good, excellent) used to classify the metrics' agreement with human raters."},{"cited_title":"Promptly Yours? A Human Subject Study on Prompt Inference in AI-Generated Art","cited_arxiv_id":"2410.08406","evidence_quote":"Prior single-shot prompt-inference study that this work extends by allowing multiple iterations of prompt refinement."},{"cited_title":"Capability-aware prompt refor- mulation learning for text-to-image generation","cited_arxiv_id":null,"evidence_quote":"Capability-aware prompt reformulation work that motivates the study's focus on user-driven refinement and its comparison with AI-led prompt adjustment."}],"review_version":1}