REVIEW 4 major objections 5 minor 33 references
A Picture is Worth a Thousand Prompts? Efficacy of Iterative Human-Driven Prompt Refinement in Image Regeneration Tasks
T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Iterative human-driven prompt refinement substantially improves how closely AI-generated images match a target visual, with gains concentrated in the early iterations and moderate agreement between perceptual/CLIP similarity metrics and…
desk verdict A small, honest user study showing iterative human prompt refinement helps in image regeneration, but the 'substantial' claim outruns the small effects and order-biased rankings. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the iterative prompt-refinement loop: the user inspects the target image, writes a prompt, generates an image, compares the output to the target, and edits the prompt for the next round. The study measures this loop with three tools: Intraclass Correlation Coefficient (ICC) to quantify agreement between each ISM's ranking and human rankings, a linear mixed-effects model with an AR(1) residual structure to test how iteration and other factors change adjusted ISM scores, and a chi-square goodness-of-fit test on the iteration from which each user's top-ranked image came. The ISMs themselves are defined by their comparison machinery: Perceptual Similarity (LPIPS) compares deep CNN feature maps, CLIP B32 and L14 compare image embeddings, and ImageHash compares Hamming distances between perceptual hashes.
What would settle it
A replication that presents the ten generated images in a randomized order, or ranks them against the target one at a time, should still show users disproportionately selecting later-iteration images if the paper's human-centric claim is right; if the preference for iterations 9 and 10 weakens or disappears, the subjective-improvement evidence would be substantially overstated.
Extended reading notes
Core claim
The study's central discovery is that human-driven iterative prompt refinement improves image-regeneration alignment, and that the improvement appears both in objective similarity scores and in users' own rankings. In a study with 20 participants, 10 target images per participant, and 10 iterations per image (2,000 prompts total), the mixed-effects model found significant score gains for iterations 1 through 6 relative to iteration 10, after which gains were no longer statistically significant. Users' top-ranked images came disproportionately from the last two iterations, with 44 of 150 top choices at iteration 10. Intraclass correlation coefficients placed Perceptual Similarity at 0.686, CLIP B32 at 0.620, and CLIP L14 at 0.527, all moderate, while ImageHash scored 0.250. The authors interpret these results as evidence that iterative refinement works, that gains plateau around the seventh iteration, and that only some ISMs are trustworthy proxies for human perception.
Load-bearing premise
The human-ranking evidence assumes that users' top-ranked images reflect genuine similarity improvement rather than a preference for whichever image was generated last, since every session used a fixed ten iterations and users ranked all ten images only after the session ended.
Editorial extensions
If this is right
- Users who iterate on prompts rather than relying on a single attempt move consistently closer to a target image, as measured by Perceptual Similarity and CLIP scores.
- Most measurable improvement occurs in iterations 1 through 6; beyond that, additional prompt edits yield no statistically significant score gains, implying a practical plateau.
- Users subjectively favor later iterations, with the most-similar image most often coming from iteration 9 or 10, confirming that the perceived benefit of iteration matches the objective trend.
- Perceptual Similarity and the two CLIP variants can serve as moderate proxies for human similarity judgment in iterative workflows, but ImageHash should not be used this way.
- Providing the numeric ISM score during the task did not change the rate of improvement, so simply showing a metric is not enough to boost user performance.
Reading between the lines
- Beyond the paper: if the fixed ten-iteration design introduced recency bias, the disproportionate preference for iterations 9 and 10 may overstate true perceptual gains; a replication with randomly ordered or pairwise rankings would separate genuine improvement from a last-seen effect.
- Beyond the paper: the plateau after iteration 6 suggests an optimal-stopping rule for image-regeneration tools, where the system could signal users when further edits are unlikely to pay off and save time and compute.
- Beyond the paper: the moderate ICC values imply that ISM-guided feedback should be presented as suggestions rather than verdicts, and richer feedback such as localized visual differences may be more helpful than a single aggregate score.
- Beyond the paper: because the visibility of the ISM score did not alter improvement, future designs might test adaptive feedback, such as showing which regions of the image are most dissimilar, instead of repeating the same global score each round.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a user study (n=20) of 'image regeneration,' in which participants iteratively edit prompts over 10 iterations per target image to recreate a target visual, with one of four image similarity metrics (ISMs) available and visible for half of the targets. The authors evaluate (RQ1) whether ISMs align with human similarity rankings via ICC, and (RQ2) whether iterative refinement improves similarity via a linear mixed-effects model on an adjusted ISM score and via a chi-square test on the iteration of users' top-ranked images. PS and CLIP variants show moderate ICC; ImageHash is excluded. The mixed model finds a significant iteration effect, with early iterations differing from iteration 10 and a plateau after iteration 6; the top-ranked images concentrate in iterations 9 and 10. The paper concludes that iterative prompt refinement substantially enhances alignment, especially early, and that select ISMs can act as feedback proxies.
Significance. If the findings are robust, this is a useful empirical contribution to an underexplored area: human-driven prompt refinement for image regeneration, with practical implications for novice prompt engineering, art restoration, and educational feedback tools. The paper has notable strengths: a structured within-subject design (10 iterations × 10 targets), a stand-alone human ranking measure, and an explicit limitations section. The central claims, however, rest on three fragile pillars: the post-hoc exclusion of the ImageHash condition, the possibly order-biased and clustered human-ranking test, and small objective effect sizes relative to the word 'substantially.' These are fixable with additional analyses, so the paper's contribution is promising but not yet established.
major comments (4)
- [Section 5.1, Table 1] The decision to exclude ImageHash (ICC = 0.250) was made after observing the same data that feed the main analysis. Because participants were randomly assigned to metrics, dropping the ImageHash condition removes 25% of the sample (5 of 20) and breaks the per-metric balance; if the ImageHash-assigned participants happened to differ in iteration behavior, the mixed model and chi-square results in Section 5.2 are no longer from a randomized comparison. Please report a sensitivity analysis that retains ImageHash or otherwise show that the conclusions are unchanged, and either specify the ICC threshold a priori or clearly frame the main analysis as exploratory.
- [Section 5.2, Table 5] The chi-square test treats 150 top-ranked choices as independent, but they come from 15 participants with 10 choices each; ignoring this clustering can inflate significance. In addition, the manuscript does not report whether the ranking display order was randomized; if images were shown chronologically, recency or position bias could inflate the counts at iterations 9 and 10. The authors acknowledge this possibility in Section 6 ('may have introduced potential biases in the later iterations'). Please provide a cluster-robust analysis (e.g., a mixed-effects multinomial or logistic model with participant as a random effect) and, if feasible, a blinded re-ranking in randomized order.
- [Section 5.2, Table 3 and Appendix C] The direction of the adjusted score is stated inconsistently. Appendix C defines the adjusted ISM score so that 'higher score meaning better similarity,' but the text describes the negative coefficients for early iterations as 'improved' and as 'significantly lower (improved) adjusted scores.' Under the stated definition, a negative coefficient relative to iteration 10 means the earlier iteration has a lower score, which is worse, not better. Please correct either the definition or the interpretation; as written, the objective evidence for Hypothesis 2.1 is internally contradictory and difficult for a reader to verify.
- [Section 5.2, Table 3 and Conclusion] Even after correcting the sign, the objective effect is modest: the cumulative difference between iteration 1 and iteration 10 is about 0.053 on a normalized 0–1 scale, or roughly 8.5% of the reference mean (0.620). The conclusion that iterative refinement 'substantially enhances alignment' is stronger than the data support; a more precise statement would describe a statistically significant but modest improvement concentrated in the early iterations.
minor comments (5)
- [Throughout] The word 'subject' is used both for human participants and for the target prompt content (e.g., cat, astronaut); Table 2's 'subject' fixed effect should be renamed to 'target subject' to avoid ambiguity.
- [Section 6] There are duplicated words that should be corrected: 'may have have introduced' and 'by by trends'.
- [Abstract] The phrase 'how such content are inspired and generated' should be 'how such content is inspired and generated' (or 'such contents are').
- [Data Availability] The manuscript does not state whether data or code will be made available; a data availability statement would improve reproducibility.
- [References] The reference to Mañas et al. contains an unnormalized tilde glyph in the author name; please ensure the LaTeX/PDF rendering is correct.
Circularity Check
No significant circularity: the paper's claims rest on empirical measurements, not on definitions or self-citation chains.
full rationale
The paper does not derive its central result from an input that already contains it. The claim that iterative prompt refinement improves image-target alignment is assessed through two independent empirical channels: a linear mixed-effects model on objective ISM scores and a chi-square test on users' top-ranked images. The ISMs (Perceptual Similarity, CLIP B32, CLIP L14, ImageHash) are external, pre-existing metrics, and the paper first tests their alignment with human rankings via ICC before using the moderate-agreement metrics in subsequent analysis; this is a validation step, not a definitional equivalence. The human-centric top-ranked-image analysis does not use ISM scores at all, so it cannot be forced by the mixed-effects model. The paper's self-citations (e.g., Trinh et al. 2024) are used for dataset construction methodology and background context, not as load-bearing proof of the main finding; no uniqueness theorem or ansatz is imported from prior work. Acknowledged limitations such as the fixed iteration count and potential recency bias in rankings are threats to validity or generalizability, not circular reasoning. No step reduces by construction to its own inputs, so the appropriate score is 0.
Assumptions & free parameters
free parameters (1)
- ICC retention threshold =
0.5
assumptions (4)
- domain assumption Human subjective rankings are the appropriate ground truth for image similarity.
- domain assumption The mixed-effects model's AR(1) covariance and random intercepts correctly capture within-participant and within-prompt dependence.
- domain assumption Fixed seed and Stable Diffusion 3.0 parameters make target and user images comparable and representative of text-to-image generation.
- domain assumption Top-ranked-image choices across sessions are independent for the chi-square test.
Cite this review
Pith. "Pith review of A Picture is Worth a Thousand Prompts? Efficacy of Iterative Human-Driven Prompt Refinement in Image Regeneration Tasks." pith.science (2026). https://pith.science/paper/3DVJ57KW
@misc{pith2026250420340,
author = {Pith},
title = {Pith review of: A Picture is Worth a Thousand Prompts? Efficacy of Iterative Human-Driven Prompt Refinement in Image Regeneration Tasks},
year = {2026},
howpublished = {\url{https://pith.science/paper/3DVJ57KW}},
note = {Machine review of arXiv:2504.20340}
}
read the original abstract
With AI-generated content becoming ubiquitous across the web, social media, and other digital platforms, it is vital to examine how such content are inspired and generated. The creation of AI-generated images often involves refining the input prompt iteratively to achieve desired visual outcomes. This study focuses on the relatively underexplored concept of image regeneration using AI, in which a human operator attempts to closely recreate a specific target image by iteratively refining their prompt. Image regeneration is distinct from normal image generation, which lacks any predefined visual reference. A separate challenge lies in determining whether existing image similarity metrics (ISMs) can provide reliable, objective feedback in iterative workflows, given that we do not fully understand if subjective human judgments of similarity align with these metrics. Consequently, we must first validate their alignment with human perception before assessing their potential as a feedback mechanism in the iterative prompt refinement process. To address these research gaps, we present a structured user study evaluating how iterative prompt refinement affects the similarity of regenerated images relative to their targets, while also examining whether ISMs capture the same improvements perceived by human observers. Our findings suggest that incremental prompt adjustments substantially improve alignment, verified through both subjective evaluations and quantitative measures, underscoring the broader potential of iterative workflows to enhance generative AI content creation across various application domains.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Improv- ing image generation with better captions
[Betker et al., 2023] James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al. Improv- ing image generation with better captions. Computer Sci- ence. https://cdn. openai. com/papers/dall-e-3. pdf, 2(3):8,
2023
-
[6]
A cognitive process theory of writ- ing
[Flower, 1981] L Flower. A cognitive process theory of writ- ing. Composition and communication,
work page 1981
-
[8]
Gen- erative adversarial networks
[Goodfellow et al., 2020] Ian Goodfellow, Jean Pouget- Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Gen- erative adversarial networks. Communications of the ACM, 63(11):139–144,
work page 2020
-
[9]
Genassist: Making image generation accessible
[Huh et al., 2023] Mina Huh, Yi-Hao Peng, and Amy Pavel. Genassist: Making image generation accessible. In Pro- ceedings of the 36th Annual ACM Symposium on User In- terface Software and Technology, pages 1–17,
work page 2023
-
[10]
Human image generation: A comprehensive survey
[Jia et al., 2024] Zhen Jia, Zhang Zhang, Liang Wang, and Tieniu Tan. Human image generation: A comprehensive survey. ACM Computing Surveys, 56(11):1–39,
work page 2024
-
[11]
[Koo and Li, 2016] Terry K Koo and Mae Y Li. A guide- line of selecting and reporting intraclass correlation co- efficients for reliability research. Journal of chiropractic medicine, 15(2):155–163,
work page 2016
-
[14]
Iterative prompt learning for unsupervised backlit image enhance- ment
[Liang et al., 2023] Zhexin Liang, Chongyi Li, Shangchen Zhou, Ruicheng Feng, and Chen Change Loy. Iterative prompt learning for unsupervised backlit image enhance- ment. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 8094–8103,
work page 2023
-
[15]
A word is worth a thousand pictures: Prompts as ai design material
[Kulkarni et al., 2023] Chinmay Kulkarni, Stefania Druga, Minsuk Chang, Alex Fiannaca, Carrie Cai, and Michael Terry. A word is worth a thousand pictures: Prompts as ai design material. arXiv preprint arXiv:2303.12647,
arXiv 2023
Show all 33 references
-
[16]
Self-refine: Iterative refinement with self-feedback
[Madaan et al., 2024] Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems , 36,
2024
-
[17]
Improving text-to-image consistency via automatic prompt optimization
[Ma˜nas et al., 2024] Oscar Ma ˜nas, Pietro Astolfi, Melissa Hall, Candace Ross, Jack Urbanek, Adina Williams, Aishwarya Agrawal, Adriana Romero-Soriano, and Michal Drozdzal. Improving text-to-image consistency via automatic prompt optimization. arXiv preprint arXiv:2403.17804,
2024 arXiv
-
[18]
Generative ai has a visual plagiarism problem, May
[Marcus and Southen, 2024] Gary Marcus and Reid Southen. Generative ai has a visual plagiarism problem, May
2024
-
[19]
Midjourney Model Ver- sions
[MidJourney, 2024] MidJourney. Midjourney Model Ver- sions. https://docs.midjourney.com/docs/model-versions,
2024
-
[20]
The consequence of ignoring a level of nesting in multilevel analysis
[Moerbeek, 2004] Mirjam Moerbeek. The consequence of ignoring a level of nesting in multilevel analysis. Multi- variate behavioral research, 39(1):129–149,
2004
-
[21]
Generative AI: the risks and the unknowns — oecd.ai
[OECD.AI, 2025] OECD.AI. Generative AI: the risks and the unknowns — oecd.ai. https://oecd.ai/en/genai/issues/ risks-and-unknowns,
2025
-
[24]
High-resolution image synthesis with latent diffusion models
[Rombach et al., 2022] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In IEEE/CVF CVPR,
2022
-
[25]
Image super-resolution via iterative refinement
[Saharia et al., 2022] Chitwan Saharia, Jonathan Ho, William Chan, Tim Salimans, David J Fleet, and Mo- hammad Norouzi. Image super-resolution via iterative refinement. IEEE transactions on pattern analysis and machine intelligence, 45(4):4713–4726,
2022
-
[26]
A perceptually based comparison of image similarity met- rics
[Sinha and Russell, 2011] Pawan Sinha and Richard Russell. A perceptually based comparison of image similarity met- rics. Perception, 40(11):1269–1281,
2011
-
[28]
Promptly yours? a human subject study on prompt inference in ai-generated art
[Trinh et al., 2024] Khoi Trinh, Joseph Spracklen, Raveen Wijewickrama, Bimal Viswanath, Murtuza Jadliwala, and Anindya Maiti. Promptly yours? a human subject study on prompt inference in ai-generated art. arXiv preprint arXiv:2410.08406,
2024 arXiv
-
[29]
Exploring clip for assessing the look and feel of images
[Wang et al., 2023] Jianyi Wang, Kelvin CK Chan, and Chen Change Loy. Exploring clip for assessing the look and feel of images. In AAAI Conference on Artificial In- telligence, volume 37,
2023
-
[30]
Capability-aware prompt refor- mulation learning for text-to-image generation
[Zhan et al., 2024] Jingtao Zhan, Qingyao Ai, Yiqun Liu, Jia Chen, and Shaoping Ma. Capability-aware prompt refor- mulation learning for text-to-image generation. In Pro- ceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retriev...
2024
-
[31]
The unreasonable effectiveness of deep features as a perceptual metric
[Zhang et al., 2018] Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In IEEE/CVF CVPR,
2018
-
[33]
Computer Science
• β6: Fixed effect of Metric Type (categorical) Random Effects: • bj: Random intercept for session id, capturing vari- ability across sessions, withbj∼N (0,σ 2 b,j) • bk: Random intercept for target prompt, capturing variability across prompts, withbk∼N (0,σ 2 b,k) Repeated Me...
-
[1981]
Shift-tolerant perceptual similarity metric
[Ghildyal and Liu, 2022] Abhijay Ghildyal and Feng Liu. Shift-tolerant perceptual similarity metric. In European Conference on Computer Vision, pages 91–107. Springer,
2022
-
[2004]
Prompt refinement or fine-tuning? best practices for using llms in computational social science tasks
[Møller and Aiello, 2024] Anders Giovanni Møller and Luca Maria Aiello. Prompt refinement or fine-tuning? best practices for using llms in computational social science tasks. arXiv preprint arXiv:2408.01346,
2024 arXiv
-
[2011]
Exploring the im- pact of ai-generated image tools on professional and non- professional users in the art and design fields
[Tang et al., 2024] Yuying Tang, Ningning Zhang, Mari- ana Ciancia, and Zhigang Wang. Exploring the im- pact of ai-generated image tools on professional and non- professional users in the art and design fields. In Com- panion Publication of the 2024 Conference on Computer- Sup...
2024
-
[2016]
Looks like it
[Krawetz, 2011] Neal Krawetz. Looks like it. https://www.hackerfactor.com/blog/index.php?/archives/ 432-Looks-Like-It.html, May
2011
-
[2018]
Con- trolled
A Chosen AI-Image Generation Model Both the target and user images are generated using the Stable Diffusion 3.0 model1 from Stability AI. We chose this model it was the most current model when we began experimen- tal design and setup, and has demonstrated improved perfor- manc...
2024
-
[2020]
Read, revise, repeat: A system demonstration for human-in-the-loop iterative text revision
[Du et al., 2022] Wanyu Du, Zae Myung Kim, Vipul Raheja, Dhruv Kumar, and Dongyeop Kang. Read, revise, repeat: A system demonstration for human-in-the-loop iterative text revision. In Proceedings of the First Workshop on Intelligent and Interactive Writing Assistants (In2Writi...
2022
-
[2021]
Prompting ai art: An investigation into the creative skill of prompt engineering
[Oppenlaender et al., 2024] Jonas Oppenlaender, Rhema Linder, and Johanna Silvennoinen. Prompting ai art: An investigation into the creative skill of prompt engineering. International Journal of Human–Computer Interaction , pages 1–23,
2024
-
[2022]
Stable diffusion 3: re- search paper–stability ai
[Esser and others, 2024] P Esser et al. Stable diffusion 3: re- search paper–stability ai. Stability AI,
2024
-
[2023]
Imagehash
[Buchner, 2024] Johannes Buchner. Imagehash. https://pypi. org/project/ImageHash/,
2024
-
[2024]
[Chan and Murphy, 2020] Cherice Chan and Dillon Murphy
Accessed: 2024-10-13. [Chan and Murphy, 2020] Cherice Chan and Dillon Murphy. Priming in action: How we are influenced without even knowing
2024
-
[2025]
https: //openai.com/research/clip,
[OpenAI, 2021] CLIP: Connecting text and images. https: //openai.com/research/clip,
2021
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.