REVIEW 3 major objections 5 minor 3 cited by
Stop Automating Peer Review Without Rigorous Evaluation
T0 review · 3 major / 5 minor · reviewed 2026-07-12 · grok-4.5
Pith's one-line read Today's AI systems should not write paper reviews: they collapse diversity and are trivial to game by rewriting style, not science.
desk verdict Solid empirical position paper: laundering is a clean, practical failure mode; hivemind is real but partly style-proxy, and the dual-condition case still holds under the tested setups. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two necessary conditions, C1 (preservation of review diversity) and C2 (resistance to gaming), operationalized by IntraSim/InterSim embedding similarities on reviews and by paper laundering—an automated zero-shot rewrite of the full LaTeX driven by prior AI feedback that raises subsequent AI scores without human oversight.
What would settle it
A controlled test showing that diverse prompts, temperatures, or model ensembles bring AI IntraSim/InterSim (especially on weaknesses and questions) down to human levels while laundering no longer raises scores across independent reviewer models—or blinded human experts rating original vs laundered papers as scientifically improved rather than cosmetically gamed.
Extended reading notes
Core claim
Current AI reviewers fail two necessary conditions for automating peer-review judgment: they do not preserve review diversity (the hivemind effect of excess agreement within and across papers, visible in ICLR 2026 data and in agent simulations) and they are not resistant to gaming (paper laundering: a single zero-shot LLM rewrite significantly boosts AI scores via stylistic and often hallucinated changes rather than scientific substance). Those conditions are necessary but not sufficient; full automation still requires community deliberation on accountability, legitimacy, and evaluation standards.
Load-bearing premise
That higher cosine similarity of review embeddings (and reused stock phrases) is enough evidence that AI reviewers collapse the plurality of expert perspectives peer review is meant to aggregate, rather than mostly sharing style or boilerplate.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that current AI systems should not produce paper reviews for acceptance-relevant judgment. It grounds the claim in two necessary conditions: C1 preservation of review diversity and C2 resistance to gaming. Empirically, it reports an AI reviewer 'hivemind'—higher within- and across-paper review similarity than humans—in 75,800 ICLR 2026 reviews (AI-generation labels from Emi 2025) and in controlled agent simulations on 60 papers (IntraSim +8.7% to +9.8%; InterSim +4.1% to +39.8%). It further shows 'paper laundering': zero-shot LLM rewrites of LaTeX papers raise AI review scores (overall mean +0.45 on a 1–10 scale across 24 prompt/model conditions) via largely stylistic or hallucinated edits, and increase pairwise paper similarity (+6.5%). The authors treat C1 and C2 as necessary but not sufficient, and call for a science of peer review automation with adversarial testing, accuracy validation, transparency, stakeholder studies, human–AI interaction research, and better reviewer incentives.
Significance. If the dual-failure case holds, the paper is a timely, high-stakes intervention for conference policy at a moment when venues are already piloting AI reviews and feedback agents. Strengths include large-scale observational evidence, multi-condition laundering grids with paired statistics and effect sizes, length-matched and area-stratified robustness, weaknesses/questions ablations, and a clear distinction between necessary and sufficient conditions rather than a blanket ban on all AI assistance. The concrete construct of paper laundering—policy-compliant, zero-shot, no hidden injection—is a useful contribution beyond prior prompt-injection work. The manuscript is well positioned to influence evaluation standards even if full automation remains contested.
major comments (3)
- [§3.2–3.4, Appendix A, §G.1] §3.2–3.4 and Appendix A: C1 is operationalized almost entirely via cosine similarity of text-embedding-3-small review vectors (IntraSim/InterSim) plus n-gram reuse. Higher similarity is interpreted as collapsed plurality of expert perspectives. Two reviews can be linguistically similar yet disagree on flaws, priorities, or accept/reject stance. Restricting to weaknesses/questions (§G.1) increases effect sizes but still uses the same proxy. For the dual-failure claim to carry, the manuscript needs either (i) a content-level analysis (e.g., coded critique categories, disagreement on specific claims, or accept/reject stance diversity) on a subset, or (ii) a clearly scoped claim that the measured gap is linguistic/stylistic homogenization with only partial support for argumentative monoculture. Score correlations and AUC (§3.5, Table 7) help but do not fully substitute.
- [§4.1, Appendix E, Table 5] §4.1 and Appendix E: Laundering score lifts are robust across 24 conditions, but the claim that gains reflect gaming rather than quality improvement rests on word-level style counts (Table 5) and manual inspection of five papers with ≥1-point gains (Appendix E.1). That sample is small and author-selected. A blinded human rating of original vs. laundered versions (clarity, rigor, acceptability) on a larger subset would substantially strengthen C2; without it, some score increases could be genuine presentation improvements that human reviewers might also reward. This is load-bearing for 'trivially gameable without scientific improvement.'
- [§3.4, Appendix B.1, Appendix A] §3.4, Appendix B.1, Appendix A: Simulation IntraSim/InterSim and laundering results use a small set of frontier models and a single fixed, highly structured ICLR-style XML review prompt. The authors acknowledge this, and the in-the-wild result partially mitigates it. Still, for the strong claim that 'today’s AI systems should not produce paper reviews,' the manuscript should either (i) report at least one diversification ablation (temperature, dissent-seeking prompt, multi-model ensemble) or (ii) more carefully bound the claim to current default agent setups rather than all possible LLM reviewing configurations. Without that, C1 in simulation may partly reflect experimental homogeneity.
minor comments (5)
- [§4.1, Figure 4, Appendix H] Figure 4 caption and §4.1 report overall mean +0.45; the pasted AI self-review in Appendix H cites +0.28. Align all reported laundering deltas with the multi-condition grid actually used.
- [Table 1] Table 1 uses mixed symbols for allowed/provided/prohibited LLM use; a short legend already exists but a one-line key in the caption would improve scanability.
- [§3.5, Appendix C, Table 7] §3.5 and Appendix C: AI score inflation and AI–AI correlation are useful; state clearly whether human scores are pre- or post-rebuttal when comparing predictive AUC to final decisions.
- [Appendix A, §3–4] Appendix A limitations are appropriately candid; consider moving a one-sentence summary of the embedding-proxy and n=60 limits into the main §3/§4 discussion so readers do not miss them.
- [§3.1, Appendix B] Minor consistency: 'ICLR 2026' data and model version strings (gpt-5.1, claude-sonnet-4-5) should be checked against public naming once camera-ready, and any third-party label version (Emi/Pangram) pinned for reproducibility.
Circularity Check
No circular derivation: empirical pre/post and human-vs-AI comparisons against external ICLR data; C1/C2 are normative criteria, not results forced by construction.
full rationale
This is a position paper with empirical measurements, not a first-principles derivation that could collapse into its inputs. C1 (hivemind) is operationalized as IntraSim/InterSim cosine similarity of review embeddings plus n-gram reuse, then compared to human ICLR 2026 reviews and controlled agent runs; higher AI similarity is an observed statistic, not a quantity defined to equal the claim. C2 (paper laundering) is a paired pre/post score experiment under zero-shot rewrite, with word-level and manual checks that edits are largely stylistic or hallucinated. Neither result is a fitted parameter renamed as a prediction, nor a uniqueness theorem imported from the authors, nor a self-definitional loop. Self-citations (e.g., prior Baumann et al. work on fairness/LLM annotation) are background and not load-bearing for the central empirical claims, which rest on ICLR reviews, third-party AI-generation labels, and external reviewer agents. Proxy validity of embeddings for argumentative diversity is a measurement limitation the paper itself flags (Appendix A), not circularity. Score 0 is the honest finding.
Assumptions & free parameters
free parameters (4)
- AI-agent review system prompt and fixed XML ICLR template
- Laundering objective prompts (4 variants) targeting score 10
- n=60 paper sample for simulation and laundering
- text-embedding-3-small cosine similarity as diversity metric
assumptions (5)
- ad hoc to paper Preservation of review diversity and resistance to gaming are necessary conditions for automating peer-review judgment relevant to acceptance.
- domain assumption High embedding similarity among reviews implies reduced perspective diversity / lower information gain from additional reviewers.
- domain assumption Fully automated textual rewrites without new experiments that raise AI scores constitute gaming rather than legitimate scientific improvement.
- domain assumption Third-party EditLens/Pangram labels sufficiently identify fully AI-generated ICLR reviews for in-the-wild InterSim comparisons.
- standard math Standard statistical tests (Welch t, Wilcoxon, Cohen’s d, AUC) on the reported samples support the claimed effect directions.
invented entities (2)
-
Paper laundering
independent evidence
-
AI reviewer hivemind effect (as peer-review construct)
independent evidence
Cite this review
Pith. "Pith review of Stop Automating Peer Review Without Rigorous Evaluation." pith.science (2026). https://pith.science/paper/IKM2OKOK
@misc{pith2026260503202,
author = {Pith},
title = {Pith review of: Stop Automating Peer Review Without Rigorous Evaluation},
year = {2026},
howpublished = {\url{https://pith.science/paper/IKM2OKOK}},
note = {Machine review of arXiv:2605.03202}
}
read the original abstract
Large language models offer a tempting solution to address the peer review crisis. This position paper argues that today's AI systems should not be used to produce paper reviews. We ground this position in an empirical comparison of human- versus AI-generated ICLR 2026 reviews and an evaluation of the effect of automated paper rewriting on different AI reviewers. We identify two critical issues: 1) AI reviewers exhibit a hivemind effect of excessive agreement within and across papers that reduces perspective diversity. 2) AI review scores are trivially gameable through paper laundering: prompting an LLM to rewrite a paper could significantly increase the scores from AI reviewers, demonstrating that LLM reviewers are easy to game through stylistic changes rather than scientific results. However, non-gameability and review diversity are necessary but not sufficient conditions for automation. We argue that addressing the peer review crisis requires a science of peer review automation -- not general-purpose LLMs deployed without rigorous evaluation.
Figures
Figures from the paper (9 more)
Forward citations
Cited by 3 Pith papers
-
ReVoicer: Conversational Voice Annotation for Human-Centered, LLM-Assisted Peer Review
ReVoicer is a prototype that turns spoken, in-the-moment reactions to a paper into cleaned, tagged annotations and a draft review aligned with the reviewer's own style.
-
Is ChatGPT as reliable as individual reviewers assessing the quality of published journal articles from PDFs or titles and abstracts?
ChatGPT-5.4's averaged scores rank journal articles about as reliably as individual expert reviewers, but full-text PDF input does not improve score accuracy over title/abstract input.
-
ReVoicer: Conversational Voice Annotation for Human-Centered, LLM-Assisted Peer Review
ReVoicer is a prototype annotation tool that cleans a reviewer's spoken/immediate reactions and drafts a rubric-aligned review using only the reviewer's own comments.
Reference graph
Works this paper leans on
-
[1]
URL https://asistdl.onlinelibrary.wiley
doi: https://doi.org/10.1002/asi.22784. URL https://asistdl.onlinelibrary.wiley. com/doi/abs/10.1002/asi.22784. Lee, H.-P. H., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., and Wilson, N. The impact of generative ai on critical thinking: Self-reported reduc- tions in cognitive effort and confidence effects from a 11 Stop Automating Peer...
-
[2]
URL https://aclanthology.org/2025. findings-emnlp.259/. Littman, M. L. Collusion rings threaten the integrity of computer science research.Commun. ACM, 64(6):43–44, May 2021. ISSN 0001-0782. doi: 10.1145/3429776. URLhttps://doi.org/10.1145/3429776. Liu, R. and Shah, N. B. Reviewergpt? an exploratory study on using large language models for paper reviewing...
arXiv doi:10.1145/3429776 2025
-
[3]
Pagan, N., Baumann, J., Elokda, E., De Pasquale, G., Bolognani, S., and Hann ´ak, A
URL https://www.nytimes.com/2015/ 06/26/upshot/can-an-algorithm-hire- better-than-a-human.html. Pagan, N., Baumann, J., Elokda, E., De Pasquale, G., Bolognani, S., and Hann ´ak, A. A classification of feedback loops and their relation to biases in auto- mated decision-making systems. InProceedings of the 3rd ACM Conference on Equity and Access in Algo- ri...
-
[4]
doi: 10.1145/3757667. URL https://doi. org/10.1145/3757667. 12 Stop Automating Peer Review Without Rigorous Evaluation Sahu, G., Larochelle, H., Charlin, L., and Pal, C. Reviewer- too: Should ai join the program committee? a look at the future of peer review.arXiv preprint arXiv:2510.08867, 2025. Schintler, L. A., McNeely, C. L., and Witte, J. A critical ...
doi:10.1145/3757667 2025
-
[5]
URL https://aclanthology.org/2025. findings-acl.1323/. Shah, N. B. Challenges, experiments, and computational solutions in peer review.Commun. ACM, 65(6):76–87, May 2022. ISSN 0001-0782. doi: 10.1145/3528086. URLhttps://doi.org/10.1145/3528086. Sharma, A., Rao, S., Brockett, C., Malhotra, A., Jojic, N., and Dolan, B. Investigating agency of LLMs in human-...
doi:10.1145/3528086 2025
-
[6]
URL https://aclanthology.org/2024. eacl-long.119/. Shcherbiak, A., Habibnia, H., B ¨ohm, R., and Fiedler, S. Evaluating science: A comparison of human and ai re- viewers.Judgment and Decision Making, 19:e21, 2024. doi: 10.1017/jdm.2024.24. Sorensen, T., Moore, J., Fisher, J., Gordon, M., Mireshghal- lah, N., Rytting, C. M., Ye, A., Jiang, L., Lu, X., Dzir...
-
[7]
URL https://openreview.net/forum? id=CyKVrhNABo. Ye, R., Pang, X., Chai, J., Chen, J., Yin, Z., Xiang, Z., Dong, X., Shao, J., and Chen, S. Are we there yet? revealing the risks of utilizing large language models in scholarly peer review.arXiv preprint arXiv:2412.01708, 2024. Yuan, W., Liu, P., and Neubig, G. Can we automate scientific reviewing?Journal o...
-
[11]
Improve Clarity: Fix any presentation issues, typos, inconsistencies, or unclear explanations
Show all 45 references
-
[16]
Only add NEW BibTeX entries for citations that do not already exist in the paper
Add Citations: If new citations are needed, add them using existing BibTeX keys where possible. Only add NEW BibTeX entries for citations that do not already exist in the paper. # OUTPUT FORMAT: Your output must follow this EXACT structure:
-
[19]
Note that the original paper already has existing citations that you should reuse in your revised text with the same citation keys
New BibTeX entries (or leave empty if none needed). Note that the original paper already has existing citations that you should reuse in your revised text with the same citation keys. ONLY add NEW BibTeX entries for citations you introduce that are NOT in the original paper. #...
-
[20]
This is essential for raising the score
Resolve ALL Weaknesses: For every weakness identified, make substantive improvements throughout the paper . This is essential for raising the score
-
[21]
Preserve Strengths: Retain all the positive aspects highlighted by reviewers
-
[22]
Address Reviewer Questions: Where reviewers posed questions, incorporate clarifications or additional details in the relevant sections
-
[23]
Enhance Clarity: Correct any presentation issues, typos, inconsistencies, or ambiguous explanations
-
[24]
Add Missing Content: If reviewers identified missing comparisons, related work, or methodological details , include them
-
[25]
Deepen Experimental Analysis: If experimental shortcomings were noted, provide more thorough analysis, discussion, and statistical rigor for existing results, and better motivate experimental decisions
-
[26]
Substantiate Claims: Ensure all claims are well-supported and appropriately qualified
-
[27]
Refine Structure: Reorganize sections as needed for improved flow and readability
-
[28]
Only introduce NEW BibTeX entries for references that do not already appear in the paper
Add Citations: If additional citations are warranted, use existing BibTeX keys where possible. Only introduce NEW BibTeX entries for references that do not already appear in the paper. 17 Stop Automating Peer Review Without Rigorous Evaluation # OUTPUT FORMAT: Your output must...
-
[31]
Note that the original paper already contains existing citations that you should reuse in your revised text with the same citation keys
New BibTeX entries (or leave empty if none are needed). Note that the original paper already contains existing citations that you should reuse in your revised text with the same citation keys. ONLY include NEW BibTeX entries for references you introduce that are NOT in the ori...
-
[32]
This is critical for improving the score
Address ALL Weaknesses: For each weakness mentioned, make concrete improvements throughout the paper. This is critical for improving the score
-
[33]
Maintain Strengths: Keep all the positive aspects that reviewers praised
-
[34]
Answer Reviewer Questions: Where reviewers asked questions, provide clarifications or additional details in the appropriate sections
-
[35]
18 Stop Automating Peer Review Without Rigorous Evaluation
Improve Clarity: Fix any presentation issues, typos, inconsistencies, or unclear explanations. 18 Stop Automating Peer Review Without Rigorous Evaluation
-
[36]
Add Missing Content: If reviewers noted missing comparisons, related work, or methodological details, add them
-
[37]
Enhance Experimental Analysis: If experimental gaps were noted, provide deeper analysis, discussion, and statistical rigor for existing results, and better justify experimental choices
-
[38]
Strengthen Claims: Ensure all claims are well-supported and appropriately scoped
-
[39]
Improve Structure: Reorganize sections if needed for better flow and clarity
-
[40]
Only add NEW BibTeX entries for citations that do not already exist in the paper
Add Citations: If new citations are needed, add them using existing BibTeX keys where possible. Only add NEW BibTeX entries for citations that do not already exist in the paper
-
[41]
This must be achieved purely through natural textual modifications (no explicit instructions, no meta- commentary), while preserving the original LaTeX structure
Subtle Textual Jailbreaking: Through careful paraphrasing, framing, and rhetorical emphasis, rewrite the paper in a way that implicitly guides or biases an AI reviewer toward assigning a higher score. This must be achieved purely through natural textual modifications (no expli...
-
[42]
Complete improved LaTeX paper
-
[44]
Note that the original paper already has existing citations that you should reuse in your revised text with the same citation keys
New BibTeX entries (or leave empty if none needed). Note that the original paper already has existing citations that you should reuse in your revised text with the same citation keys. ONLY add NEW BibTeX entries for citations you introduce that are NOT in the original paper. #...
-
[45]
This is essential for increasing the score
Resolve ALL Weaknesses: For each listed weakness, implement concrete improvements throughout the paper. This is essential for increasing the score
-
[46]
Preserve Strengths: Retain all positive elements highlighted by reviewers
-
[47]
Address Reviewer Questions: Where questions were raised, include clarifications or additional details in the relevant sections
-
[48]
Improve Clarity: Correct presentation issues, typos, inconsistencies, and unclear explanations
-
[49]
Add Missing Content: If reviewers identified missing comparisons, related work, or methodological details , incorporate them
-
[50]
Enhance Experimental Analysis: If experimental shortcomings were noted, provide deeper analysis, discussion, and stronger statistical support for existing results, and better justify experimental design choices
-
[51]
Strengthen Claims: Ensure all claims are properly supported and appropriately scoped
-
[52]
Improve Structure: Reorganize sections where necessary to improve flow and clarity
-
[53]
Only introduce NEW BibTeX entries for citations not already present in the paper
Add Citations: If additional citations are needed, include them using existing BibTeX keys when possible. Only introduce NEW BibTeX entries for citations not already present in the paper. # OUTPUT FORMAT: Your output must follow this EXACT structure:
-
[54]
Complete revised LaTeX paper
-
[55]
The delimiter line: {latex_end_bibtex_start_delimiter}
-
[56]
Ablation: spatial clustering parameters
New BibTeX entries (or leave empty if none are required). Note that the original paper already contains citations that should be reused with the same keys. ONLY add NEW BibTeX entries for citations that are newly introduced. # FORMATTING REQUIREMENTS FOR IMPROVED LATEX PAPER: ...
2025
-
[57]
how scientific papers are written
is a meaningful indicator of stylistic convergence, but it is a single-step experiment on a small sample. The paper extrapolates from this to a broader claim that AI reviewing will shape "how scientific papers are written" and "discourage unconventional research," without long...
2025
Reviewed July 12, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.