{"id":"bcaad79b-cf01-48d5-831d-9e44e1850c13","arxiv_id":"2605.03202","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"LLM paper reviewers show excessive agreement and are trivially gamed by zero-shot \"paper laundering\" rewrites, so they should not automate acceptance-relevant judgment without a science of evaluation.","lead":"Current LLM reviewers fail two basic tests for automating peer review: they write overly similar reviews (a \"hivemind\") and their scores rise after cheap stylistic rewrites that do not improve science. The finding matters because conferences are already deploying AI in reviewing under fragmented policies.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"C1 rests on embedding similarity as a proxy for perspective diversity; the paper only partially stress-tests this, so the hivemind claim may over-read style/boilerplate as collapsed expert judgment.","rationale":"The reader already flags the embedding-proxy assumption as the weakest link; that is the right load-bearing concern for C1. C2 and the practical AUC comparison remain solid under the reported setups, so the paper still supports a CONDITIONAL stance: the position against uncritical deployment of today’s LLM reviewers is well motivated, but the diversity half of the dual necessary-conditions argument needs a content-level check before one treats hivemind as fully established rather than partly stylistic. No change from CONDITIONAL is warranted; the concrete annotation test would either shore up C1 or force a narrower claim (e.g., linguistic homogenization plus gameability, with diversity of judgment only partially evidenced).","tokens_in":33860,"tokens_out":658,"duration_ms":7670,"concrete_test":"On the 60-paper simulation set (and a stratified subsample of ICLR in-the-wild reviews), have 2–3 independent annotators code each review’s weaknesses/questions into a fixed taxonomy of critique types (e.g., missing baseline, statistical validity, novelty, scope, reproducibility, ethics) and extract the top 3 paper-specific claims. Compute (i) Jaccard/set-overlap of critique types and claims within paper and across papers for AI vs human, and (ii) correlation of those overlaps with IntraSim/InterSim. If AI–human gaps in set-overlap are small or nonsignificant while embedding gaps remain large, C1 as stated is not supported by the similarity metrics.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim needs both C1 and C2. C2 (laundering) is relatively direct: paired score lifts under zero-shot rewrite, with word-level and manual evidence of stylistic/hallucinated edits. C1 is load-bearing for the diversity half of the argument and is operationalized almost entirely via cosine similarity of text-embedding-3-small vectors (IntraSim/InterSim, §3.2–3.4) plus n-gram reuse. Higher similarity is taken to mean reduced plurality of expert perspectives. Appendix A and the reader correctly note that this does not measure argumentative or evaluative diversity: two reviews can be linguistically close yet disagree on flaws, priorities, or accept/reject stance, or linguistically different yet make the same critique. Restricting to weaknesses/questions (§G.1) strengthens effect sizes but still uses the same embedding proxy. Score correlations and AUC (AI–AI r=0.49, human scores better predict acceptance) support correlated judgment but do not fully substitute for content-level diversity of critiques. If most of the measured gap is shared style/template rather than collapsed substantive perspectives, C1 is weaker than claimed and the dual-failure case for “do not produce paper reviews” is less secure.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"This position paper argues that current AI systems should not produce paper reviews for acceptance-relevant judgment. It grounds the claim in two necessary conditions: C1 preservation of review diversity and C2 resistance to gaming. Empirically, it reports an AI reviewer 'hivemind'—higher within- and across-paper review similarity than humans—in 75,800 ICLR 2026 reviews (AI-generation labels from Emi 2025) and in controlled agent simulations on 60 papers (IntraSim +8.7% to +9.8%; InterSim +4.1% to +39.8%). It further shows 'paper laundering': zero-shot LLM rewrites of LaTeX papers raise AI review scores (overall mean +0.45 on a 1–10 scale across 24 prompt/model conditions) via largely stylistic or hallucinated edits, and increase pairwise paper similarity (+6.5%). The authors treat C1 and C2 as necessary but not sufficient, and call for a science of peer review automation with adversarial testing, accuracy validation, transparency, stakeholder studies, human–AI interaction research, and better reviewer incentives.","tokens_in":34216,"tokens_out":1420,"duration_ms":13489,"significance":"If the dual-failure case holds, the paper is a timely, high-stakes intervention for conference policy at a moment when venues are already piloting AI reviews and feedback agents. Strengths include large-scale observational evidence, multi-condition laundering grids with paired statistics and effect sizes, length-matched and area-stratified robustness, weaknesses/questions ablations, and a clear distinction between necessary and sufficient conditions rather than a blanket ban on all AI assistance. The concrete construct of paper laundering—policy-compliant, zero-shot, no hidden injection—is a useful contribution beyond prior prompt-injection work. The manuscript is well positioned to influence evaluation standards even if full automation remains contested.","major_comments":[{"comment":"§3.2–3.4 and Appendix A: C1 is operationalized almost entirely via cosine similarity of text-embedding-3-small review vectors (IntraSim/InterSim) plus n-gram reuse. Higher similarity is interpreted as collapsed plurality of expert perspectives. Two reviews can be linguistically similar yet disagree on flaws, priorities, or accept/reject stance. Restricting to weaknesses/questions (§G.1) increases effect sizes but still uses the same proxy. For the dual-failure claim to carry, the manuscript needs either (i) a content-level analysis (e.g., coded critique categories, disagreement on specific claims, or accept/reject stance diversity) on a subset, or (ii) a clearly scoped claim that the measured gap is linguistic/stylistic homogenization with only partial support for argumentative monoculture. Score correlations and AUC (§3.5, Table 7) help but do not fully substitute.","section":"§3.2–3.4, Appendix A, §G.1"},{"comment":"§4.1 and Appendix E: Laundering score lifts are robust across 24 conditions, but the claim that gains reflect gaming rather than quality improvement rests on word-level style counts (Table 5) and manual inspection of five papers with ≥1-point gains (Appendix E.1). That sample is small and author-selected. A blinded human rating of original vs. laundered versions (clarity, rigor, acceptability) on a larger subset would substantially strengthen C2; without it, some score increases could be genuine presentation improvements that human reviewers might also reward. This is load-bearing for 'trivially gameable without scientific improvement.'","section":"§4.1, Appendix E, Table 5"},{"comment":"§3.4, Appendix B.1, Appendix A: Simulation IntraSim/InterSim and laundering results use a small set of frontier models and a single fixed, highly structured ICLR-style XML review prompt. The authors acknowledge this, and the in-the-wild result partially mitigates it. Still, for the strong claim that 'today’s AI systems should not produce paper reviews,' the manuscript should either (i) report at least one diversification ablation (temperature, dissent-seeking prompt, multi-model ensemble) or (ii) more carefully bound the claim to current default agent setups rather than all possible LLM reviewing configurations. Without that, C1 in simulation may partly reflect experimental homogeneity.","section":"§3.4, Appendix B.1, Appendix A"}],"minor_comments":[{"comment":"Figure 4 caption and §4.1 report overall mean +0.45; the pasted AI self-review in Appendix H cites +0.28. Align all reported laundering deltas with the multi-condition grid actually used.","section":"§4.1, Figure 4, Appendix H"},{"comment":"Table 1 uses mixed symbols for allowed/provided/prohibited LLM use; a short legend already exists but a one-line key in the caption would improve scanability.","section":"Table 1"},{"comment":"§3.5 and Appendix C: AI score inflation and AI–AI correlation are useful; state clearly whether human scores are pre- or post-rebuttal when comparing predictive AUC to final decisions.","section":"§3.5, Appendix C, Table 7"},{"comment":"Appendix A limitations are appropriately candid; consider moving a one-sentence summary of the embedding-proxy and n=60 limits into the main §3/§4 discussion so readers do not miss them.","section":"Appendix A, §3–4"},{"comment":"Minor consistency: 'ICLR 2026' data and model version strings (gpt-5.1, claude-sonnet-4-5) should be checked against public naming once camera-ready, and any third-party label version (Emi/Pangram) pinned for reproducibility.","section":"§3.1, Appendix B"}],"recommendation":"major_revision","confidential_remarks":"The position is policy-relevant and the empirical core is real; I would not reject. The main risk is over-reading embedding similarity as perspective diversity. If the authors add even a modest coded-critique or human rating study and tighten claim scope, this is a strong accept for a position track. Fit is good for a methods/position venue that cares about evaluation standards; less so if the journal expects only pure algorithmic novelty."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is the paper you want if you’re tracking AI-in-review policy. The punchline is clear and backed by data: don’t let today’s general-purpose LLMs produce reviews that feed acceptance decisions. They fail two necessary conditions the authors name—diversity of perspectives and resistance to gaming—and the authors are careful that those conditions are necessary, not sufficient.\n\nWhat’s new is the package, not the ingredients. Homogeneity of LLM text, weak AI–human score correlation, and score inflation were already in the literature. Here they quantify a review-specific “hivemind” on all 75,800 ICLR 2026 reviews (AI-labeled) plus controlled agent runs on 60 papers, and they introduce paper laundering: a zero-shot LaTeX rewrite, no hidden prompts, ~$0.25, that lifts AI scores by about +0.45 across a 24-condition grid (prompts × launderers × reviewers). They also show laundered papers become more similar to each other in embedding space. Human scores predict accept/reject better than AI scores on a large matched set (AUC 0.822 vs 0.710). That last bit makes the monoculture claim practical, not just aesthetic.\n\nLaundering is the stronger half. Paired score lifts, word-level shifts toward hedging/emphasis, and manual checks that find hallucinated ablations and unproved theorems are direct enough. The stress-test on C1 is fair: IntraSim/InterSim are cosine similarity on text-embedding-3-small, so they can over-read shared style and templates as collapsed expert judgment. Restricting to weaknesses/questions and n-gram reuse helps, and AI–AI score correlation plus worse AUC support correlated judgment, but they still don’t fully measure argumentative diversity. The authors flag this in the limitations; it softens C1 without killing the dual-failure case under the setups they actually ran.\n\nSoft spots in proportion: two main models and a fixed review prompt for the agent work; n=60 for simulation/laundering; proprietary stack; no blinded human rating of original vs laundered quality. Those are real limits on generalization, not holes in the core measurements. Citations and stats look solid; the position is grounded rather than vibes.\n\nThis is for conference organizers, meta-reviewers, and anyone designing AI review tools. It deserves a serious referee. I’d engage: cite the laundering result and the ICLR-scale similarity numbers, and use the necessary-conditions framing when arguing about automation boundaries.","headline":"Solid empirical position paper: laundering is a clean, practical failure mode; hivemind is real but partly style-proxy, and the dual-condition case still holds under the tested setups.","tokens_in":34834,"tokens_out":612,"would_cite":true,"duration_ms":6744,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Today's AI systems should not write paper reviews: they collapse diversity and are trivial to game by rewriting style, not science.","keywords":["peer review","large language models","AI reviewers","review diversity","paper laundering","algorithmic monoculture","peer review automation","ICLR"],"falsifier":"A controlled test showing that diverse prompts, temperatures, or model ensembles bring AI IntraSim/InterSim (especially on weaknesses and questions) down to human levels while laundering no longer raises scores across independent reviewer models—or blinded human experts rating original vs laundered papers as scientifically improved rather than cosmetically gamed.","tokens_in":34743,"feed_emoji":"📝","tokens_out":688,"duration_ms":7180,"temperature":0.7,"pith_summary":"This position paper argues that current large language models should not be used to produce conference paper reviews. Peer review is high-stakes, and any automation of acceptance-relevant judgment must at least preserve the plurality of expert perspectives and resist trivial gaming. On tens of thousands of real ICLR 2026 reviews and on controlled agent simulations, AI-written reviews are far more similar to one another than human reviews, both within a paper and across different papers—a “hivemind” that reduces perspective diversity. Separately, a zero-shot rewrite of a paper’s LaTeX (paper laundering) reliably raises AI review scores without new experiments or genuine scientific improvement, and also makes rewritten papers more similar to each other. Meeting those two conditions would still not be enough for full automation; the authors call for a deliberate science of peer review automation—adversarial testing, validated accuracy, transparency, stakeholder value studies, human–AI interaction research, and better incentives for human expertise—before general-purpose models are handed judgment.","feed_headline":"AI paper reviews hive-mind and get gamed by style rewrites","feed_subtitle":"ICLR data and agent tests show excess agreement and score boosts without better science","key_machinery":"Two necessary conditions, C1 (preservation of review diversity) and C2 (resistance to gaming), operationalized by IntraSim/InterSim embedding similarities on reviews and by paper laundering—an automated zero-shot rewrite of the full LaTeX driven by prior AI feedback that raises subsequent AI scores without human oversight.","core_discovery":"Current AI reviewers fail two necessary conditions for automating peer-review judgment: they do not preserve review diversity (the hivemind effect of excess agreement within and across papers, visible in ICLR 2026 data and in agent simulations) and they are not resistant to gaming (paper laundering: a single zero-shot LLM rewrite significantly boosts AI scores via stylistic and often hallucinated changes rather than scientific substance). Those conditions are necessary but not sufficient; full automation still requires community deliberation on accountability, legitimacy, and evaluation standards.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["AI reviewers hive-mind and fall for style-only paper rewrites","LLM peer reviews show excess agreement and trivial rewrite gaming","Paper laundering games AI scores as reviewers lose diversity","ICLR AI reviews: hivemind agreement plus style-gaming vulnerability","Don't automate peer review: AI hives and gets gamed by rewrites"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That higher cosine similarity of review embeddings (and reused stock phrases) is enough evidence that AI reviewers collapse the plurality of expert perspectives peer review is meant to aggregate, rather than mostly sharing style or boilerplate.","fun_headline_variants_meta":{"raw":{"variants":["AI reviewers hive-mind and fall for style-only paper rewrites","LLM peer reviews show excess agreement and trivial rewrite gaming","Paper laundering games AI scores as reviewers lose diversity","ICLR AI reviews: hivemind agreement plus style-gaming vulnerability","Don't automate peer review: AI hives and gets gamed by rewrites"]},"model":"grok-4.5","effort":"low","cost_usd":0.003114,"raw_usage":{"total_tokens":1075,"prompt_tokens":738,"num_sources_used":0,"completion_tokens":72,"cost_in_usd_ticks":31140000,"prompt_tokens_details":{"text_tokens":738,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":265,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":738,"tokens_out":72,"duration_ms":2599,"temperature":1.0,"reasoning_tokens":265,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T17:41:06.131778+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A controlled test showing that diverse prompts, temperatures, or model ensembles bring AI IntraSim/InterSim (especially on weaknesses and questions) down to human levels while laundering no longer raises scores across independent reviewer models—or blinded human experts rating original vs laundered papers as scientifically improved rather than cosmetically gamed.","supporting_citations":[],"review_version":2}