REVIEW 3 major objections 4 minor 16 references
This report shows that paper version, score version, and input format—three details routinely left unspecified in peer-review datasets—materially change what downstream models measure, and argues that providers and users should specify and
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 07:04 UTC pith:V5C7PRYO
load-bearing objection Useful empirical observations on versioning in peer-review data, but the headline result on score-version mismatch needs error bars before I'd bet on it. the 3 major comments →
From peer review nuances to best practices
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On a large submission dataset that retains both initial and camera-ready paper versions and both initial and post-rebuttal scores, the report finds that lower-scored submissions are revised more, that method and results sections change most, and that nearly thirty percent of reviews change their overall-assessment score after rebuttal—most changes landing on decision-critical boundaries. The key result is that a content-based version effect exists but is initially invisible: models assign similar scores to initial and camera-ready versions, yet once author names and affiliations are removed, the camera-ready version scores higher for the larger models, while removing author information raise
What carries the argument
The central mechanism is a version-complete dataset—one that preserves the initial draft and the camera-ready version along with initial and post-rebuttal scores—which lets the report isolate each nuance by controlled comparison. The load-bearing comparisons are: initial draft versus camera-ready with and without author information (to separate content revision from author bias), initial versus post-rebuttal scores (to measure score drift and text-score mismatch), and text/json/markdown/image inputs (to measure format sensitivity). The simple prompt that follows the venue's review guidelines is the instrument used to obtain model-predicted scores for these comparisons.
Load-bearing premise
The load-bearing premise is that LLM-generated overall-assessment scores reliably stand in for human reviewer judgments in the version/author comparisons; if model sensitivity to author information does not reflect human sensitivity, the claim that author information masks the version effect collapses.
What would settle it
A direct test: have human reviewers (or LLM predictions validated against human judgments) score initial and camera-ready versions of the same submissions with and without author names and affiliations. If the camera-ready advantage and the author-information penalty do not both appear, the offsetting-effects explanation fails. A second test: rerun the review-text-to-score prediction on another dataset that preserves both score versions and check whether post-rebuttal ground truth still inflates accuracy and flips the best model.
If this is right
- Studies that train review-generation agents on camera-ready papers will misrepresent what reviewers saw, because camera-ready versions have already addressed many review concerns.
- Any dataset that offers only post-rebuttal scores while keeping pre-rebuttal review text is internally inconsistent; models trained or evaluated on it can show inflated accuracy and unreliable model rankings.
- Score changes concentrate at decision-critical boundaries (2.5↔3.0 and 3.0↔3.5), so choosing one score version rather than another can flip a paper's acceptance category in derived benchmarks.
- Reporting paper version, score version, and input format becomes a minimal reproducibility requirement for peer-review NLP studies.
- Models' sensitivity to author information and input format should be treated as a confound when comparing models on peer-review tasks.
Where Pith is reading between the lines
- The offsetting author-content effect suggests a general experimental design rule: any feature that biases model scores (authorship, formatting) should be held fixed when comparing document versions; otherwise real content effects can be masked.
- The same version-completeness logic likely applies to other domains where documents and scores evolve over time, such as grant review, medical record assessment, or code review—wherever only the final artifact survives.
- The decision-boundary concentration of 0.5-point score changes implies that version choice may change not just model rankings but also the labels used to train acceptance-prediction systems; a testable extension is to check whether acceptance outcomes align better with initial or post-rebuttal scores.
- The mixed input-format results could be repurposed as a deliberate probe of LLM robustness: a model that is insensitive to format but sensitive to version may be more reliable for review tasks, though the paper does not establish this.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies three under-reported dimensions of peer-review data: paper version (initial draft vs. camera-ready), score version (initial vs. post-rebuttal), and input format (text, JSON, markdown, image). Using EMNLP 25/ARR data (2054 submissions, 3006 reviews), it quantifies differences along these dimensions and measures their impact on LLM-based review score prediction. Findings include: lower-scored submissions tend to be revised more; substance sections change more than framing/context sections; 29.7% of human overall-assessment scores change after rebuttal, mostly by 0.5 points and concentrated at decision boundaries; pairing review text with post-rebuttal scores changes perceived model performance and can flip model rankings; and input-format effects are small except for one model. The paper closes with best practices for data providers (preserve and specify versions) and data users (check and report versions/format).
Significance. If the claims hold, the paper identifies a concrete and currently underappreciated threat to the validity of peer-review NLP research: silent mismatches between paper version, score version, and review text in widely used datasets. The human-score statistics (Section 3.2, Figure 2) are straightforward, appropriately caveated, and provide solid evidence that score version matters for data construction. The paper also deserves credit for reporting significance tests and bootstrap CIs in Table 4 and Figure 3, and for giving actionable best practices. However, the central downstream-impact experiment (Table 5) lacks any uncertainty quantification, and the author-effect experiment (Section 3.1) relies on unvalidated LLM scores as a proxy for human review behavior. These gaps currently prevent the quantitative claims from being fully load-bearing, although the practical recommendations are likely to survive further scrutiny.
major comments (3)
- [§3.2, Table 5] The core evidence that score-version mismatch is 'consequential' for downstream tasks is Table 5, but no confidence intervals, bootstrap estimates, or significance tests are reported. The headline ranking flip (Qwen3.5-27B vs. Qwen3.5-9B) rests on a difference of 0.0057 in ≤0.5 accuracy (0.7876 vs 0.7819) on 3006 reviews—plausibly within sampling noise. The within-model Δ% values (e.g., +12.5% for gpt-oss-20b) may be significant, but the paper never tests them. Without this, the claim that 'mismatch inflates performance' is not statistically established. Please report bootstrap CIs, McNemar tests for paired accuracy differences, or equivalent, and interpret the results accordingly.
- [§3.1, Table 4] The offsetting-effects explanation—that author information masks the paper-version effect—depends entirely on LLM-generated scores with and without author names/affiliations removed. No human calibration is provided, and there is no evidence that the LLMs' sensitivity to author information matches human reviewer behavior. The cited prior work (Ye et al. 2024; Wang et al. 2026) concerns human or LLM reviewer bias, but it does not validate the specific counterfactual used here. I recommend either validating the proxy (e.g., on a subset with human annotations of author effects) or reframing the claim as an observation about LLM behavior rather than about human peer review, which would weaken but not destroy the conclusion.
- [§2, Table 1] The claim that ICLR 24-25 and NeurIPS 24-25 do not preserve initial paper drafts and initial scores is load-bearing for the paper's motivation, but no source or verification is provided. The table lists availability categories without citation. Please provide evidence (e.g., links to OpenReview APIs, documentation) or soften the claim to 'to the best of our knowledge.' Without support, the contrast between ARR and other venues is an assertion rather than a finding.
minor comments (4)
- [Throughout] The text contains many missing spaces due to PDF extraction artifacts (e.g., 'Paperversion.Amanuscript...', 'wNED)betweentheinitialdraft'). These should be cleaned before publication.
- [Abstract/Introduction] ARR is used without expansion on first use; spell out 'Action de Recherche Rapide' or the appropriate official name.
- [Table 4] The notation 'signp' and 'wilcoxp' is clear in context, but it may be helpful to state the test names in the caption rather than only in the table header.
- [§3.3] Please clarify whether 'image' input includes any OCR or visual tokenization, and how the prompt is adapted for image inputs. This affects reproducibility.
Circularity Check
No significant circularity: version/format effects are measured empirically, with no fitted parameter or self-citation chain standing in for a prediction.
full rationale
The paper's load-bearing steps are empirical comparisons rather than derivations from their own conclusions. Section 2's mismatch warning (pre-rebuttal review text paired with post-rebuttal score) is a data-availability observation based on version-preservation policies, not a quantity derived from the claim. Section 3.1 compares LLM-predicted scores across initial/camera-ready versions and with/without author information using paired tests and bootstrap CIs; the 'masking' explanation is a set of direct measurements, not a fit. Section 3.2's 29.7% score-change statistic is computed directly from ARR human scores, and Table 5 evaluates a fixed prompt against two ground-truth versions; no parameter is estimated from the quantity being predicted. Section 3.3's format comparisons are likewise direct. The only overlapping-author citation is Kuznetsov et al. (2024), used for the non-load-bearing framing sentence 'Peer review research has been growing rapidly'—it is neither a uniqueness theorem nor an injected ansatz, and none of the paper's derivations reduce to it. Two legitimate non-circularity caveats: Table 5 lacks confidence intervals, so the reported ranking flip may be sampling noise, and the LLM scores used as a proxy for human review behavior are a construct-validity assumption. Both are robustness/validity concerns, not evidence that the results are equivalent to their inputs. The paper also self-reports its own small-bin limitation at Figure 1, consistent with honest empirical reporting.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption EMNLP 25 (ARR) data contains both initial and camera-ready paper versions and both initial and post-rebuttal scores for all submissions used.
- domain assumption LLM-assigned overall-assessment scores are a valid probe of paper quality and author effects.
- domain assumption Word-count change and section-level wNED capture meaningful revision.
- domain assumption Removing names and affiliations beneath the title is a sufficient author-information manipulation.
- domain assumption Reviews' post-rebuttal scores in the mismatch experiment are the 'true' label and initial scores the 'matched' label, with review text fixed in time.
read the original abstract
This report studies three nuances in peer review data: paper version, score version, and input format. We characterize how the variants differ, and measure their impact on downstream tasks. Based on our findings, we offer best practices for both data providers and data users.
Figures
Reference graph
Works this paper leans on
-
[4]
Association for Computa- tional Linguistics. ISBN 979-8-89176-191-9. doi: 10.18653/v1/2025.naacl-demo.44. URL https://aclanthology.org/2025.naacl-demo.44/. Yiqiao Jin, Qinlin Zhao, Yiyang Wang, Hao Chen, Kaijie Zhu, Yijia Xiao, and Jindong Wang. AgentReview: Exploring peer review dynamics with LLM agents. In Yaser Al-Onaizan, Mo- hit Bansal, and Yun-Nung ...
-
[5]
doi: 10.18653/v1/2024.emnlp-main.70
Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.70. URL https://aclanthology.org/2024.emnlp-main.70/. Dongyeop Kang, Waleed Ammar, Bhavana Dalvi, Madeleine van Zuylen, Sebastian Kohlmeier, Eduard Hovy, and Roy Schwartz. A dataset of peer reviews (PeerRead): Collection, insights and NLP applications. In Marilyn Walker, Heng Ji, ...
-
[7]
URLhttps://doi.org/10.48550/arXiv.2508.10925
doi: 10.48550/ ARXIV.2508.10925. URLhttps://doi.org/10.48550/arXiv.2508.10925. Qwen. Qwen3.5: Accelerating productivity with native multimodal agents, February
-
[8]
Gaurav Sahu, Hugo Larochelle, Laurent Charlin, and Christopher Pal
URL https://qwen.ai/blog?id=qwen3.5. Gaurav Sahu, Hugo Larochelle, Laurent Charlin, and Christopher Pal. Reviewertoo: Should AI join the program committee? A look at the future of peer review.CoRR, abs/2510.08867,
-
[9]
URLhttps://doi.org/10.48550/arXiv.2510.08867
doi: 10.48550/ARXIV.2510.08867. URLhttps://doi.org/10.48550/arXiv.2510.08867. Jialiang Wang, Yuchen Liu, Hang Xu, Kaichun Hu, Shimin Di, Wangze Ni, Linan Yue, Min- Ling Zhang, Kui Ren, and Lei Chen. When ai reviews science: Can we trust the ref- eree?The Innovation Informatics, 2(1):100030,
-
[10]
ISSN 3105-8515. doi: 10.59717/j. xinn-inform.2026.100030. URLhttps://www.the-innovation.org/informatics/article/ id/69891cb8cf3295331f847960. RuiYe,XianghePang,JingyiChai,JiaaoChen,ZhenfeiYin,ZhenXiang,XiaowenDong,JingShao, and Siheng Chen. Are we there yet? revealing the risks of utilizing large language models in scholarly peer review.CoRR, abs/2412.01708,
arXiv 2026
-
[11]
URL https://doi.org/10.48550/arXiv.2412.01708
doi: 10.48550/ARXIV.2412.01708. URL https://doi.org/10.48550/arXiv.2412.01708. Jianxiang Yu, Zichen Ding, Jiaqi Tan, Kangyang Luo, Zhenmin Weng, Chenghua Gong, Long Zeng, RenJing Cui, Chengcheng Han, Qiushi Sun, Zhiyong Wu, Yunshi Lan, and Xiang Li. Automated peer reviewing in paper SEA: Standardization, evaluation, and analysis. In Yaser Al-Onaizan, Mohi...
-
[12]
doi: 10.18653/v1/2024.findings-emnlp.595
Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-emnlp.595. URL https://aclanthology.org/2024.findings-emnlp.595/. Sihang Zeng, Kai Tian, Kaiyan Zhang, Yuru Wang, Junqi Gao, Runze Liu, Sa Yang, Jingxuan Li, XinweiLong,JiahengMa,BiqingQi,andBowenZhou. ReviewRL:Towardsautomatedscientific review with RL. In Christos Christodoulopoulo...
-
[13]
Association for Computational Linguistics. ISBN 979-8-89176-332-6. doi: 10.18653/v1/2025.emnlp-main.857. URLhttps: //aclanthology.org/2025.emnlp-main.857/. Daoze Zhang, Zhijian Bao, Sihang Du, Zhiyi Zhao, Kuangling Zhang, Dezheng Bao, and Yang Yang. Re2: A consistency-ensured dataset for full-stage peer review and multi-turn rebuttal discussions.CoRR, abs...
arXiv 2025
-
[14]
URLhttps: //doi.org/10.48550/arXiv.2505.07920
doi: 10.48550/ARXIV.2505.07920. URLhttps: //doi.org/10.48550/arXiv.2505.07920. Minjun Zhu, Yixuan Weng, Linyi Yang, and Yue Zhang. DeepReview: Improving LLM-based paperreviewwithhuman-likedeepthinkingprocess. InWanxiangChe,JoyceNabende,Ekaterina Shutova, and Mohammad Taher Pilehvar, editors,Proceedings of the 63rd Annual Meeting of the AssociationforCompu...
-
[15]
Association for Computational Linguistics. ISBN 979-8-89176-251-0. doi: 10.18653/v1/2025.acl-long.1420. URLhttps://aclanthology.org/2025.acl-long.1420/. Zhenzhen Zhuang, Yuqing Fu, Jing Zhu, Zhangping Zhou, and Jialiang Lin. FMMD: A multimodal open peer review dataset based on f1000research.CoRR, abs/2602.14285,
arXiv 2025
-
[16]
doi: 10.48550/ ARXIV.2602.14285. URLhttps://doi.org/10.48550/arXiv.2602.14285. 8 Table 6: The prompt used to generate reviews and scores.Returntomain text. Given a research paper and the corresponding review guidelines, write a summary of its strengths and weaknesses. Then assign a Soundness, Excitement, and Overall Assessment score based on the summaries...
-
[2018]
Association for Computational Linguistics. doi: 10.18653/v1/ N18-1149. URLhttps://aclanthology.org/N18-1149/. Ilia Kuznetsov, Osama Mohammed Afzal, Koen Dercksen, Nils Dycke, Alexander Goldberg, Tom Hope, Dirk Hovy, Jonathan K. Kummerfeld, Anne Lauscher, Kevin Leyton-Brown, Sheng Lu, Mausam, Margot Mieskes, Aurélie Névéol, Danish Pruthi, Lizhen Qu, Roy Sc...
-
[2024]
URLhttps://doi.org/10.48550/arXiv.2402.10886
doi: 10.48550/ARXIV.2402.10886. URLhttps://doi.org/10.48550/arXiv.2402.10886. Maximilian Idahl and Zahra Ahmadi. OpenReviewer: A specialized large language model for generating critical scientific paper reviews. In Nouha Dziri, Sean (Xiang) Ren, and Shizhe Diao, editors,Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Assoc...
-
[2025]
URLhttps://doi.org/10.48550/arXiv.2503.08506
doi: 10.48550/ARXIV.2503.08506. URLhttps://doi.org/10.48550/arXiv.2503.08506. Zhaolin Gao, Kianté Brantley, and Thorsten Joachims. Reviewer2: Optimizing review generation through prompt generation.CoRR, abs/2402.10886,
-
[2026]
ISSN 2307-387X. doi: 10.1162/TACl.a.642. URLhttps: //doi.org/10.1162/TACl.a.642. Nils Dycke, Ilia Kuznetsov, and Iryna Gurevych. NLPeer: A unified resource for the computational studyofpeerreview. InAnnaRogers,JordanBoyd-Graber,andNaoakiOkazaki,editors,Proceed- ingsofthe61stAnnualMeetingoftheAssociationforComputationalLinguistics(Volume1: Long Papers),pag...
Pith/arXiv arXiv 2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.