Pith. sign in

REVIEW 4 major objections 5 minor 31 references

Reviewer scores are not comparable across ML research areas: at any given score, acceptance odds vary up to 8-fold by topic.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 15:08 UTC pith:SSF73N6N

load-bearing objection A genuinely new ICLR-scale audit of score-conditional acceptance, but the 8x headline is a pooled-raw-score artifact that needs year-stratified resampling before it can carry the abstract. the 4 major comments →

arxiv 2607.27209 v1 pith:SSF73N6N submitted 2026-04-30 cs.DL cs.AIcs.LG

Reviewer Scores Are Not Comparable Across Research Areas in ML Peer Review

classification cs.DL cs.AIcs.LG
keywords peer reviewreviewer scoresacceptance biasresearch topicsICLRscore calibrationmeasurement designfairness in review
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tries to establish that the standard peer-review instrument of ML conferences, the mean reviewer score, does not carry the same meaning across research areas. Using ICLR 2021–2026 data on 50,289 papers grouped into 219 topics, it shows that papers receiving the same reviewer score face acceptance probabilities that differ by up to 8× depending on their topic, and that topic identity predicts acceptance beyond the score (statistical test p≈5×10⁻³). The authors argue the cause is structural rather than personal bias: a single fixed numerical scale cannot encode both within-area relative quality and cross-area absolute quality when reviewer pools are non-uniform, so area chairs substitute community priors for score-based decisions. They rule out scoring culture, expert-reviewer standards, rational area-chair reweighting, and quality dilution as alternative explanations. If the paper is right, equal scores do not imply equal acceptance probability, and publication decisions are shaped partly by community interest in a topic rather than by reviewer-assessed quality alone.

Core claim

The paper's central discovery is that reviewer scores and acceptance decisions are decoupled along topic lines: at identical mean reviewer scores, acceptance rates differ by up to 8× across research topics. For example, at a mean score near 5.0, policy-optimization and LLM-reasoning papers reach 52.6% acceptance while CNN and adversarial-robustness papers sit at 6.6%. This decoupling holds at every score level and is largest in the borderline zone where area-chair discretion is highest. The authors identify the root cause as a measurement design failure: because reviewer assignment is structurally non-uniform in expertise depth, scoring culture, and novelty baselines, absolute scores are inc

What carries the argument

The central instrument is the same-score-band comparison: for narrow windows of mean reviewer score (±0.125 points around 4.0–6.0), the paper compares acceptance rates across topics with at least 50 papers in the band. This is the model-free result that requires no assumptions about paper quality, and it directly demonstrates that equal reviewer scores yield unequal acceptance probabilities. A second mechanism is the per-topic acceptance threshold τ̂, estimated by fitting a logistic curve P(accept)=σ(β₀+β₁·score) within each topic and solving for the 50% acceptance score; the spread of these thresholds (0.812 points) shows that the decision bar itself varies by topic. Together they carry the

Load-bearing premise

The analysis assumes that reviewers in different areas are not just using the 1–10 scale differently, so that equal numerical scores reflect equal reviewer-perceived quality; if a 5 in one area means what a 6 does in another, the 8× acceptance gap is the system working correctly.

What would settle it

A decisive falsifier would be to re-analyze the same ICLR data after converting each paper's mean reviewer score to a within-area percentile rank; if the 8× same-score acceptance gap at score ≈5.0 collapses toward 1× when scores are expressed as area-relative percentiles, then the gap is driven by scale-usage differences rather than by a measurement design failure.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the paper is correct, mean reviewer score should not be treated as a comparable measure across research areas; cross-area comparisons of acceptance bars based on scores are invalid.
  • Publishing topic-stratified, score-conditional acceptance rates becomes a first-class fairness metric that program committees can adopt without changing the review workflow.
  • Calibrated review signals, such as reviewer calibration profiles and author self-rankings, are motivated as supplements to raw scores.
  • The documented hype-cycle premium implies that emerging topics receive an acceptance boost at equal scores, while declining topics face a penalty, which should persist unless interventions are introduced.
  • Analogous patterns at other ML venues remain undetectable because of opt-in review disclosure, giving venues an additional transparency argument.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Our inference: if the 8× same-score gap is real, then a paper's topic label is a strategic choice variable, and authors may relabel marginal work toward high-acceptance topics; transparent publication of score-conditional rates could either expose or exacerbate this gaming.
  • Our inference: the structural argument generalizes beyond ML—any large conference or grant review that uses a fixed numerical scale across heterogeneous reviewer communities should exhibit similar score incomparability, even if the effect is smaller.
  • Our inference: a direct test of the measurement-design claim would be to run the same analysis using within-area percentile rankings of scores instead of raw scores; if the gap collapses, the incomparability is a scale-usage artifact rather than a structural measurement failure.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper analyzes ICLR 2021–2026 data (50,289 papers, 219 BERTopic-derived topics) to argue that acceptance probabilities are not comparable across research areas at equal mean reviewer scores. The headline evidence is that, in the same raw-score band (±0.125 around 4.0–6.0), top-5 topics are accepted up to 8× more often than bottom-5 topics (Table 1/Figure 3). The authors supplement this with a cluster-robust likelihood-ratio test showing topic identity adds variance beyond score and year (χ²=341.2, df=277, p=1.8×10⁻²), a per-topic acceptance-threshold analysis, and robustness checks intended to rule out scoring culture, expert-reviewer standards, AC reweighting, and quality dilution. They conclude that the root cause is a measurement design failure—non-uniform reviewer pools making absolute scores incomparable—and propose publishing topic-stratified score-conditional acceptance rates plus calibrated review signals.

Significance. If the 8× same-score gap is real, this is an important result for ML peer-review research and for conference policy: it would mean that the implicit contract that equal reviewer scores carry equal accept/reject probabilities is violated in a systematic, area-dependent way. The paper is commendably transparent about several limitations: Figure 4 is explicitly labeled circular-by-construction and illustrative, the NeurIPS validation is honestly described as selection-biased, and the observational, non-causal scope is stated. The descriptive core—that topic-level acceptance rates vary substantially and that topic identity predicts acceptance beyond score and year—is supported by the raw data and the cluster-robust LRT. However, the load-bearing '8×' claim and the 'scoring culture is ruled out' conclusion rest on analyses with unresolved confounding and incomplete methodology. The central claim is defensible but not yet established; a focused revision could make it compelling.

major comments (4)
  1. [§2.3, Table 1/Figure 3] The headline same-score-band analysis pools six years of raw mean reviewer scores. The paper itself documents in §3.2 that the global acceptance threshold fell by 0.33–1.09 points as submissions tripled, and the top topics at score ≈5.0 (policy optimization, LLM reasoning, chain-of-thought) are concentrated in 2025–2026 while the bottom topics (CNN, adversarial robustness, residual connections) are concentrated earlier. Table 1 therefore cannot separate a topic effect from a year/cohort effect. The §2.1 LRT includes year dummies, but Table 1 does not, and Appendix C(b) does not stratify by year. Please provide year-stratified same-score-band comparisons, or repeat the analysis on within-year aligned scores, or fit a model with topic×year interactions; the 8× claim should be re-estimated under year controls.
  2. [§2.3/Table 1] The 8× ratio is an extreme-groups statistic: the top-5 and bottom-5 of up to 76 topics are selected after sorting by in-band acceptance rate. This selection induces winner's-curse inflation, and no bootstrap or permutation confidence interval is reported for this specific statistic. The permutation test in Figure 5(c) addresses a different max/min statistic (observed 11.4×), not the Table 1 ratios. Please report bootstrap or permutation intervals for the top-5/bottom-5 ratio at each score band, and consider reporting all topics in the band rather than only the extremes.
  3. [§2.2 / Appendix C(b), Figure 5(b)] The 'scoring culture' alternative—that a raw 5 in a strict-scoring community denotes the same quality as a 6 in a lenient one—is the most direct threat to the interpretation that the same-score gap is a measurement failure. The paper's response is that z-scoring within area leaves '79% of the gap' persisting, but Appendix C(b) does not define the spread measure, report uncertainty, stratify by year, or specify whether the normalization is at primary-area or topic level despite the multi-label structure. This is not yet a falsification of scoring culture. Please give a complete specification: topic-level, year-stratified normalization; the exact metric used to measure the gap; and confidence intervals for the 79% figure.
  4. [§2.3, 'AC rate control' paragraph] The argument that AC rate control 'cannot explain within-score-band gaps' is logically strained. If area chairs set different per-area thresholds to compensate for scoring differences, then at any fixed raw score acceptance odds will differ across areas—which is precisely the pattern in Table 1. Shifting a *uniform* threshold would not produce such differences, but per-area calibration would. The paragraph conflates these two cases. The subscore analysis later addresses rational reweighting, but this paragraph as written does not rule out AC rate control.
minor comments (5)
  1. [§5.1] Typo: 'it ges unmeasured' should be 'it goes unmeasured'.
  2. [Figure 3] The score-5.0 panel lists 'CoT reasoning chain' twice among the displayed topics; one is presumably intended to be a different topic. Please check the labels and ensure the top-5 and bottom-5 counts are each five distinct topics.
  3. [Figure 2 caption] The caption says the 20 lowest and 20 highest thresholds are shown 'out of 278 qualifying topics,' while the main text and the figure title reference the 146 forest-plot-eligible topics. Please clarify which denominator applies to the threshold analysis.
  4. [Appendix B and Appendix E] Minor editorial: 'analyzes' should be 'analyses'; also the taxonomy count transition (219 / 278 / 146 / 322) is initially confusing—Table 3 helps, but a sentence in the main text explaining the four counts would improve readability.
  5. [Table 1] The '—' for the ratio at score 4.0 is explained, but the 12pp gap with a 0.0% bottom-5 acceptance rate deserves a sentence in the main text noting that the ratio is undefined exactly because of zero events in the bottom group.

Circularity Check

1 steps flagged

Only disclosed illustrative circularity; central score-acceptance decoupling claim is independent.

specific steps
  1. self definitional [Appendix A.4 (Figure 4 technical note); also Section 2.3 footnote 1]
    "Figure 4 uses a split-by-AR-group design that introduces a circularity by construction: topics are sorted by six-year AR to form groups, and then in-band ARs (which contribute to the six-year AR) are plotted."

    The illustrative mechanism splits topics into top-20/bottom-20 by six-year AR and then reports that these groups differ in acceptance rate within score bands. Since the grouping variable is the same outcome (AR) as the plotted in-band AR, the high-AR group's higher in-band acceptance is guaranteed to some degree by the selection rule: in-band AR is a constituent of the six-year AR used to define the groups. The authors explicitly acknowledge this ('circularity by construction') and relegate the figure to illustration, citing Table 1 as the primary model-free evidence, so this circular step is not load-bearing for the paper's central 8x claim.

full rationale

The central claim is not circular. The same-score-band gaps in Table 1/Figure 3 are computed from raw score bands and acceptance outcomes; topic labels are content-derived (BERTopic), not defined by the outcome. The likelihood-ratio test (chi2=341.2, df=277, p=5.1e-3; cluster-robust p=1.8e-2) is an independent test of whether topic dummies add explanatory power beyond score and year, and the paper's own caveats (observational, ICLR-only, NeurIPS selection bias) do not hide a definitional reduction. The only explicit circularity is Figure 4's high/low-AR grouping, which the authors themselves disclose as 'correlated with AR by construction' and label as illustrative, explicitly deferring to Table 1 as the primary evidence; it therefore does not infect the headline claim. The top-5/bottom-5 design in Table 1 is a selection on the outcome and is vulnerable to winner's curse and lacks a null distribution for the 8x ratio, but this is a statistical robustness concern, not definitional circularity. No load-bearing self-citation chain exists: the data-source reference [Yang et al., 2025] is an external repository (authors Jing Yang et al.), not the present authors' own derivation, and the paper validates acceptance rates against official ICLR figures within 1 percentage point. Score 2 reflects only the disclosed, non-load-bearing illustrative circularity.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 0 invented entities

No new physical or conceptual entities are postulated. 'Community priors' and 'measurement design failure' are labels for observed statistical patterns, not independently evidenced mechanisms. The main free parameters are binning and inclusion thresholds that directly affect the headline 8x statistic, plus topic-taxonomy hyperparameters and post-fit logistic filters.

free parameters (6)
  • Score-band half-width = 0.125
    Papers are grouped into mean-score windows of ±0.125 around 4.0, 4.5, 5.0, 5.5, 6.0. This bin width is chosen by hand; sensitivity checks with ±0.0625 and ±0.25 are reported as qualitatively identical, but the headline 8x statistic depends on this binning.
  • Minimum in-band papers per topic = 50
    Topics must have ≥50 papers in a score band to enter Table 1/Figure 3, chosen to keep the Wilson CI half-width below 14 percentage points. This directly determines which topics are comparable and shapes the top-5/bottom-5 extreme groups.
  • Logistic post-fit filters = β1 > 0.3; τ̂ ∈ [3.0, 8.5]
    Per-topic threshold regressions retain only fits with positive discrimination and thresholds in a plausible range; 44 of 322 topics are excluded. These filters shape the threshold analysis (Figure 2) but not the model-free score-band analysis.
  • BERTopic min_cluster_size = 10
    The 219-topic taxonomy is built with HDBSCAN min_cluster_size=10 after UMAP reduction to 10 dimensions on OpenAI embeddings. This clustering choice determines the topic partition underlying all topic-level analyses.
  • Forest-plot inclusion thresholds = ≥250 matched papers; ≥30 borderline papers
    Figure 1 includes only 146 of 219 topics meeting these filters; the descriptive AR spread from 18.3% to 46.0% depends on which topics are shown.
  • Within-year aligned-score normalization = z-score per year, rescaled to 1–10
    Threshold estimates use within-year standardized reviewer means to remove year-level drift. The choice of this alignment affects the per-topic bar τ̂ but not the raw same-score-band Table 1.
axioms (5)
  • domain assumption ICLR full-review disclosure data in PaperCopilot (Yang et al., 2025) accurately records submissions, scores, and decisions for 2021–2026.
    All analysis depends on this dataset. The paper validates pooled acceptance rates against official ICLR figures within 1pp, but does not independently audit per-topic data quality or the multi-label regex matching.
  • domain assumption Mean reviewer rating is the operative quality signal and the right conditioning variable for testing score-acceptance comparability.
    The paper's contract framing assumes acceptance should track reviewer scores. If area chairs legitimately use subscores, confidence, or written reviews beyond the mean, same-score comparisons are incomplete; Section 3 tries to address subscores but not all signals.
  • domain assumption BERTopic-derived topics and regex matching identify meaningful, stable research areas.
    Topic labels are auto-generated plus manual overrides (Appendix E), and one paper can belong to multiple topics, so topic-level estimates are not independent. If topic boundaries are noisy, score-conditional AR differences may be misattributed.
  • domain assumption Within-area z-scoring removes scoring-culture differences, so the residual 79% persistence counts as evidence against the culture explanation.
    Figure 5(b) is used to reject the scoring-culture explanation, but z-scoring by overlapping topic labels is a crude correction, and a partial/null result is treated as a falsification.
  • domain assumption The logistic form P(accept)=σ(β0+β1·score) correctly describes each topic's decision function for threshold estimation.
    Used to compute τ̂ in Section 2.2. If the true decision function is not logistic, threshold spreads could be artifacts. The model-free score-band analysis avoids this but has its own binning assumptions.

pith-pipeline@v1.3.0-alltime-deepseek · 21437 in / 16955 out tokens · 169913 ms · 2026-08-02T15:08:57.549051+00:00 · methodology

0 comments
read the original abstract

Peer review at ML conferences increasingly relies on reviewer scores as the primary decision instrument. As submissions have scaled from thousands to tens of thousands per year, no systematic audit has examined whether this instrument functions uniformly across research areas, or whether acceptance outcomes are in practice shaped by forces that reviewer scores neither capture nor control. This position paper argues that acceptance outcomes are shaped by forces beyond reviewer scores, and that the underlying cause is a measurement design failure, not individual bias. When a fixed numerical scale aggregates quality judgments across communities with structurally non-uniform reviewer pools, absolute scores become incomparable across areas, and area chairs must substitute community priors for score-based decisions. Using ICLR 2021--2026 data covering 50,289 papers across 219 research topics, we show that at any given reviewer score, a paper's acceptance probability varies by up to 8x depending on its topic. We rule out scoring culture, expert reviewer standards, rational area chair reweighting, and quality dilution as alternative explanations. We call on program committees to adopt inherently calibrated review signals and publish topic-stratified, score-conditional acceptance rates as a first-class fairness metric.

Figures

Figures reproduced from arXiv: 2607.27209 by Binyan Xu, Fan Yang, Kehuan Zhang, Xilin Dai.

Figure 1
Figure 1. Figure 1: Acceptance rate disparities across 146 qualifying BERTopic topics at ICLR 2021–2026. Each bar shows the acceptance rate (with 95% Wilson CI) for one topic. Left panel: top 50% of topics by AR (highest AR first). Right panel: bottom 50% of topics by AR (lowest AR last). Only topics with ≥250 total matched papers and ≥30 borderline papers are included. Dashed horizontal line indicates the pooled average (30.… view at source ↗
Figure 2
Figure 2. Figure 2: Per-topic acceptance threshold analysis (ICLR 2021–2026). Each bar shows τˆ, the mean reviewer score at which a paper has a 50% acceptance probability (estimated by per-topic logistic regression on within-year normalized scores). Blue: 20 lowest-threshold topics (easiest bar); red: 20 highest-threshold topics (hardest bar), out of 278 qualifying topics. Spread: 0.812 points (chain-of-thought: τˆ = 4.938; a… view at source ↗
Figure 3
Figure 3. Figure 3: At identical reviewer scores, acceptance rates differ by up to 8× across topics (ICLR 2021–2026). For each of five score bands (±0.125 around 4.0, 4.5, 5.0, 5.5, 6.0), the top-5 (blue) and bottom-5 (red) topics by acceptance rate are shown, among all topics with ≥50 papers in the band. The dashed line marks the overall acceptance rate in each band. At score ≈4.5 and ≈5.0 the ratio between top-5 and bottom-… view at source ↗
Figure 4
Figure 4. Figure 4: Illustrative mechanism: topics with low overall AR face a double disadvantage (ICLR 2021–2026). Left: Distribution of mean normalized reviewer scores for the top-20 and bottom-20 AR topics (n ≥ 50, grouped by overall six-year AR). Low-AR topics receive scores 0.622 points lower on average (t = 9.56, p < 0.001). Right: At each score band, the mean acceptance rate of the two groups. The high-AR group has hig… view at source ↗
Figure 5
Figure 5. Figure 5: Robustness and confound analysis. Four tests of alternative explanations. (a) Reviewer confidence vs. acceptance threshold: no significant correlation (r = −0.102, p = 0.626), rejecting the expert-reviewer explanation. (b) Z-score normalized acceptance rates: 79% of the gap persists after accounting for scoring culture differences. (c) Submission volume growth vs. AR change: no significant correlation (r =… view at source ↗
Figure 6
Figure 6. Figure 6: NeurIPS 2024 external validation. NeurIPS 2024 data is available for approximately 4,830 of ∼15,650 total submissions. Due to NeurIPS’s opt-in review disclosure policy, the available data is heavily biased toward accepted papers (only 268 rejected papers). This selection bias precludes reliable borderline analysis. The figure illustrates the acceptance threshold patterns visible in the available NeurIPS da… view at source ↗
Figure 7
Figure 7. Figure 7: Promotional language analysis: null result. Using the Millar et al. [2022] promotional language lexicon (139 words across 8 categories: superlatives, hedges, novelty claims, etc.), we computed the promotional language score for each paper title. Correlation between promotional language score and acceptance rate across areas: p = 0.94 (not significant). The acceptance rate disparities documented in this pap… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

31 extracted references · 7 linked inside Pith

  1. [1]

    We obtain Λ = 341.2 , p= 5.1×10 −3

    follows a χ2 distribution with 277 degrees of freedom under H0 (topic dummies add no explanatory power). We obtain Λ = 341.2 , p= 5.1×10 −3. Because papers can belong to multiple topics, standard errors and p-values are validated with a cluster-robust sandwich estimator clustering at the paper level, yielding p= 1.8×10 −2. Both specifications rejectH 0 at...

  2. [4]

    Abeba Birhane, Pratyusha Kalluri, Dallas Card, William Agnew, Ravit Dotan, and Michelle Bao

    URLhttps://arxiv.org/abs/2306.03262. Abeba Birhane, Pratyusha Kalluri, Dallas Card, William Agnew, Ravit Dotan, and Michelle Bao. The values encoded in machine learning research.Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency,

  3. [7]

    Maarten R

    URLhttps://arxiv.org/abs/2311.09497. Maarten R. Grootendorst. Bertopic: Neural topic modeling with a class-based tf-idf procedure.ArXiv, abs/2203.05794,

  4. [9]

    doi: 10.18653/v1/ N18-1149

    Association for Computational Linguistics. doi: 10.18653/v1/ N18-1149. URLhttps://aclanthology.org/N18-1149/. Norman Kaplan, R. Merton, and Norman Wyman Storer. The sociology of science: Theoretical and empirical investigations.Journal for the Scientific Study of Religion, 14:70,

  5. [14]

    Weixin Liang, Yuhui Zhang, Hancheng Cao, Binglu Wang, Daisy Ding, Xinyu Yang, Kailas V odra- halli, Siyu He, Daniel Scott Smith, Yian Yin, Daniel A

    URL https://api.semanticscholar.org/CorpusID:279511028. Weixin Liang, Yuhui Zhang, Hancheng Cao, Binglu Wang, Daisy Ding, Xinyu Yang, Kailas V odra- halli, Siyu He, Daniel Scott Smith, Yian Yin, Daniel A. McFarland, and James Zou. Can large lan- guage models provide useful feedback on research papers? a large-scale empirical analysis.ArXiv, abs/2310.01783,

  6. [16]

    Neil Millar, Bojan Batalo, and Brian Budgell

    URL https: //arxiv.org/abs/2512.04448. Neil Millar, Bojan Batalo, and Brian Budgell. Trends in the use of promotional language (hype) in abstracts of successful national institutes of health grant applications, 1985-2020.JAMA Network Open, 5,

  7. [17]

    Joelle Pineau, Philippe Vincent-Lamarre, Koustuv Sinha, Vincent Larivière, Alina Beygelzimer, Flo- rence d’Alché Buc, Emily B

    doi: 10.1073/pnas.1714379115. Joelle Pineau, Philippe Vincent-Lamarre, Koustuv Sinha, Vincent Larivière, Alina Beygelzimer, Flo- rence d’Alché Buc, Emily B. Fox, and H. Larochelle. Improving reproducibility in machine learning research (a report from the neurips 2019 reproducibility program).Journal of Machine Learn- ing Research, 22:164:1–164:20,

  8. [19]

    doi: 10.18653/v1/2020.findings-emnlp.112

    Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.112. URL https: //aclanthology.org/2020.findings-emnlp.112/. Anna Sophie Rogers, Marzena Karpinska, Jordan L. Boyd-Graber, and Naoaki Okazaki. Program chairs’ report on peer review at ACL 2023.Proceedings of the 61st Annual Meeting of the Association for Computational Linguist...

  9. [21]

    Buxin Su, Jiayao Zhang, Natalie Collina, Yuling Yan, Didong Li, Kyunghyun Cho, Jianqing Fan, Aaron Roth, and Weijie J Su

    doi: 10.1145/3449149. Buxin Su, Jiayao Zhang, Natalie Collina, Yuling Yan, Didong Li, Kyunghyun Cho, Jianqing Fan, Aaron Roth, and Weijie J Su. The ICML 2023 ranking experiment: Examining author self- assessments in peer review.Journal of the American Statistical Association,

  10. [22]

    arXiv:2408.13430

    doi: 10.1080/ 01621459.2025.2510006. arXiv:2408.13430. Weijie Su. You are the best reviewer of your own papers: The isotonic mechanism.Opera- tions Research, 74:804–824,

  11. [24]

    doi: 10.48550/arXiv.2010. 05137. Jian Wang, Reinhilde Veugelers, Reinhilde Veugelers, Reinhilde Veugelers, Paula E. Stephan, and Paula E. Stephan. Bias against novelty in science: A cautionary tale for users of bibliometric indicators.Research Policy, 46:1416–1436,

  12. [26]

    12 A Main Results Figures This appendix collects all figures referenced in the main text

    URL https://api.semanticscholar.org/ CorpusID:282102210. 12 A Main Results Figures This appendix collects all figures referenced in the main text. Each figure is accompanied by a technical note detailing data filters, estimation procedures, and design choices relevant to reproduction and interpretation. A.1 Acceptance Rate Disparities Across Topics 0.0 0....

  13. [29]

    auto” (directly from BERTopic keywords) or source=“override

    this available subset is dominated by accepted papers: we identify approximately 4,562 accepted papers with public reviews but only 268 rejected papers, a ratio of 17:1 in favor of accepts, compared to the true NeurIPS 2024 accept/reject ratio of approximately 1:4. This inversion means that any threshold or score-conditional AR analysis on the available N...

  14. [30]

    provide a coarser but author- assigned taxonomy. The BERTopic regex taxonomy has three advantages: (i) it spans all six years, 19 7.5 5.0 2.5 0.0 2.5 5.0 7.5 AR deviation vs papers without this pattern (pp) less is more (n=28) redefine* (n=14) paradigm(*) (n=91) all you need (n=63) tame/taming X (n=52) magic(al) (n=11) harness* (n=71) empower* (n=90) unle...

  15. [31]

    novel”, “superior

    promotional language lexicon (139 words across 8 categories: superlatives, hedges, novelty claims, etc.), we computed the promotional language score for each paper title. Correlation between promotional language score and acceptance rate across areas: p= 0.94 (not significant). The acceptance rate disparities documented in this paper are not explained by ...

  16. [322]

    same score

    The aligned score is the within-year standardized reviewer mean (z-scored to zero mean and unit variance across all papers in that year), then rescaled to the original 1–10 range, to remove year-level threshold drift driven by submission volume growth. Figure 2 displays the 20 topics with the lowest estimated threshold (easiest bar) and the 20 with the hi...

  17. [1927]

    org/CorpusID:121572396

    URLhttps://api.semanticscholar. org/CorpusID:121572396. Jing Yang, Qiyao Wei, and Jiaxin Pei. Paper copilot: Tracking the evolution of peer review in ai conferences.ArXiv, abs/2510.13201,

  18. [1975]

    Amir Hossein Kargaran, Nafiseh Nikeghbal, Jing Yang, and Nedjma Djouhra Ousidhoum

    URL https://api.semanticscholar.org/CorpusID:112124451. Amir Hossein Kargaran, Nafiseh Nikeghbal, Jing Yang, and Nedjma Djouhra Ousidhoum. Insights from the iclr peer review and rebuttal process.ArXiv, abs/2511.15462,

  19. [2013]

    doi: 10.1002/asi. 22784. Yifei Li, Xiaoting Xu, Dongqing Lyu, Zhen Zhang, Juan Xie, and Ying Cheng. Developing a criteria framework for peer review: A critical interpretive synthesis.Learned Publishing, 38,

  20. [2015]

    doi: 10.1145/2732417

    ISSN 0001-0782. doi: 10.1145/2732417. URL https://doi.org/10.1145/2732417. Carole J. Lee, Cassidy R. Sugimoto, Guo Zhang, and Blaise Cronin. Bias in peer review.Journal of the American Society for Information Science and Technology, 64(1):2–17,

  21. [2016]

    Alina Beygelzimer, Yann N

    URL https://api.semanticscholar.org/CorpusID:19641078. Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan. Has the machine learning review process become more arbitrary as the field has grown? the neurips 2021 consistency experiment,

  22. [2017]

    David Tran, Alex Valtchanov, Keshav Ganapathy, Raymond Feng, Eric Slud, Micah Goldblum, and Tom Goldstein

    doi: 10.1073/pnas.1707323114. David Tran, Alex Valtchanov, Keshav Ganapathy, Raymond Feng, Eric Slud, Micah Goldblum, and Tom Goldstein. An open review of OpenReview: A critical analysis of the machine learning conference review process.arXiv preprint arXiv:2010.05137,

  23. [2018]

    Corinna Cortes and Neil D

    URL https://api.semanticscholar.org/CorpusID:126306204. Corinna Cortes and Neil D. Lawrence. Inconsistency in conference peer review: Revisiting the 2014 neurips experiment,

  24. [2019]

    doi: 10.1145/3316774

    ISSN 0001-0782. doi: 10.1145/3316774. URLhttps://doi.org/10.1145/3316774. Jianglin Ma, Ben Yao, Xiang Li, and Yazhou Zhang. Has acl lost its crown? a decade-long quantitative analysis of scale and impact across leading ai conferences,

  25. [2020]

    Anna Rogers and Isabelle Augenstein

    URL https://api.semanticscholar.org/CorpusID: 214693121. Anna Rogers and Isabelle Augenstein. What can we do to improve peer review in NLP? In Trevor Cohn, Yulan He, and Yang Liu, editors,Findings of the Association for Computa- tional Linguistics: EMNLP 2020, pages 1256–1262, Online, November

  26. [2021]

    issue-attention cycle

    URLhttps://arxiv.org/abs/2109.09774. Anthony M R Downs. Up and down with ecology—the “issue-attention cycle”. volume 28, pages 38–50,

  27. [2022]

    11 Ivan Stelmakh, Nihar B

    doi: 10.1145/3528086. 11 Ivan Stelmakh, Nihar B. Shah, and Aarti Singh. Peerreview4all: Fair and accurate reviewer assignment in peer review. InInternational Conference on Algorithmic Learning Theory,

  28. [2023]

    doi: 10.18653/v1/2023.acl-long.734

    Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.734. URL https://aclanthology.org/2023.acl-long.734/. Ariful Azad and Afeefa Banu. Publication trends in artificial intelligence conferences: The rise of super prolific authors,

  29. [2024]

    Stefano Balietti, Robert L

    URLhttps://arxiv.org/abs/2412.07793. Stefano Balietti, Robert L. Goldstone, and Dirk Helbing. Peer review and competition in the art exhibition game.Proceedings of the National Academy of Sciences, 113:8414 – 8419,

  30. [2025]

    Thomas S

    URLhttps://arxiv.org/abs/2505.04966. Thomas S. Kuhn and David Hawkins. The structure of scientific revolutions.American Jour- nal of Physics, 31:554–555,

  31. [2026]

    URL https://arxiv.org/abs/2509. 25701. Dongyeop Kang, Waleed Ammar, Bhavana Dalvi, Madeleine van Zuylen, Sebastian Kohlmeier, Eduard Hovy, and Roy Schwartz. A dataset of peer reviews (PeerRead): Collection, insights and NLP applications. In Marilyn Walker, Heng Ji, and Amanda Stent, editors,Proceedings of the 2018 Conference of the North American Chapter ...