{"id":"21e0ca18-3cd3-476b-823e-3a6a7adfc7d5","arxiv_id":"2607.27209","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"An audit of 50,289 ICLR papers shows acceptance odds vary up to 8x across topics at equal reviewer scores, indicating scores are not comparable across research areas.","lead":"This paper analyzes six years of ICLR review data and finds that papers with identical average reviewer scores can have very different acceptance odds depending on their research topic—up to 8 times. It argues the problem is a structural flaw in how scores are used, not reviewer bias, and urges conferences to publish score-conditional acceptance rates by area.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline 8× same-score gap is computed on raw scores pooled across 2021–2026; with documented threshold drift and topic popularity shifts across years, the gap may be a year artifact rather than a topic effect.","rationale":"Good-faith reading: the paper is transparent about limitations (Figure 4's circularity, observational status) and its LRT controls for year. The descriptive finding likely has some reality, and the transparency proposal is low-cost and defensible. However, the central quantitative claim is presented as model-free when it is not adjusted for the very time effects the paper itself documents. A year-stratified recomputation would either confirm the 8× gap or reveal it as an artifact. This does not overturn the paper, but it strengthens the case for keeping the verdict CONDITIONAL rather than moving to ACCEPT. The reader's weakest_assumption focused on scoring culture; my concern overlaps (both about what 'same score' means) but is distinct in mechanism, hence partial agreement.","tokens_in":21907,"tokens_out":12134,"duration_ms":139710,"concrete_test":"Recompute Table 1 separately for each ICLR year: for score bands centered at 4.5, 5.0, and 5.5 (±0.125), require ≥20 papers per topic per year, take top/bottom acceptance-rate topics, and report the ratio with a bootstrap 95% CI. Repeat using within-year z-scored mean reviewer scores instead of raw scores. If the 8× ratio falls below ~2× or is not significant in most years, the headline overstates topic identity and the 'measurement design failure' inference is not supported by the model-free evidence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing weak point is the headline model-free evidence (Table 1 / Figure 3). It pools six ICLR years and uses raw mean reviewer scores, while the paper itself documents a large temporal drift: Section 3.2 reports the global acceptance threshold fell by 0.33–1.09 points as submissions tripled. The top-5 topics at score ≈5.0 (policy optimization, LLM reasoning, CoT) are concentrated in 2025–2026; the bottom-5 (CNN, GAN, adversarial robustness, residual connections) are concentrated in earlier years. If later cohorts were accepted more often at the same raw score, pooling creates a topic gap that is really a year effect. The §2.1 likelihood-ratio test includes year dummies, but Table 1 does not; the Appendix C(b) z-score normalization is reported only as '79% persists' with no year stratification, no definition of the spread measure, and no uncertainty. Additionally, selecting the top/bottom 5 of 55 topics invites winner's-curse inflation, and no bootstrap/permutation interval is given for the 8× ratio (the permutation test in Figure 5(c) applies to a different max/min statistic). This is load-bearing because the abstract's 8× claim and the 'rule out scoring culture' conclusion rest on this table.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes ICLR 2021–2026 data (50,289 papers, 219 BERTopic-derived topics) to argue that acceptance probabilities are not comparable across research areas at equal mean reviewer scores. The headline evidence is that, in the same raw-score band (±0.125 around 4.0–6.0), top-5 topics are accepted up to 8× more often than bottom-5 topics (Table 1/Figure 3). The authors supplement this with a cluster-robust likelihood-ratio test showing topic identity adds variance beyond score and year (χ²=341.2, df=277, p=1.8×10⁻²), a per-topic acceptance-threshold analysis, and robustness checks intended to rule out scoring culture, expert-reviewer standards, AC reweighting, and quality dilution. They conclude that the root cause is a measurement design failure—non-uniform reviewer pools making absolute scores incomparable—and propose publishing topic-stratified score-conditional acceptance rates plus calibrated review signals.","tokens_in":22218,"tokens_out":5387,"duration_ms":60776,"significance":"If the 8× same-score gap is real, this is an important result for ML peer-review research and for conference policy: it would mean that the implicit contract that equal reviewer scores carry equal accept/reject probabilities is violated in a systematic, area-dependent way. The paper is commendably transparent about several limitations: Figure 4 is explicitly labeled circular-by-construction and illustrative, the NeurIPS validation is honestly described as selection-biased, and the observational, non-causal scope is stated. The descriptive core—that topic-level acceptance rates vary substantially and that topic identity predicts acceptance beyond score and year—is supported by the raw data and the cluster-robust LRT. However, the load-bearing '8×' claim and the 'scoring culture is ruled out' conclusion rest on analyses with unresolved confounding and incomplete methodology. The central claim is defensible but not yet established; a focused revision could make it compelling.","major_comments":[{"comment":"The headline same-score-band analysis pools six years of raw mean reviewer scores. The paper itself documents in §3.2 that the global acceptance threshold fell by 0.33–1.09 points as submissions tripled, and the top topics at score ≈5.0 (policy optimization, LLM reasoning, chain-of-thought) are concentrated in 2025–2026 while the bottom topics (CNN, adversarial robustness, residual connections) are concentrated earlier. Table 1 therefore cannot separate a topic effect from a year/cohort effect. The §2.1 LRT includes year dummies, but Table 1 does not, and Appendix C(b) does not stratify by year. Please provide year-stratified same-score-band comparisons, or repeat the analysis on within-year aligned scores, or fit a model with topic×year interactions; the 8× claim should be re-estimated under year controls.","section":"§2.3, Table 1/Figure 3"},{"comment":"The 8× ratio is an extreme-groups statistic: the top-5 and bottom-5 of up to 76 topics are selected after sorting by in-band acceptance rate. This selection induces winner's-curse inflation, and no bootstrap or permutation confidence interval is reported for this specific statistic. The permutation test in Figure 5(c) addresses a different max/min statistic (observed 11.4×), not the Table 1 ratios. Please report bootstrap or permutation intervals for the top-5/bottom-5 ratio at each score band, and consider reporting all topics in the band rather than only the extremes.","section":"§2.3/Table 1"},{"comment":"The 'scoring culture' alternative—that a raw 5 in a strict-scoring community denotes the same quality as a 6 in a lenient one—is the most direct threat to the interpretation that the same-score gap is a measurement failure. The paper's response is that z-scoring within area leaves '79% of the gap' persisting, but Appendix C(b) does not define the spread measure, report uncertainty, stratify by year, or specify whether the normalization is at primary-area or topic level despite the multi-label structure. This is not yet a falsification of scoring culture. Please give a complete specification: topic-level, year-stratified normalization; the exact metric used to measure the gap; and confidence intervals for the 79% figure.","section":"§2.2 / Appendix C(b), Figure 5(b)"},{"comment":"The argument that AC rate control 'cannot explain within-score-band gaps' is logically strained. If area chairs set different per-area thresholds to compensate for scoring differences, then at any fixed raw score acceptance odds will differ across areas—which is precisely the pattern in Table 1. Shifting a *uniform* threshold would not produce such differences, but per-area calibration would. The paragraph conflates these two cases. The subscore analysis later addresses rational reweighting, but this paragraph as written does not rule out AC rate control.","section":"§2.3, 'AC rate control' paragraph"}],"minor_comments":[{"comment":"Typo: 'it ges unmeasured' should be 'it goes unmeasured'.","section":"§5.1"},{"comment":"The score-5.0 panel lists 'CoT reasoning chain' twice among the displayed topics; one is presumably intended to be a different topic. Please check the labels and ensure the top-5 and bottom-5 counts are each five distinct topics.","section":"Figure 3"},{"comment":"The caption says the 20 lowest and 20 highest thresholds are shown 'out of 278 qualifying topics,' while the main text and the figure title reference the 146 forest-plot-eligible topics. Please clarify which denominator applies to the threshold analysis.","section":"Figure 2 caption"},{"comment":"Minor editorial: 'analyzes' should be 'analyses'; also the taxonomy count transition (219 / 278 / 146 / 322) is initially confusing—Table 3 helps, but a sentence in the main text explaining the four counts would improve readability.","section":"Appendix B and Appendix E"},{"comment":"The '—' for the ratio at score 4.0 is explained, but the 12pp gap with a 0.0% bottom-5 acceptance rate deserves a sentence in the main text noting that the ratio is undefined exactly because of zero events in the bottom group.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is potentially significant for ML peer-review research, but the main quantitative claim currently depends on a year-confounded extreme-groups comparison and an underspecified scoring-culture correction. I would be willing to review a revision that supplies year-stratified same-score evidence with proper uncertainty quantification. The self-disclosed limitations and the candid NeurIPS validation attempt are a strength; the paper should be judged on the strength of the revised central evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about arXiv:2607.27209. First: it's the first topic-level, score-conditional acceptance audit across six years of ICLR, and the structural claim holds up: topic identity predicts acceptance beyond reviewer scores even with year dummies in a cluster-robust LRT (p=0.018). That is a genuinely new empirical contribution. Second: the headline \"8x acceptance gap at the same score\" is not yet load-bearing. Table 1 pools all six years on raw scores, while the same paper documents a global threshold drop of 0.33–1.09 points as submissions tripled. The hot topics are concentrated in 2025–2026 and the cold topics in earlier years, so the 8x is very likely inflated by a year effect. The stress-test concern is on target.\n\nWhat the paper does well: it is unusually candid. It discloses that Figure 4 is circular by construction and labels it illustrative only. It says clearly that the analysis is observational and ICLR-only, and it explains why NeurIPS data can't be used. The transparency proposal — publish score-conditional acceptance rates by area — is low-cost and sensible, and it doesn't depend on the 8x figure being exact. I also checked the data-source circularity concern: PaperCopilot is a third-party dataset with no author overlap, so that worry doesn't land.\n\nSoft spots: the \"79% persists after z-scoring\" sentence in Section 2.2 is the main evidence against the scoring-culture explanation, and it has no methodology, no year stratification, and no uncertainty. The top/bottom-5 design has no bootstrap or permutation interval on the ratio, so winner's curse is unaddressed. Several \"ruled out\" alternatives (expert reviewers, quality dilution) rely on non-significant correlations, which is weak evidence for absence. And the LRT itself uses 277 dummy variables on multi-label topic membership, so the raw p=5e-3 is less impressive than it looks; the cluster-robust p=0.018 is the number to trust.\n\nWho should read it: peer-review researchers, ML conference policy people, anyone working on fairness in evaluation. It deserves a serious referee: the dataset and the core finding are worth engaging with, but the authors need to re-run the same-score analysis year-stratified or with year-blocked resampling, report the missing intervals, and soften the \"rule out\" language. My own verdict is conditional — the transparency recommendation stands on its own; the 8x magnitude needs to be retested before it goes in an abstract.","headline":"A genuinely new ICLR-scale audit of score-conditional acceptance, but the 8x headline is a pooled-raw-score artifact that needs year-stratified resampling before it can carry the abstract.","tokens_in":22723,"tokens_out":4068,"would_cite":true,"duration_ms":44395,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Reviewer scores are not comparable across ML research areas: at any given score, acceptance odds vary up to 8-fold by topic.","keywords":["peer review","reviewer scores","acceptance bias","research topics","ICLR","score calibration","measurement design","fairness in review"],"falsifier":"A decisive falsifier would be to re-analyze the same ICLR data after converting each paper's mean reviewer score to a within-area percentile rank; if the 8× same-score acceptance gap at score ≈5.0 collapses toward 1× when scores are expressed as area-relative percentiles, then the gap is driven by scale-usage differences rather than by a measurement design failure.","tokens_in":21769,"feed_emoji":"📊","tokens_out":3466,"duration_ms":36253,"temperature":0.7,"pith_summary":"This paper tries to establish that the standard peer-review instrument of ML conferences, the mean reviewer score, does not carry the same meaning across research areas. Using ICLR 2021–2026 data on 50,289 papers grouped into 219 topics, it shows that papers receiving the same reviewer score face acceptance probabilities that differ by up to 8× depending on their topic, and that topic identity predicts acceptance beyond the score (statistical test p≈5×10⁻³). The authors argue the cause is structural rather than personal bias: a single fixed numerical scale cannot encode both within-area relative quality and cross-area absolute quality when reviewer pools are non-uniform, so area chairs substitute community priors for score-based decisions. They rule out scoring culture, expert-reviewer standards, rational area-chair reweighting, and quality dilution as alternative explanations. If the paper is right, equal scores do not imply equal acceptance probability, and publication decisions are shaped partly by community interest in a topic rather than by reviewer-assessed quality alone.","feed_headline":"Same reviewer score, 8x different odds of acceptance by topic","feed_subtitle":"ICLR audit of 50,289 papers finds equal scores carry unequal acceptance across topics—and calls it a measurement design failure.","key_machinery":"The central instrument is the same-score-band comparison: for narrow windows of mean reviewer score (±0.125 points around 4.0–6.0), the paper compares acceptance rates across topics with at least 50 papers in the band. This is the model-free result that requires no assumptions about paper quality, and it directly demonstrates that equal reviewer scores yield unequal acceptance probabilities. A second mechanism is the per-topic acceptance threshold τ̂, estimated by fitting a logistic curve P(accept)=σ(β₀+β₁·score) within each topic and solving for the 50% acceptance score; the spread of these thresholds (0.812 points) shows that the decision bar itself varies by topic. Together they carry the","core_discovery":"The paper's central discovery is that reviewer scores and acceptance decisions are decoupled along topic lines: at identical mean reviewer scores, acceptance rates differ by up to 8× across research topics. For example, at a mean score near 5.0, policy-optimization and LLM-reasoning papers reach 52.6% acceptance while CNN and adversarial-robustness papers sit at 6.6%. This decoupling holds at every score level and is largest in the borderline zone where area-chair discretion is highest. The authors identify the root cause as a measurement design failure: because reviewer assignment is structurally non-uniform in expertise depth, scoring culture, and novelty baselines, absolute scores are inc","pith_inferences":["Our inference: if the 8× same-score gap is real, then a paper's topic label is a strategic choice variable, and authors may relabel marginal work toward high-acceptance topics; transparent publication of score-conditional rates could either expose or exacerbate this gaming.","Our inference: the structural argument generalizes beyond ML—any large conference or grant review that uses a fixed numerical scale across heterogeneous reviewer communities should exhibit similar score incomparability, even if the effect is smaller.","Our inference: a direct test of the measurement-design claim would be to run the same analysis using within-area percentile rankings of scores instead of raw scores; if the gap collapses, the incomparability is a scale-usage artifact rather than a structural measurement failure."],"forward_implications":["If the paper is correct, mean reviewer score should not be treated as a comparable measure across research areas; cross-area comparisons of acceptance bars based on scores are invalid.","Publishing topic-stratified, score-conditional acceptance rates becomes a first-class fairness metric that program committees can adopt without changing the review workflow.","Calibrated review signals, such as reviewer calibration profiles and author self-rankings, are motivated as supplements to raw scores.","The documented hype-cycle premium implies that emerging topics receive an acceptance boost at equal scores, while declining topics face a penalty, which should persist unless interventions are introduced.","Analogous patterns at other ML venues remain undetectable because of opt-in review disclosure, giving venues an additional transparency argument."],"fun_headline_variants":["Same score, 8x acceptance gap across ML topics","Equal scores, unequal odds: up to 8x gap by topic","Reviewer scores incomparable across areas: 8x acceptance gap","Same reviewer score, up to 8x odds difference across topics","Scores don't mean equal odds: 8x acceptance gap by topic"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The analysis assumes that reviewers in different areas are not just using the 1–10 scale differently, so that equal numerical scores reflect equal reviewer-perceived quality; if a 5 in one area means what a 6 does in another, the 8× acceptance gap is the system working correctly.","fun_headline_variants_meta":{"raw":{"variants":["Same score, 8x acceptance gap across ML topics","Equal scores, unequal odds: up to 8x gap by topic","Reviewer scores incomparable across areas: 8x acceptance gap","Same reviewer score, up to 8x odds difference across topics","Scores don't mean equal odds: 8x acceptance gap by topic"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000561,"raw_usage":{"total_tokens":2496,"prompt_tokens":737,"completion_tokens":1759,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":481,"completion_tokens_details":{"reasoning_tokens":1667}},"tokens_in":481,"tokens_out":1759,"duration_ms":12599,"temperature":1.0,"reasoning_tokens":1667,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T15:08:57.549051+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive falsifier would be to re-analyze the same ICLR data after converting each paper's mean reviewer score to a within-area percentile rank; if the 8× same-score acceptance gap at score ≈5.0 collapses toward 1× when scores are expressed as area-relative percentiles, then the gap is driven by scale-usage differences rather than by a measurement design failure.","supporting_citations":[],"review_version":1}