Pith. sign in

REVIEW 3 major objections 4 minor 66 references

At equal reviewer scores, borderline papers without a top-institution author are accepted less often at a major machine-learning conference, yet the accepted and rejected papers show no better downstream outcomes—evidence the edge comes fro

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 00:15 UTC pith:HC5TOQFZ

load-bearing objection The audit design is novel and the null is honest, but the headline prestige gap doesn't survive the paper's own permutation test in-sample and the out-of-sample leg is too thin to carry it. the 3 major comments →

arxiv 2607.26280 v1 pith:HC5TOQFZ submitted 2026-07-28 cs.DL

Bias at the Borderline: Who Gets the Benefit of the Doubt in Peer Review?

classification cs.DL
keywords peer reviewdiscriminationoutcome testborderline decisionsinstitutional prestigeICLRdouble-blind reviewpreprint identifiability
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper studies the discretionary calls that decide borderline papers at ICLR, a machine-learning conference whose full review record is public. It establishes that, at identical reviewer scores, papers with no author from a top-25 institution are accepted 0.5 to 1.6 percentage points less often; the gap sits in the area-chair decision after scores, and almost entirely among submissions identifiable through a pre-decision preprint. Running a pre-registered robust outcome test on five downstream quality measures and both sides of the decision, however, the paper finds no disparity consistent with a higher bar—none of the 27 tests survives correction. The paper reads the joint pattern as area chairs using revealed institutional prestige as a prior, a statistical discrimination that outcome tests are structurally blind to and that double-blind review exists to prevent.

Core claim

The central claim is that at ICLR 2019–2025, among the 10,416 borderline submissions, the equal-score acceptance gap against papers without top-25-institution authors (−0.5 percentage points in the discovery half, −1.6 in the confirmatory half) is real but does not reflect a higher bar on any measured outcome. The accepted low-prestige papers do not go on to outperform; the rejected pile does not contain systematically better papers. Instead the disparity concentrates among papers whose authors were identifiable via a pre-decision preprint (−3.4 points against −0.2) and reappears in the pre-registered 2026 cohort. The paper's stated conclusion is that the evidence is consistent with area cha

What carries the argument

Two devices carry the argument. The borderline band—submissions within half a within-year standard deviation of the fitted 50% acceptance threshold—localizes the analysis at the margin where discretion operates. The robust outcome test concludes discrimination only when a benchmark test (equal-score acceptance gap) and an outcome test (downstream quality among accepted and rejected papers) agree directionally; the paper applies it with five outcomes (citations, disruption, two novelty measures, eventual venue) and, on the reject side, traces rejected submissions to their eventual publication.

Load-bearing premise

The null assumes the five measured outcomes faithfully capture the quality an area chair targets, and that the monotone-likelihood-ratio condition holds; if prestige itself inflates these outcomes, or if the relevant quality is unmeasured, a real higher bar would leave exactly the observed null.

What would settle it

Compute the equal-score acceptance gap among borderline papers with a pre-decision preprint enrolled in an enforced anonymous-posting regime; if the −3.4-point gap persists while downstream outcomes stay balanced, the prestige-prior account fails. Conversely, in the ICLR 2026 cohort, a benchmark-concordant positive outcome disparity for low-prestige accepted or rejected papers (e.g., higher citations after tracing) would flip the null toward a higher-bar conclusion.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A conference's preprint policy is a fairness lever: the entire equal-score gap lives where authors are identifiable before the decision.
  • Audits of discretionary gatekeeping should restrict to the margin, pre-register outcomes, and examine the rejected side when it leaves a public trace.
  • The verdict 'none' means no robust-test evidence of a higher bar, not proof that the discretion is benign.
  • The out-of-sample ICLR 2026 cohort recovers the decision-stage prestige gap, so the benchmark result is not a within-sample artifact.
  • Because an accurate prior equalizes marginal outcomes, the outcome-test null is compatible with statistical discrimination grounded in prestige.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the prestige-prior account is right, enforcing anonymous preprinting or delaying preprint visibility until after decisions should shrink or eliminate the −3.4-point gap; this is a testable policy experiment.
  • The outcome null may be partly produced by the same process being tested: prestige inflates citations and venue placement, so a real higher bar could be masked by the Matthew effect; the paper's own negative citation disparities are consistent with that.
  • The same borderline-band, reject-side template could be carried to journals and grant panels that publish rejection records; the paper's design suggests where to look (marginal decisions) and what to demand (concordance, not single-test significance).
  • The gender axis remains under-powered and instrument-dependent; a benchmark disparity with a defensible outcome signature on fresh data would flip that verdict.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper audits the discretionary area-chair stage of ICLR peer review using public OpenReview records for 2019–2025 plus the pre-registered 2026 cohort. It defines a borderline band around the fitted acceptance threshold and estimates three stage-decomposition estimands: the total in-band acceptance gap, the reviewer-score gap, and the equal-score decision-stage gap (the benchmark test). It then applies a five-outcome family (citations, disruption, two novelty measures, eventual venue) to accepted and rejected borderline papers, with a pre-registered 27-cell confirmatory family and a concordance rule inspired by the Gaebler–Goel robust outcome test. The headline affirmative finding is a 0.5–1.6 percentage point lower acceptance rate for borderline papers without a top-25-institution author at equal reviewer scores, concentrated in the decision stage and, suggestively, among submissions with pre-decision arXiv preprints; the result nominally reappears on ICLR 2026. The outcome-test leg returns no benchmark-concordant evidence of a higher bar, which the paper interprets not as exoneration but as consistent with statistical discrimination through identity leakage.

Significance. If the benchmark gap were robust, this would be a valuable template for auditing discretionary gatekeeping: it brings the robust outcome-test logic to peer review, exploits the public record of rejected submissions for a reject-side outcome test, pre-registers a large confirmatory family, and reports unusually extensive robustness analysis (permutation tests, CR2 adjustments, Lee–Manski bounds, revision-extent controls, band-width and prestige-coding grids). The null outcome leg is carefully bounded and the dual-use and proxy-validity caveats are stated honestly. The main weakness is that the central affirmative claim does not survive the paper's own finite-sample permutation test in the main sample, and the out-of-sample leg is a single cohort with no multiplicity correction and no outcome data. The design and transparency are strengths, but the headline claim is currently stronger than the evidence.

major comments (3)
  1. [§5.1 and Appendix F] The paper's central affirmative claim — an equal-score prestige gap at the discretionary stage in ICLR 2019–2025 — does not survive its own stringent permutation test: p_perm = 0.67 (discovery) and 0.20 (confirmatory), pooled 0.27, as reported in §5.1. The text itself says the max-rule q-values 'overstate in-sample certainty.' Since this benchmark gap is the foundation of the 'identity leakage' interpretation, the abstract's '0.5 to 1.6 percentage point lower rate' and the phrase 'survive false-discovery correction' should be presented with the finite-sample failure prominently attached. Either provide an inference procedure whose conditions are met and that sustains the main-sample claim, or explicitly designate the in-sample gap as suggestive and rest the affirmative case on the out-of-sample cohort.
  2. [§5.5 and Table 11] The out-of-sample replication on ICLR 2026 is a single cohort with three pre-registered axes; the one significant prestige result (p_perm = 0.033) is not multiplicity-corrected. A Bonferroni correction across the three axes gives p ≈ 0.099, and a Fisher combination of the two in-sample halves plus 2026 (0.67, 0.20, 0.033) gives p ≈ 0.09. Moreover, 2026 has no outcome data, so it can support only the benchmark leg, not the robust-test conjunction. The sentence calling the 2026 result 'the strongest single evidence' therefore overstates the strength. Please pre-specify and report a multiplicity correction for the 2026 axes and temper the language accordingly.
  3. [§5.2 and §7] The claim that the gap 'concentrates almost entirely' among arXiv-identifiable submissions is based on a pooled-band interaction outside the confirmatory family, with no multiplicity correction, and is subject to selection into posting. The paper acknowledges these limits in §5.2, but the Discussion then elevates the pattern to a central interpretive finding ('the strongest clue,' 'nearly the whole disparity sits there'). Either move this analysis into the pre-registered family with appropriate corrections, or present it consistently as hypothesis-generating and not as part of the core evidentiary claim. As written, the abstract and Discussion give it more weight than the confirmatory status allows.
minor comments (4)
  1. [Abstract and §5.5] 'Reappears out-of-sample' should specify that the 2026 replication covers only the benchmark side; downstream outcomes are structurally unavailable, so the robust-test conclusion is not replicated out-of-sample.
  2. [§3.3 and Table 6] The prestige axis resolves for essentially no 2019–2020 submissions, so headline statements about 'ICLR 2019–2025' for prestige are effectively 2021–2025. Please state this explicitly wherever the pooled window is quoted.
  3. [Figure 2 and §5.2] The interaction estimates are labeled suggestive, but the figure caption should also state that the two significant interaction p-values are not corrected for the multiple axes and partitions examined.
  4. [§5.5] With a single venue-year cluster, the 2026 point estimate has no sandwich-based uncertainty; consider reporting a permutation or bootstrap confidence interval for β_AC in addition to the p-value.

Circularity Check

0 steps flagged

No significant circularity: the audit is pre-registered, the out-of-sample cohort is genuinely future data, and no estimand reduces by construction to a fitted input or self-citation.

full rationale

The paper's derivation chain is an observational audit against external public data (OpenReview, Semantic Scholar, CS-rankings), not a derivation that identifies its conclusion with its inputs. The central benchmark estimand β_AC is a regression-adjusted acceptance gap conditional on reviewer scores; it is not fitted to the outcome family, and the outcome family is a separately pre-registered set of downstream proxies. The ICLR 2026 result is a genuine out-of-sample replication: the cohort was collected after the plan was frozen, downstream outcomes are structurally unavailable, and the permutation test carries the inference. The robust-test verdict is explicitly conjunctive and one-directional; the paper repeatedly states that the outcome null is not proof of equal treatment and that proxies are not quality itself (Sections 4.1, 7; Appendix A). No load-bearing self-citation appears: the robust-test theorem is Gaebler and Goel [24], external work, and the paper explicitly lists its departures from that theorem, calling the design 'inspired by' rather than a direct instantiation. The acknowledged limitations (MLRP validity, proxy contestability, trace selection, few clusters, name-based gender inference) are assumptions and sensitivity analyses, not circular definitions. In particular, the 'revealed prestige prior' interpretation is presented as one reading consistent with the joint pattern, not as an inference forced by the equations. The in-sample permutation-test weakness and the fragility of the 2026 p-value are statistical-evidence concerns, not circularity. I therefore find no step in which a prediction reduces to its own input, and no self-citation chain carries the central claim.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 0 invented entities

The central claims rest on standard statistical machinery (BH, permutation, spline regression), domain assumptions about proxy validity and MLRP, and hand-set thresholds (band half-width, top-25 cut). No invented entities; the prestige-prior mechanism is interpretive. The paper is transparent about most of these, but the proxy-validity assumption is load-bearing for the null and cannot be verified with the released aggregates.

free parameters (4)
  • borderline band half-width = 0.5 within-year SD
    Hand-chosen threshold fixed at discovery and carried into confirmatory analysis; sample definition (10,416 borderline). Robustness over 0.25-1.0 SD attenuates but preserves verdicts (Section 3.2, Appendix G).
  • top-25 prestige boundary = CS-rankings top-25
    Hand-chosen cut defining G=1; boundary-specific effect: top-10 coding flips sign; ranks 11-25 drive the gap (Section 6).
  • Q1 citation window = 3 years
    Pre-registered window; 2- and 5-year variants shift estimates at third decimal (Appendix D, Appendix M).
  • MLRP sensitivity index k, epsilon = k=0.5, epsilon=1e-3
    Stylized pre-committed parameters in breakdown value rho*; heuristic not derived (Appendix N).
axioms (4)
  • domain assumption MLRP holds for the conditioning signal (mean reviewer score)
    Required for the robust outcome test's directional validity; paper reports data-implied rho~1 vs rho*=15.1 and labels the index heuristic (Section 4.1, Appendix N).
  • domain assumption The five outcome measures proxy the quality the AC targets
    Outcome test validity depends on proxies capturing the latent quality; paper acknowledges citations/venue are socially produced and could be prestige-shaped (Section 3.4, Section 7).
  • domain assumption OpenReview public records accurately reflect review scores/decisions
    All analyses use the public record; errors in the record would propagate (Section 3.1).
  • domain assumption Rejected papers' eventual published versions represent the as-rejected paper
    Reject-side outcome test requires this; paper measures revision extent and adds controls, but cannot fully remove revision effects (Section 6, Appendix O).

pith-pipeline@v1.3.0-alltime-deepseek · 27493 in / 12706 out tokens · 108996 ms · 2026-08-01T00:15:22.219509+00:00 · methodology

0 comments
read the original abstract

We study peer review at ICLR, a large machine-learning conference whose complete review record, including rejected submissions, is public. Reviewers score each submission; for the borderline band whose scores do not settle an outcome, an area chair makes a discretionary accept-or-reject call. We ask whether that call is even-handed: do authors from prestigious institutions, WEIRD countries, or all-male teams get the benefit of the doubt at the margin? Across ICLR 2019-2025 (31,711 submissions; 10,416 borderline), borderline papers without a top-25-institution author are accepted at a 0.5 to 1.6 percentage point lower rate at the same reviewer scores. The gap arises at the discretionary stage, reappears out-of-sample in the pre-registered ICLR 2026 cohort, and concentrates almost entirely among submissions identifiable through a pre-decision arXiv preprint (-3.4 vs. -0.2 points). Equal scores need not mean equal papers: an area chair may respond to quality the scores miss. We apply a robust outcome test, which concludes discrimination only when the group accepted at a lower rate also realizes better downstream outcomes. We measure five outcomes (citations, disruption, two forms of novelty, eventual venue) on both sides of the decision, including the first "ones that got away" test of rejected submissions. Our headline result is a null: across a pre-registered family of 27 tests, no disparity concordant with the decision-rate gap survives correction; we find no evidence that any group faced a higher bar on the outcomes we measure. That null is not an exoneration. A pre-decision preprint pierces the blind through policy-permitted means, and the acceptance gap lives almost entirely in that porosity, consistent with area chairs using revealed institutional prestige as a prior: statistical discrimination that outcome tests may not detect, and a practice double-blind review exists to prevent.

Figures

Figures reproduced from arXiv: 2607.26280 by Hazem Ibrahim, Talal Rahwan, Yasir Zaki.

Figure 1
Figure 1. Figure 1: The discretionary regime and the stage decomposition. (A) Within-year-normalized mean reviewer scores, centered at the fitted [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The equal-score gap tracks identifiability. Group [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: All 27 pre-registered outcome cells: the disparity [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: The equal-score acceptance gap across band half-widths, per axis; whiskers are 95% intervals, filled markers survive BH-FDR [PITH_FULL_IMAGE:figures/full_fig_p021_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Design sensitivity. Monte-Carlo power curves (1,000 [PITH_FULL_IMAGE:figures/full_fig_p023_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: The robust outcome test as paired forests, one row block per group axis. Left: the benchmark test (equal-score acceptance gap [PITH_FULL_IMAGE:figures/full_fig_p025_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

66 extracted references · 24 canonical work pages

  1. [1]

    Aksnes, Liv Langfeldt, and Paul Wouters

    Dag W. Aksnes, Liv Langfeldt, and Paul Wouters. 2019. Citations, citation indicators, and research quality: An overview of basic concepts and theories.SAGE Open9, 1 (2019). https://doi.org/10.1177/2158244019829575

  2. [2]

    Shamena Anwar and Hanming Fang. 2006. An alternative test of racial prejudice in motor vehicle searches: Theory and evidence.American Economic Review96, 1 (2006), 127–151. https://doi.org/10.1257/000282806776157579

  3. [3]

    David Arnold, Will Dobbie, and Crystal S. Yang. 2018. Racial bias in bail decisions.The Quarterly Journal of Economics133, 4 (2018), 1885–1932. https://doi.org/10.1093/qje/qjy012 16 Hazem Ibrahim, Talal Rahwan, and Yasir Zaki

  4. [4]

    Kenneth J. Arrow. 1973. The theory of discrimination. InDiscrimination in Labor Markets, Orley Ashenfelter and Albert Rees (Eds.). Princeton University Press, Princeton, NJ, 3–33

  5. [5]

    Ian Ayres. 2002. Outcome tests of racial disparities in police practices.Justice Research and Policy4, 1–2 (2002), 131–142. https://doi.org/10.3818/ JRP.4.1.2002.131

  6. [6]

    Gary S. Becker. 1971.The Economics of Discrimination(2 ed.). University of Chicago Press, Chicago, IL

  7. [7]

    Yoav Benjamini and Yosef Hochberg. 1995. Controlling the false discovery rate: a practical and powerful approach to multiple testing.Journal of the Royal Statistical Society: Series B (Methodological)57, 1 (1995), 289–300

  8. [8]

    Marianne Bertrand and Sendhil Mullainathan. 2004. Are Emily and Greg more employable than Lakisha and Jamal? A field experiment on labor market discrimination.American Economic Review94, 4 (2004), 991–1013. https://doi.org/10.1257/0002828042002561

  9. [9]

    Dauphin, Percy Liang, and Jennifer Wortman Vaughan

    Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan. 2021. The NeurIPS 2021 consistency experiment. NeurIPS Blog. Fuller write-up: arXiv:2306.03262

  10. [10]

    Homanga Bharadhwaj, Dylan Turpin, Animesh Garg, and Ashton Anderson. 2020. De-anonymization of authors through arXiv submissions during double-blind review.arXiv preprint arXiv:2007.00177(2020)

  11. [11]

    Rebecca M. Blank. 1991. The effects of double-blind versus single-blind reviewing: Experimental evidence from The American Economic Review. American Economic Review81, 5 (1991), 1041–1067

  12. [12]

    Aislinn Bohren, Alex Imas, and Michael Rosenberg

    J. Aislinn Bohren, Alex Imas, and Michael Rosenberg. 2019. The dynamics of discrimination: Theory and evidence.American Economic Review109, 10 (2019), 3395–3436. https://doi.org/10.1257/aer.20171829

  13. [13]

    Lutz Bornmann. 2011. Scientific peer review.Annual Review of Information Science and Technology45, 1 (2011), 197–245. https://doi.org/10.1002/ aris.2011.1440450112

  14. [14]

    Lutz Bornmann and Hans-Dieter Daniel. 2008. What do citation counts measure? A review of studies on citing behavior.Journal of Documentation 64, 1 (2008), 45–80. https://doi.org/10.1108/00220410810844150

  15. [15]

    Boudreau, Eva C

    Kevin J. Boudreau, Eva C. Guinan, Karim R. Lakhani, and Christoph Riedl. 2016. Looking across and looking beyond the knowledge frontier: Intellectual distance, novelty, and resource allocation in science.Management Science62, 10 (2016), 2765–2783. https://doi.org/10.1287/mnsc.2015.2285

  16. [16]

    Budden, Tom Tregenza, Lonnie W

    Amber E. Budden, Tom Tregenza, Lonnie W. Aarssen, Julia Koricheva, Roosa Leimu, and Christopher J. Lortie. 2008. Double-blind review favours increased representation of female authors.Trends in Ecology & Evolution23, 1 (2008), 4–6. https://doi.org/10.1016/j.tree.2007.07.008

  17. [17]

    Canay, Magne Mogstad, and Jack Mountjoy

    Ivan A. Canay, Magne Mogstad, and Jack Mountjoy. 2024. On the use of outcome tests for detecting bias in decision making.The Review of Economic Studies91, 4 (2024), 2135–2167. https://doi.org/10.1093/restud/rdad082

  18. [18]

    David Card, Stefano DellaVigna, Patricia Funk, and Nagore Iriberri. 2020. Are referees and editors in economics gender neutral?The Quarterly Journal of Economics135, 1 (2020), 269–327. https://doi.org/10.1093/qje/qjz035

  19. [19]

    Larremore

    Aaron Clauset, Samuel Arbesman, and Daniel B. Larremore. 2015. Systematic inequality and hierarchy in faculty hiring networks.Science Advances 1, 1 (2015), e1400005. https://doi.org/10.1126/sciadv.1400005

  20. [20]

    Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S. Weld. 2020. SPECTER: Document-level representation learning using citation-informed transformers. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2270–2282

  21. [21]

    David Firth. 1993. Bias reduction of maximum likelihood estimates.Biometrika80, 1 (1993), 27–38. https://doi.org/10.1093/biomet/80.1.27

  22. [22]

    Bergstrom, Katy Börner, James A

    Santo Fortunato, Carl T. Bergstrom, Katy Börner, James A. Evans, Dirk Helbing, Staša Milojévić, Alexander M. Petersen, Filippo Radicchi, Roberta Sinatra, Brian Uzzi, Alessandro Vespignani, Ludo Waltman, Dashun Wang, and Albert-László Barabási. 2018. Science of science.Science359, 6379 (2018), eaao0185. https://doi.org/10.1126/science.aao0185

  23. [23]

    Funk and Jason Owen-Smith

    Russell J. Funk and Jason Owen-Smith. 2017. A dynamic network measure of technological change.Management Science63, 3 (2017), 791–817

  24. [24]

    Gaebler and Sharad Goel

    Johann D. Gaebler and Sharad Goel. 2025. A simple, statistically robust test of discrimination.Proceedings of the National Academy of Sciences122, 10 (2025), e2416348122. https://doi.org/10.1073/pnas.2416348122

  25. [25]

    Ginther, Walter T

    Donna K. Ginther, Walter T. Schaffer, Joshua Schnell, Beth Masimore, Faye Liu, Laurel L. Haak, and Raynard Kington. 2011. Race, ethnicity, and NIH research awards.Science333, 6045 (2011), 1015–1019. https://doi.org/10.1126/science.1196783

  26. [26]

    Heine, and Ara Norenzayan

    Joseph Henrich, Steven J. Heine, and Ara Norenzayan. 2010. The weirdest people in the world?Behavioral and Brain Sciences33, 2-3 (2010), 61–83. https://doi.org/10.1017/S0140525X0999152X

  27. [27]

    Kulkarni, Sebastian Munoz-Najar Galvez, Bryan He, Dan Jurafsky, and Daniel A

    Bas Hofstra, Vivek V. Kulkarni, Sebastian Munoz-Najar Galvez, Bryan He, Dan Jurafsky, and Daniel A. McFarland. 2020. The diversity–innovation paradox in science.Proceedings of the National Academy of Sciences117, 17 (2020), 9284–9291. https://doi.org/10.1073/pnas.1915378117

  28. [28]

    D. G. Horvitz and D. J. Thompson. 1952. A generalization of sampling without replacement from a finite universe.J. Amer. Statist. Assoc.47, 260 (1952), 663–685. https://doi.org/10.1080/01621459.1952.10483446

  29. [29]

    Jürgen Huber, Sabiou Inoua, Rudolf Kerschbamer, Christian König-Kersting, Stefan Palan, and Vernon L. Smith. 2022. Nobel and novice: Author prominence affects peer review.Proceedings of the National Academy of Sciences119, 41 (2022), e2205779119. https://doi.org/10.1073/pnas.2205779119

  30. [30]

    2021.What marginal outcome tests can tell us about racially biased decision-making

    Peter Hull. 2021.What marginal outcome tests can tell us about racially biased decision-making. Working Paper 28503. National Bureau of Economic Research. https://doi.org/10.3386/w28503

  31. [31]

    Rodney Kinney, Chloe Anastasiades, Russell Authur, Iz Beltagy, Jonathan Bragg, et al. 2023. The Semantic Scholar open data platform.arXiv preprint arXiv:2301.10140(2023)

  32. [32]

    Jon Kleinberg, Himabindu Lakkaraju, Jure Leskovec, Jens Ludwig, and Sendhil Mullainathan. 2018. Human decisions and machine predictions.The Quarterly Journal of Economics133, 1 (2018), 237–293. https://doi.org/10.1093/qje/qjx032 Bias at the Borderline: Who Gets the Benefit of the Doubt in Peer Review? Evidence from ICLR 17

  33. [33]

    John Knowles, Nicola Persico, and Petra Todd. 2001. Racial bias in motor vehicle searches: Theory and evidence.Journal of Political Economy109, 1 (2001), 203–229. https://doi.org/10.1086/318603

  34. [34]

    Daniël Lakens. 2017. Equivalence tests: A practical primer for t tests, correlations, and meta-analyses.Social Psychological and Personality Science8, 4 (2017), 355–362. https://doi.org/10.1177/1948550617697177

  35. [35]

    John Langford and Mark Guzdial. 2015. The arbitrariness of reviews, and advice for school administrators.Commun. ACM58, 4 (2015), 12–13. https://doi.org/10.1145/2732417

  36. [36]

    Lee, Cassidy R

    Carole J. Lee, Cassidy R. Sugimoto, Guo Zhang, and Blaise Cronin. 2013. Bias in peer review.Journal of the American Society for Information Science and Technology64, 1 (2013), 2–17. https://doi.org/10.1002/asi.22784

  37. [37]

    David S. Lee. 2009. Training, wages, and sample selection: Estimating sharp bounds on treatment effects.The Review of Economic Studies76, 3 (2009), 1071–1102. https://doi.org/10.1111/j.1467-937X.2009.00536.x

  38. [38]

    Lee and Thomas Lemieux

    David S. Lee and Thomas Lemieux. 2010. Regression discontinuity designs in economics.Journal of Economic Literature48, 2 (2010), 281–355. https://doi.org/10.1257/jel.48.2.281

  39. [39]

    You Cannot Sound Like GPT

    Haley Lepp and Daniel Scott Smith. 2025. “You Cannot Sound Like GPT”: Signs of language discrimination and resistance in computer science publishing. InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency. 3162–3181. https://doi.org/10.1145/3715275. 3732202

  40. [40]

    Charles F. Manski. 1990. Nonparametric bounds on treatment effects.The American Economic Review80, 2 (1990), 319–323. Papers and Proceedings

  41. [41]

    Emaad Manzoor and Nihar B. Shah. 2021. Uncovering latent biases in text: Method and application to peer review. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 4767–4775. https://doi.org/10.1609/aaai.v35i6.16608

  42. [42]

    Robert K. Merton. 1968. The Matthew effect in science.Science159, 3810 (1968), 56–63. https://doi.org/10.1126/science.159.3810.56

  43. [43]

    Hug, Mininder S

    Kanu Okike, Kevin T. Hug, Mininder S. Kocher, and Seth S. Leopold. 2016. Single-blind vs double-blind peer review in the setting of author prestige. JAMA316, 12 (2016), 1315–1316. https://doi.org/10.1001/jama.2016.11014

  44. [44]

    Michael Park, Erin Leahey, and Russell J. Funk. 2023. Papers and patents are becoming less disruptive over time.Nature613 (2023), 138–144. https://doi.org/10.1038/s41586-022-05543-x

  45. [45]

    Peters and Stephen J

    Douglas P. Peters and Stephen J. Ceci. 1982. Peer-review practices of psychological journals: The fate of published articles, submitted again. Behavioral and Brain Sciences5, 2 (1982), 187–195. https://doi.org/10.1017/S0140525X00011183

  46. [46]

    Edmund S. Phelps. 1972. The statistical theory of racism and sexism.American Economic Review62, 4 (1972), 659–661

  47. [47]

    Pier, Markus Brauer, Amarette Filut, Anna Kaatz, Joshua Raclaw, Mitchell J

    Elizabeth L. Pier, Markus Brauer, Amarette Filut, Anna Kaatz, Joshua Raclaw, Mitchell J. Nathan, Cecilia E. Ford, and Molly Carnes. 2018. Low agreement among reviewers evaluating the same NIH grant applications.Proceedings of the National Academy of Sciences115, 12 (2018), 2952–2957. https://doi.org/10.1073/pnas.1714379115

  48. [48]

    Emma Pierson, Camelia Simoiu, Jan Overgoor, Sam Corbett-Davies, Daniel Jenson, Amy Shoemaker, Vignesh Ramachandran, Phoebe Barghouty, Cheryl Phillips, Ravi Shroff, and Sharad Goel. 2020. A large-scale analysis of racial disparities in police stops across the United States.Nature Human Behaviour4, 7 (2020), 736–745. https://doi.org/10.1038/s41562-020-0858-1

  49. [49]

    Charvi Rastogi, Ivan Stelmakh, Xinwei Shen, Marina Meila, Federico Echenique, Shuchi Chawla, and Nihar B. Shah. 2022. To arXiv or not to arXiv: A study quantifying pros and cons of posting preprints online. arXiv:2203.17259

  50. [50]

    Ross, Cary P

    Joseph S. Ross, Cary P. Gross, Mayur M. Desai, Yuling Hong, Augustus O. Grant, Stephen R. Daniels, Vladimir C. Hachinski, Raymond J. Gibbons, Timothy J. Gardner, and Harlan M. Krumholz. 2006. Effect of blinded peer review on abstract acceptance.JAMA295, 14 (2006), 1675–1680. https://doi.org/10.1001/jama.295.14.1675

  51. [51]

    Davidson, Veniamin Veselovsky, and Robert West

    Giuseppe Russo Latona, Manoel Horta Ribeiro, Tim R. Davidson, Veniamin Veselovsky, and Robert West. 2025. The AI review lottery: Widespread AI-assisted peer reviews boost paper scores and acceptance rates.Proceedings of the ACM on Human-Computer Interaction9, 7 (CSCW) (2025). https://doi.org/10.1145/3757667

  52. [52]

    Lucía Santamaría and Helena Mihaljević. 2018. Comparison and benchmark of name-to-gender inference services.PeerJ Computer Science4 (2018), e156. https://doi.org/10.7717/peerj-cs.156

  53. [53]

    Schuirmann

    Donald J. Schuirmann. 1987. A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability.Journal of Pharmacokinetics and Biopharmaceutics15, 6 (1987), 657–680. https://doi.org/10.1007/BF01068419

  54. [54]

    Nihar B. Shah. 2022. Challenges, experiments, and computational solutions in peer review.Commun. ACM65, 6 (2022), 76–87. https://doi.org/10. 1145/3528086

  55. [55]

    Camelia Simoiu, Sam Corbett-Davies, and Sharad Goel. 2017. The problem of infra-marginality in outcome tests for discrimination.Annals of Applied Statistics11, 3 (2017), 1193–1216. https://doi.org/10.1214/17-AOAS1058

  56. [56]

    Flaminio Squazzoni, Giangiacomo Bravo, Mike Farjam, Ana Marušić, Bahar Mehmani, Michael Willis, Aliaksandr Birukou, Pierpaolo Dondio, and Francisco Grimaldo. 2021. Peer review and gender bias: A study on 145 scholarly journals.Science Advances7, 2 (2021), eabd0299. https: //doi.org/10.1126/sciadv.abd0299

  57. [57]

    Ivan Stelmakh, Charvi Rastogi, Ryan Liu, Shuchi Chawla, Federico Echenique, and Nihar B. Shah. 2023. Cite-seeing and reviewing: A study on citation bias in peer review.PLOS ONE18, 7 (2023), e0283980. https://doi.org/10.1371/journal.pone.0283980

  58. [58]

    Shah, Aarti Singh, and Hal Daumé III

    Ivan Stelmakh, Nihar B. Shah, Aarti Singh, and Hal Daumé III. 2021. Prior and prejudice: The novice reviewers’ bias against resubmissions in conference peer review.Proceedings of the ACM on Human-Computer Interaction5, CSCW1, Article 75 (2021). https://doi.org/10.1145/3449149 18 Hazem Ibrahim, Talal Rahwan, and Yasir Zaki

  59. [59]

    Misha Teplitskiy, Daniel Acuna, Aïda Elamrani-Raoult, Konrad Körding, and James Evans. 2018. The sociology of scientific validity: How professional networks shape judgement in peer review.Research Policy47, 9 (2018), 1825–1841. https://doi.org/10.1016/j.respol.2018.06.014

  60. [60]

    Andrew Tomkins, Min Zhang, and William D. Heavlin. 2017. Reviewer bias in single- versus double-blind peer review.Proceedings of the National Academy of Sciences114, 48 (2017), 12708–12713. https://doi.org/10.1073/pnas.1707323114

  61. [61]

    David Tran, Alex Valtchanov, Keshav Ganapathy, Raymond Feng, Eric Slud, Micah Goldblum, and Tom Goldstein. 2020. An open review of OpenReview: A critical analysis of the machine learning conference review process. InNeurIPS 2020 Workshop on Navigating the Broader Impacts of AI Research. Workshop paper, arXiv:2010.05137

  62. [62]

    Brian Uzzi, Satyam Mukherjee, Michael Stringer, and Ben Jones. 2013. Atypical combinations and scientific impact.Science342, 6157 (2013), 468–472

  63. [63]

    Jian Wang, Reinhilde Veugelers, and Paula Stephan. 2017. Bias against novelty in science: A cautionary tale for users of bibliometric indicators. Research Policy46, 8 (2017), 1416–1436. https://doi.org/10.1016/j.respol.2017.06.006

  64. [64]

    Way, Allison C

    Samuel F. Way, Allison C. Morgan, Daniel B. Larremore, and Aaron Clauset. 2019. Productivity, prominence, and the effects of academic environment. Proceedings of the National Academy of Sciences116, 22 (2019), 10729–10733. https://doi.org/10.1073/pnas.1817431116

  65. [65]

    Witteman, Michael Hendricks, Sharon Straus, and Cara Tannenbaum

    Holly O. Witteman, Michael Hendricks, Sharon Straus, and Cara Tannenbaum. 2019. Are gender gaps due to evaluations of the applicant or the science? A natural experiment at a national funding agency.The Lancet393, 10171 (2019), 531–540. https://doi.org/10.1016/S0140-6736(18)32611-4

  66. [66]

    confirmatory

    Lingfei Wu, Dashun Wang, and James A. Evans. 2019. Large teams develop and small teams disrupt science and technology.Nature566 (2019), 378–382. A Relation to the robust-test theorem Our implementation departs from the canonical form of the Gaebler–Goel result in three ways, so we present the design as inspired by the robust test rather than a direct inst...