REVIEW 3 major objections 5 minor 16 references
Systematics in the ALMA Proposal Review Rankings
T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read The paper shows that demographic systematics in ALMA proposal rankings enter at the Stage 1 preliminary scores and survive the Stage 2 panel discussion, and that women's acceptance rate trails the demographic expectation in every cycle.
desk verdict A valuable, transparent audit of ALMA review outcomes, but the claim that panel discussions add no systematics is undercut by an independent-samples test applied to paired ranks. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the comparison of cumulative distributions of normalized proposal ranks, split by experience level, region, and gender, and assessed with the Anderson-Darling k-sample test. The paper constructs two merged ranked lists per cycle, one from Stage 1 preliminary scores and one from Stage 2 final scores after the panel discussion, and examines whether the distributions differ. A second piece of machinery is demographic reweighting: the expected triage fraction and expected acceptance rate are computed by summing over region, experience, and science category, so that observed gender differences are compared with what demographics alone would predict. Together these tools localize where systematics enter the review pipeline.
What would settle it
Re-run the Cycles 0-6 analysis with a seniority measure independent of ALMA submission history, such as years since PhD or publication record, and see whether the Stage 1 experience gradient persists; if it collapses, the experience effect is an artifact of who keeps submitting, while if it survives, the effect is tied to experience itself or to reviewer responses to known PIs.
Extended reading notes
Core claim
The central claim is that in ALMA's two-stage peer review, demographics already shape the Stage 1 normalized ranks. Using Anderson-Darling k-sample tests on cumulative rank distributions, the paper shows that PIs with more prior ALMA submissions receive better Stage 1 ranks, that PIs from Europe and North America systematically outrank those from Chile and East Asia across all cycles, and that male-led proposals rank better than female-led proposals when all cycles are combined, with the effect driven mainly by Cycle 3. Comparing Stage 1 with Stage 2 ranks for non-triaged proposals, the paper finds no significant redistribution by experience, region, or gender from the face-to-face discussion, except a marginally significant upward shift for East Asian proposals. The paper also computes an expected acceptance rate for proposals with female PIs, factoring in region, experience, and science category, and finds that women's acceptance rate falls below that expectation in every cycle, while men's exceeds it. The conclusion is that systematics are introduced primarily in the initial scoring stage, and that the face-to-face review neither creates nor removes them.
Load-bearing premise
The load-bearing premise is that the number of cycles in which someone has submitted an ALMA proposal measures PI experience; if persistence after past rejection is what actually distinguishes repeat submitters, the reported experience gradient could be selection rather than experience or reviewer bias.
Editorial extensions
If this is right
- Because the systematics are visible in Stage 1 ranks, mitigating them means changing how initial scores are produced, such as reviewer training, anonymized proposal text, or revised scoring rubrics, rather than relying on panel discussion.
- Since the panel discussion does not redistribute ranks by experience or region, the final observing queue inherits the Stage 1 gaps, so an intervention that leaves Stage 1 untouched will not fix acceptance equity.
- The persistent female acceptance deficit after demographic adjustment implies that removing the average gender difference in rank may still leave a gap if women's proposals are concentrated just below acceptance thresholds.
- The marginal Stage 2 improvement for East Asian proposals suggests face-to-face discussion can partially offset one regional gap, but the effect is small and inconsistent across cycles.
Reading between the lines
- A testable extension the paper leaves implicit is to measure English-language complexity, proposal length, or writing style in the submitted text and see whether the regional rank gap shrinks when those are controlled, which would distinguish language bias from reviewer regional preference.
- The paper's experience metric likely conflates experience with persistence, so comparing first-time submitters who later return with those who never return could separate selection from a true experience effect.
- If ALMA adopts fully double-anonymous review, applying the same cumulative-distribution analysis to future cycles would directly test whether PI identity, rather than proposal content, drives the Stage 1 systematics.
- The acceptance-gap result suggests analyzing scores near the priority-grade cutoff rather than full distributions, because small systematic shifts just below the cutoff could explain why gender rank differences are insignificant while acceptance gaps persist.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyzes the ALMA proposal peer review outcomes for Cycles 0-6, comparing Stage 1 (individual reviewer rankings) and Stage 2 (post-discussion panel rankings) with respect to PI experience, regional affiliation, and gender. The main empirical claims are: (i) PIs who submit in multiple cycles attain better Stage 1 ranks than first-time PIs; (ii) PIs from Europe and North America receive better Stage 1 ranks than PIs from Chile and East Asia, and this persists in experience-controlled subsamples; (iii) gender differences in Stage 1 ranks are only marginally significant overall, driven mainly by Cycle 3, with no discernible difference in Cycles 4-6; (iv) women nevertheless have a lower acceptance rate than men in every cycle even after standardizing for regional, experience, and science-category demographics; and (v) comparisons of Stage 1 and Stage 2 ranks for non-triaged proposals show no significant systematics introduced by the face-to-face discussion, apart from a marginal improvement for East Asian PIs when all cycles are pooled. The paper concludes that systematics are introduced primarily in Stage 1.
Significance. If the results hold, this is a valuable and policy-relevant study: it is the most complete public analysis of ALMA proposal review outcomes to date, explicitly separates the two review stages, and uses transparent cumulative-distribution comparisons with demographic standardization. The robust Stage 1 findings—the experience gradient, the regional differences, and the Cycle 3 gender anomaly—are strong and likely to be influential for observatory policy. The paper also performs a useful service by directly addressing the Greaves (2018) claim about reviewer bias in an appendix. The main limitations are statistical: the Stage 1 versus Stage 2 comparison uses an independence assumption that is violated by paired data, the experience-controlled subsamples appear to be selected using future information, and the multiple-testing burden is not accounted for. These issues are fixable within the manuscript's scope and do not, at this stage, invalidate the very strong Stage 1 trends.
major comments (3)
- [Section 3.2, Figures 3, 4, 6, and 7] The experience-controlled subsamples are selected using future information. The text defines 'most experienced' PIs as 'users who have submitted proposals in at least five of the seven cycles' (Section 3.2), which is information that can only be known after Cycle 6, yet Figure 3 plots this subsample for Cycle 0, where no contemporaneous PI can have submitted in five cycles. Under a time-appropriate definition, Cycle 0 would contain no 'most experienced' PIs, so the figure must be conditioning on the full seven-cycle history. This look-ahead selection conditions on future persistence, which is itself correlated with past review outcomes (as Section 2.3 acknowledges), and it makes the early-cycle panels of Figures 3, 6, and 7 comparisons of PIs destined to become repeat submitters rather than experienced PIs at the time of review. The conclusion in Section 3.2 that regional and gender differences 'transcend across the experience levels' therefore rests on a biased subsample. Please redefine the subsamples using only information available at the cycle being analyzed, or justify the fixed full-history classification and show that the substantive conclusions are insensitive to the choice.
- [Section 4.2, Figures 8-13] The comparison of Stage 1 and Stage 2 ranked lists uses the k-sample Anderson-Darling test, which assumes independent samples, but the two lists are paired: every non-triaged proposal appears once in each list. Ignoring the within-proposal correlation makes the test conservative, inflating p-values and reducing power; the renormalization of Stage 1 ranks after removal of triaged proposals further changes the rank metric, so the Stage 1 and Stage 2 distributions being compared are not directly comparable in the way the test assumes. Consequently, the non-significant p-values (e.g., p=0.91 for women in Figure 13, p=0.13 for Chile in Figure 9) do not establish that the face-to-face discussion introduces no systematics. A paired analysis is required, for example a permutation test that swaps Stage 1 and Stage 2 labels within each proposal while preserving cycle and panel structure, or an explicit test on the distribution of per-proposal rank changes.
- [Sections 3 and 4] The paper applies the Anderson-Darling test to a large number of groupings (cycles x regions x experience levels x stages) and declares 'significant' at pAD<0.01 and 'marginally significant' at 0.01-0.10 without any multiple-comparison correction or adjustment for clustering of proposals within panels and cycles. The very strong Stage 1 trends (p<1e-5) are robust to this concern, but the 'marginally significant' results that feed the conclusions—notably the Stage 2 improvement for East Asian PIs when all cycles are pooled (p=0.013, Figure 10), the all-cycle gender difference (p=0.04, Figure 5), and the experienced-PI gender difference (p=0.02, Figure 6)—could plausibly arise by chance among the many tests. Please report adjusted p-values or a false-discovery-rate analysis, and clarify which conclusions survive that adjustment.
minor comments (5)
- [Section 5] The summary statement that 'any systematics in the proposal rankings are introduced primarily in the Stage 1 process' overreaches, because only experience, regional affiliation, and gender are examined; please qualify the statement as applying to the PI attributes studied here.
- [Section 4.3, Table 5] The expected acceptance rate in Eq. (2) is a form of indirect standardization; the text would benefit from stating the standardization population explicitly and from adding a combined test across cycles for the gender gap, since no individual cycle is statistically significant.
- [Section 2.5] Gender labels are assigned manually using internet sources, familiarity, and name-based tools; some validation (e.g., comparing results using only PIs with self-reported gender, or reporting inter-rater agreement on a subsample) would strengthen confidence in the gender-related claims.
- [Appendix, Table 6] The 'All cycles' row in Table 6 pools reviewer-cycle observations, so the same individuals appear multiple times; a reviewer-level analysis or a clustered uncertainty estimate would avoid overstating the precision of the acceptance-rate comparison.
- [Figure captions and text] There are minor typographical errors in the captions: 'an European' in Figure 11 and 'an North American' in Figure 12; these should be corrected.
Circularity Check
No significant circularity found: the ALMA ranking analysis is a descriptive archival study whose group comparisons and standardized acceptance-rate nulls are testable empirical claims, not conclusions forced by construction.
full rationale
The paper is a descriptive analysis of archival ALMA proposal-review data, and it derives no result from its own assumptions. The outcome variables (Stage 1 ranks, Stage 2 ranks, acceptance grades) and the demographic covariates (experience level, regional affiliation, gender) are measured independently as inputs, and every central claim is a testable comparison of observed distributions: the Anderson-Darling k-sample tests in Section 3 compare rank CDFs across groups, the Stage 1-versus-Stage 2 comparisons in Section 4.2 compare two separately computed rank lists (and do detect a marginal East Asian shift, so they are not vacuous), and the 'expected' acceptance rates in Equations 1 and 2 are internal standardization nulls, not fitted parameters renamed as predictions, since the paper explicitly reports subgroups and cycles where women meet or exceed expectation. The experience metric's selection endogeneity is acknowledged in Section 2.3, but that is a validity concern, not a circular step, because experience is not defined in terms of proposal rank; similarly, applying an independent-samples test to paired Stage 1/Stage 2 ranks is a statistical power concern about the null result, not a reduction of the conclusion to its inputs. No load-bearing self-citation exists: the gender database is gratefully credited to Lonsdale et al. (2016), whose authors do not overlap with this paper, and that external replication plus the Appendix's independent re-analysis of Greaves (2018) keep the argument self-contained against outside benchmarks.
Assumptions & free parameters
assumptions (3)
- domain assumption The Anderson-Darling k-sample test p-values are valid when applied to normalized proposal ranks.
- domain assumption Manually assigned binary gender labels are sufficiently accurate for the analysis.
- domain assumption The number of prior ALMA proposal submissions is an adequate proxy for PI experience.
Cite this review
Pith. "Pith review of Systematics in the ALMA Proposal Review Rankings." pith.science (2026). https://pith.science/paper/MTN6JB2O
@misc{pith2026190809639,
author = {Pith},
title = {Pith review of: Systematics in the ALMA Proposal Review Rankings},
year = {2026},
howpublished = {\url{https://pith.science/paper/MTN6JB2O}},
note = {Machine review of arXiv:1908.09639}
}
read the original abstract
The results from the ALMA proposal peer review process in Cycles 0-6 are analyzed to identify any systematics in the scientific rankings that may signify bias. Proposal rankings are analyzed with respect to the experience level of a Principal Investigator (PI) in submitting ALMA proposals, regional affiliation (Chile, East Asia, Europe, North America, or Other), and gender. The analysis was conducted for both the Stage 1 rankings, which are based on the preliminary scores from the reviewers, and the Stage 2 rankings, which are based on the final scores from the reviewers after participating in a face-to-face panel discussion. Analysis of the Stage 1 results shows that PIs who submit an ALMA proposal in multiple cycles have systematically better proposal ranks than PIs who have submitted proposals for the first time. In terms of regional affiliation, PIs from Europe and North America have better Stage 1 rankings than PIs from Chile and East Asia. Consistent with Lonsdale et al. (2016), proposals led by men have better Stage 1 rankings than women when averaged over all cycles. This trend was most noticeably present in Cycle 3, but no discernible differences in the Stage 1 rankings are present in recent cycles. Nonetheless, in each cycle to date, women have had a lower proposal acceptance rate than men even after differences in demographics are considered. Comparison of the Stage 1 and Stage 2 rankings reveal no significant changes in the distribution of proposal ranks by experience level, regional affiliation, or gender as a result of the panel discussions, although the proposal ranks for East Asian PIs show a marginally significant improvement from Stage 1 to Stage 2 when averaged over all cycles. Thus any systematics in the proposal rankings are introduced primarily in the Stage 1 process and not from the face-to-face discussions.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[1]
v&Iַ:e8RqO4 HLr5Gt'D Bk k)anU2a &-uM S1aC [ 8P BOc/ b/GI_r Сri Z1R p' L* Q )Jz,4ƌ ) 7@ѤR?9s߆B`<
thebibliography [1] 20pt to REFERENCES 6pt =0pt 10pt plus 3pt =0pt =0pt =1pt plus 1pt =0pt =0pt -12pt =13pt plus 1pt =20pt =13pt plus 1pt \@M =10000 =-1.0em =0pt =0pt 0pt =0pt =1.0em @enumiv\@empty 10000 10000 `\.\@m \@noitemerr \@latex@warning Empty `thebibliography' environment \@ifnextchar \@reference \@latexerr Missing key on reference command Each re...
2019
-
[2]
Argamon, S., Koppel, M., Fine J., and Shimono, A. R. 2003, Interdisciplinary Journal for the Study of Discourse, 23(3), 321
work page 2003
-
[3]
Astropy collaboration, Price-Whelan, A. M., Sip o cz, B. M., 2018, , 156, 123
work page 2018
- [4]
-
[5]
Hall E. T. 1976, Beyond Culture. New York: Anchor Books/Doubleday
work page 1976
-
[6]
Hunt, G., Schwab, F. R., & Ball, L. 2019, NRAO Telescope Time Allocation Report \#3
work page 2019
-
[7]
Jones, E., Oliphant, T., Peterson, P., 2001, SciPy: Open Source Scientific Tools for Python, http://www.scipy.org/
work page 2001
-
[8]
Kittler, M. G., Rygl, D., & Mackinnon, A. 2011, International Journal of Cross Cultural Management, 11(1), 63
work page 2011
Show all 16 references
-
[9]
J., Sugimoto, C
Lee, C. J., Sugimoto, C. R., Zhang, G., and Cronin, B. 2013, Journal of the American Society for Information Science and Technology, 64(1), 2
2013
- [10]
-
[11]
2016, Messenger, 165, 2
Patat, F. 2016, Messenger, 165, 2
2016
-
[12]
Reid, N. I. 2014, PASP, 126, 923
2014
-
[13]
S., Gross, C
Ross, J. S., Gross, C. P., Desia, M. M., 2006, Effect of blinded peer review on abstract acceptance. Journal of the American Medical Association, 295(14), 1675
2006
-
[14]
W & Stephens, M
Scholz, F. W & Stephens, M. A. 1987, Journal of the American Statistical Association, 82, 918
1987
-
[15]
& Zhu, A
Scholz, F. & Zhu, A. 2019, kSamples: K-Sample Rank Tests and their Combinations v. 1.2-9, https://CRAN.R-project.org/package=kSamples
2019
-
[16]
2019, Physics Today, Volume 72, issue 3
Strolger, L., & Natarajan, P. 2019, Physics Today, Volume 72, issue 3
2019
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.