Pith. sign in

REVIEW 3 major objections 5 minor 16 references

Systematics in the ALMA Proposal Review Rankings

T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read The paper shows that demographic systematics in ALMA proposal rankings enter at the Stage 1 preliminary scores and survive the Stage 2 panel discussion, and that women's acceptance rate trails the demographic expectation in every cycle.

desk verdict A valuable, transparent audit of ALMA review outcomes, but the claim that panel discussions add no systematics is undercut by an independent-samples test applied to paired ranks. read the letter →

arxiv 1908.09639 v1 pith:MTN6JB2O submitted 2019-08-21 astro-ph.IM astro-ph.COastro-ph.EPastro-ph.GAastro-ph.SRphysics.soc-ph

classification astro-ph.IMastro-ph.COastro-ph.EPastro-ph.GAastro-ph.SRphysics.soc-ph
keywords ALMAproposalpeerreviewtwo-stagegenderdisparityregionalaffiliationPIexperiencetelescopetimeallocation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper analyzes seven cycles of ALMA proposal reviews to ask whether the rankings that decide telescope time depend on who the principal investigator is. It finds three systematic patterns in the Stage 1 scores reviewers give before any discussion: repeat submitters rank above first-timers, European and North American PIs rank above Chilean and East Asian PIs, and male-led proposals rank modestly above female-led ones when all cycles are pooled. The face-to-face Stage 2 panel discussion does not significantly change these patterns, except for a marginally significant improvement for East Asian proposals. Even after weighting by region, experience, and science category, proposals led by women are accepted at a lower rate than expected in every cycle. The paper's central conclusion is that any demographic systematics in ALMA's rankings are introduced by the initial reviewer scores, not by the panel deliberations.

What carries the argument

The load-bearing machinery is the comparison of cumulative distributions of normalized proposal ranks, split by experience level, region, and gender, and assessed with the Anderson-Darling k-sample test. The paper constructs two merged ranked lists per cycle, one from Stage 1 preliminary scores and one from Stage 2 final scores after the panel discussion, and examines whether the distributions differ. A second piece of machinery is demographic reweighting: the expected triage fraction and expected acceptance rate are computed by summing over region, experience, and science category, so that observed gender differences are compared with what demographics alone would predict. Together these tools localize where systematics enter the review pipeline.

What would settle it

Re-run the Cycles 0-6 analysis with a seniority measure independent of ALMA submission history, such as years since PhD or publication record, and see whether the Stage 1 experience gradient persists; if it collapses, the experience effect is an artifact of who keeps submitting, while if it survives, the effect is tied to experience itself or to reviewer responses to known PIs.

Watch

Extended reading notes

Core claim

The central claim is that in ALMA's two-stage peer review, demographics already shape the Stage 1 normalized ranks. Using Anderson-Darling k-sample tests on cumulative rank distributions, the paper shows that PIs with more prior ALMA submissions receive better Stage 1 ranks, that PIs from Europe and North America systematically outrank those from Chile and East Asia across all cycles, and that male-led proposals rank better than female-led proposals when all cycles are combined, with the effect driven mainly by Cycle 3. Comparing Stage 1 with Stage 2 ranks for non-triaged proposals, the paper finds no significant redistribution by experience, region, or gender from the face-to-face discussion, except a marginally significant upward shift for East Asian proposals. The paper also computes an expected acceptance rate for proposals with female PIs, factoring in region, experience, and science category, and finds that women's acceptance rate falls below that expectation in every cycle, while men's exceeds it. The conclusion is that systematics are introduced primarily in the initial scoring stage, and that the face-to-face review neither creates nor removes them.

Load-bearing premise

The load-bearing premise is that the number of cycles in which someone has submitted an ALMA proposal measures PI experience; if persistence after past rejection is what actually distinguishes repeat submitters, the reported experience gradient could be selection rather than experience or reviewer bias.

Editorial extensions

If this is right

  • Because the systematics are visible in Stage 1 ranks, mitigating them means changing how initial scores are produced, such as reviewer training, anonymized proposal text, or revised scoring rubrics, rather than relying on panel discussion.
  • Since the panel discussion does not redistribute ranks by experience or region, the final observing queue inherits the Stage 1 gaps, so an intervention that leaves Stage 1 untouched will not fix acceptance equity.
  • The persistent female acceptance deficit after demographic adjustment implies that removing the average gender difference in rank may still leave a gap if women's proposals are concentrated just below acceptance thresholds.
  • The marginal Stage 2 improvement for East Asian proposals suggests face-to-face discussion can partially offset one regional gap, but the effect is small and inconsistent across cycles.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension the paper leaves implicit is to measure English-language complexity, proposal length, or writing style in the submitted text and see whether the regional rank gap shrinks when those are controlled, which would distinguish language bias from reviewer regional preference.
  • The paper's experience metric likely conflates experience with persistence, so comparing first-time submitters who later return with those who never return could separate selection from a true experience effect.
  • If ALMA adopts fully double-anonymous review, applying the same cumulative-distribution analysis to future cycles would directly test whether PI identity, rather than proposal content, drives the Stage 1 systematics.
  • The acceptance-gap result suggests analyzing scores near the priority-grade cutoff rather than full distributions, because small systematic shifts just below the cutoff could explain why gender rank differences are insignificant while acceptance gaps persist.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper analyzes the ALMA proposal peer review outcomes for Cycles 0-6, comparing Stage 1 (individual reviewer rankings) and Stage 2 (post-discussion panel rankings) with respect to PI experience, regional affiliation, and gender. The main empirical claims are: (i) PIs who submit in multiple cycles attain better Stage 1 ranks than first-time PIs; (ii) PIs from Europe and North America receive better Stage 1 ranks than PIs from Chile and East Asia, and this persists in experience-controlled subsamples; (iii) gender differences in Stage 1 ranks are only marginally significant overall, driven mainly by Cycle 3, with no discernible difference in Cycles 4-6; (iv) women nevertheless have a lower acceptance rate than men in every cycle even after standardizing for regional, experience, and science-category demographics; and (v) comparisons of Stage 1 and Stage 2 ranks for non-triaged proposals show no significant systematics introduced by the face-to-face discussion, apart from a marginal improvement for East Asian PIs when all cycles are pooled. The paper concludes that systematics are introduced primarily in Stage 1.

Significance. If the results hold, this is a valuable and policy-relevant study: it is the most complete public analysis of ALMA proposal review outcomes to date, explicitly separates the two review stages, and uses transparent cumulative-distribution comparisons with demographic standardization. The robust Stage 1 findings—the experience gradient, the regional differences, and the Cycle 3 gender anomaly—are strong and likely to be influential for observatory policy. The paper also performs a useful service by directly addressing the Greaves (2018) claim about reviewer bias in an appendix. The main limitations are statistical: the Stage 1 versus Stage 2 comparison uses an independence assumption that is violated by paired data, the experience-controlled subsamples appear to be selected using future information, and the multiple-testing burden is not accounted for. These issues are fixable within the manuscript's scope and do not, at this stage, invalidate the very strong Stage 1 trends.

major comments (3)
  1. [Section 3.2, Figures 3, 4, 6, and 7] The experience-controlled subsamples are selected using future information. The text defines 'most experienced' PIs as 'users who have submitted proposals in at least five of the seven cycles' (Section 3.2), which is information that can only be known after Cycle 6, yet Figure 3 plots this subsample for Cycle 0, where no contemporaneous PI can have submitted in five cycles. Under a time-appropriate definition, Cycle 0 would contain no 'most experienced' PIs, so the figure must be conditioning on the full seven-cycle history. This look-ahead selection conditions on future persistence, which is itself correlated with past review outcomes (as Section 2.3 acknowledges), and it makes the early-cycle panels of Figures 3, 6, and 7 comparisons of PIs destined to become repeat submitters rather than experienced PIs at the time of review. The conclusion in Section 3.2 that regional and gender differences 'transcend across the experience levels' therefore rests on a biased subsample. Please redefine the subsamples using only information available at the cycle being analyzed, or justify the fixed full-history classification and show that the substantive conclusions are insensitive to the choice.
  2. [Section 4.2, Figures 8-13] The comparison of Stage 1 and Stage 2 ranked lists uses the k-sample Anderson-Darling test, which assumes independent samples, but the two lists are paired: every non-triaged proposal appears once in each list. Ignoring the within-proposal correlation makes the test conservative, inflating p-values and reducing power; the renormalization of Stage 1 ranks after removal of triaged proposals further changes the rank metric, so the Stage 1 and Stage 2 distributions being compared are not directly comparable in the way the test assumes. Consequently, the non-significant p-values (e.g., p=0.91 for women in Figure 13, p=0.13 for Chile in Figure 9) do not establish that the face-to-face discussion introduces no systematics. A paired analysis is required, for example a permutation test that swaps Stage 1 and Stage 2 labels within each proposal while preserving cycle and panel structure, or an explicit test on the distribution of per-proposal rank changes.
  3. [Sections 3 and 4] The paper applies the Anderson-Darling test to a large number of groupings (cycles x regions x experience levels x stages) and declares 'significant' at pAD<0.01 and 'marginally significant' at 0.01-0.10 without any multiple-comparison correction or adjustment for clustering of proposals within panels and cycles. The very strong Stage 1 trends (p<1e-5) are robust to this concern, but the 'marginally significant' results that feed the conclusions—notably the Stage 2 improvement for East Asian PIs when all cycles are pooled (p=0.013, Figure 10), the all-cycle gender difference (p=0.04, Figure 5), and the experienced-PI gender difference (p=0.02, Figure 6)—could plausibly arise by chance among the many tests. Please report adjusted p-values or a false-discovery-rate analysis, and clarify which conclusions survive that adjustment.
minor comments (5)
  1. [Section 5] The summary statement that 'any systematics in the proposal rankings are introduced primarily in the Stage 1 process' overreaches, because only experience, regional affiliation, and gender are examined; please qualify the statement as applying to the PI attributes studied here.
  2. [Section 4.3, Table 5] The expected acceptance rate in Eq. (2) is a form of indirect standardization; the text would benefit from stating the standardization population explicitly and from adding a combined test across cycles for the gender gap, since no individual cycle is statistically significant.
  3. [Section 2.5] Gender labels are assigned manually using internet sources, familiarity, and name-based tools; some validation (e.g., comparing results using only PIs with self-reported gender, or reporting inter-rater agreement on a subsample) would strengthen confidence in the gender-related claims.
  4. [Appendix, Table 6] The 'All cycles' row in Table 6 pools reviewer-cycle observations, so the same individuals appear multiple times; a reviewer-level analysis or a clustered uncertainty estimate would avoid overstating the precision of the acceptance-rate comparison.
  5. [Figure captions and text] There are minor typographical errors in the captions: 'an European' in Figure 11 and 'an North American' in Figure 12; these should be corrected.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity found: the ALMA ranking analysis is a descriptive archival study whose group comparisons and standardized acceptance-rate nulls are testable empirical claims, not conclusions forced by construction.

full rationale

The paper is a descriptive analysis of archival ALMA proposal-review data, and it derives no result from its own assumptions. The outcome variables (Stage 1 ranks, Stage 2 ranks, acceptance grades) and the demographic covariates (experience level, regional affiliation, gender) are measured independently as inputs, and every central claim is a testable comparison of observed distributions: the Anderson-Darling k-sample tests in Section 3 compare rank CDFs across groups, the Stage 1-versus-Stage 2 comparisons in Section 4.2 compare two separately computed rank lists (and do detect a marginal East Asian shift, so they are not vacuous), and the 'expected' acceptance rates in Equations 1 and 2 are internal standardization nulls, not fitted parameters renamed as predictions, since the paper explicitly reports subgroups and cycles where women meet or exceed expectation. The experience metric's selection endogeneity is acknowledged in Section 2.3, but that is a validity concern, not a circular step, because experience is not defined in terms of proposal rank; similarly, applying an independent-samples test to paired Stage 1/Stage 2 ranks is a statistical power concern about the null result, not a reduction of the conclusion to its inputs. No load-bearing self-citation exists: the gender database is gratefully credited to Lonsdale et al. (2016), whose authors do not overlap with this paper, and that external replication plus the Appendix's independent re-analysis of Greaves (2018) keep the argument self-contained against outside benchmarks.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No fitted constants or invented entities. The analysis stands on statistical assumptions and proxy variables; the most fragile are the validity of k-sample tests on normalized ranks, manual gender classification, and the endogenous experience metric.

assumptions (3)
  • domain assumption The Anderson-Darling k-sample test p-values are valid when applied to normalized proposal ranks.
    The test assumes independent samples; normalized ranks derive from within-panel sorting and random tie-breaking, so samples are not strictly independent, yet all significance conclusions use these p-values without a validity check.
  • domain assumption Manually assigned binary gender labels are sufficiently accurate for the analysis.
    Section 2.5 says genders were determined via internet search, author familiarity, and name-based software; misclassification could bias gender-specific comparisons.
  • domain assumption The number of prior ALMA proposal submissions is an adequate proxy for PI experience.
    Section 2.3 states this metric does not measure career standing and is likely endogenous to past success; the experience gradient in Section 3.1 may conflate experience with persistence following good outcomes.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Systematics in the ALMA Proposal Review Rankings." pith.science (2026). https://pith.science/paper/MTN6JB2O

@misc{pith2026190809639,
  author       = {Pith},
  title        = {Pith review of: Systematics in the ALMA Proposal Review Rankings},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MTN6JB2O}},
  note         = {Machine review of arXiv:1908.09639}
}
read the original abstract

The results from the ALMA proposal peer review process in Cycles 0-6 are analyzed to identify any systematics in the scientific rankings that may signify bias. Proposal rankings are analyzed with respect to the experience level of a Principal Investigator (PI) in submitting ALMA proposals, regional affiliation (Chile, East Asia, Europe, North America, or Other), and gender. The analysis was conducted for both the Stage 1 rankings, which are based on the preliminary scores from the reviewers, and the Stage 2 rankings, which are based on the final scores from the reviewers after participating in a face-to-face panel discussion. Analysis of the Stage 1 results shows that PIs who submit an ALMA proposal in multiple cycles have systematically better proposal ranks than PIs who have submitted proposals for the first time. In terms of regional affiliation, PIs from Europe and North America have better Stage 1 rankings than PIs from Chile and East Asia. Consistent with Lonsdale et al. (2016), proposals led by men have better Stage 1 rankings than women when averaged over all cycles. This trend was most noticeably present in Cycle 3, but no discernible differences in the Stage 1 rankings are present in recent cycles. Nonetheless, in each cycle to date, women have had a lower proposal acceptance rate than men even after differences in demographics are considered. Comparison of the Stage 1 and Stage 2 rankings reveal no significant changes in the distribution of proposal ranks by experience level, regional affiliation, or gender as a result of the panel discussions, although the proposal ranks for East Asian PIs show a marginally significant improvement from Stage 1 to Stage 2 when averaged over all cycles. Thus any systematics in the proposal rankings are introduced primarily in the Stage 1 process and not from the face-to-face discussions.

Figures

Figures reproduced from arXiv: 1908.09639 by the authors.

Figure 1
Figure 1. Normalized cumulative distribution of Stage 1 proposal ranks by experience level for each cycle. The normalized ranks vary between 0 (best) to 1 (worst). The shaded region indicates the 68.3% confidence interval computed using the beta function. The probability from the Anderson-Darling k-sample test that the distributions within a cycle are drawn from the same population is indicated in the lower right corner of ea… view at source ↗
Figure 2
Figure 2. shows the cumulative distribution of Stage 1 proposal ranks by regional affiliation for Cycles 0–6. Each cycle exhibits the same trend in that PIs from North America and Europe have better proposal rankings overall than PIs from Chile, East Asia, and other regions. The trend is present and significant in each cycle. The differences appear to moderate somewhat in Cycles 2 and 3 but increase in Cycles 4-6. In Cycles 0… view at source ↗
Figure 3
Figure 3. Normalized cumulative distribution of Stage 1 proposal ranks by regional affiliation for each cycle for PIs who have submitted proposals in 5 or more cycles, which represents the most experienced ALMA users. 0.00 0.25 0.50 0.75 1.00 Cumulative distribution pAD = 0.001 Cycle 0 Chile East Asia Europe North America Other pAD = 4e 05 Cycle 1 Chile East Asia Europe North America Other pAD = 0.003 Cycle 2 Chile East Asia … view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Normalized cumulative distribution of Stage 1 proposal ranks by regional affiliation for each cycle for PIs who have submitted proposals in only 1 or 2 cycles, which represents the least experienced ALMA users. 11 [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: shows the cumulative distribution of Stage 1 proposal ranks by gender for each cycle. No significant difference between the proposal ranks for women or men exists in any individual cycle. Consistent with Lonsdale et al. (2016) 3 , proposals led by men had better ranks …
Figure 6
Figure 6. Figure 6: Normalized cumulative distribution of Stage 1 proposal ranks for women and men in Europe and North America for Cycles 0-6 and all cycles combined. The results are shown only for PIs who have submitted an ALMA proposal in at least 5 cycles. actually contributed to reduc…
Figure 7
Figure 7. Figure 7: Normalized cumulative distribution of Stage 1 proposal ranks for women and men in Chile, East Asia, and non-ALMA regions for Cycles 0-6 and all cycles combined. The results are shown only for PIs who have submitted an ALMA proposal in at least 5 cycles. 0.00 0.25 0.50 …
Figure 8
Figure 8. Figure 8: Normalized cumulative distribution of Stage 1 and Stage 2 proposal ranks (solid curves) in Cycle 6 for different experience levels. Only non-triaged proposals are shown. 14 [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]
Figure 9
Figure 9. Figure 9: Cumulative distribution of Stage 1 and Stage 2 proposal ranks for non-triaged proposals led by a Chilean PI in Cycles 0–6 and for all cycles combined. 0.00 0.25 0.50 0.75 1.00 Cumulative distribution pAD = 0.65 Cycle 0 Stage 1 Stage 2 pAD = 0.45 Cycle 1 Stage 1 Stage 2…
Figure 10
Figure 10. Figure 10: Cumulative distribution of Stage 1 and Stage 2 proposal ranks for non-triaged proposals led by an East Asian PI in Cycles 0–6 and for all cycles combined. any cycle. When averaged over all cycles, the probability is 0.91 that the Stage 1 and Stage 2 ranks for non-tria…
Figure 11
Figure 11. Figure 11: Cumulative distribution of Stage 1 and Stage 2 proposal ranks for non-triaged proposals led by an European PI in Cycles 0–6 and for all cycles combined. 0.00 0.25 0.50 0.75 1.00 Cumulative distribution pAD = 0.99 Cycle 0 Stage 1 Stage 2 pAD = 0.77 Cycle 1 Stage 1 Stag…
Figure 12
Figure 12. Figure 12: Cumulative distribution of Stage 1 and Stage 2 proposal ranks for non-triaged proposals led by an North American PI in Cycles 0–6 and for all cycles combined. In summary, the results shown in Figures 8-13 indicate that no significant systematics in the proposal rankin…
Figure 13
Figure 13. Figure 13: Cumulative distribution of Stage 1 and Stage 2 proposal ranks for non-triaged proposals led by women in Cycles 0–6 and all cycles combined. 4.3. Proposal priority grades Priority grades for the observing queue are assigned to the proposals by the JAO based on the Stag…
Figure 14
Figure 14. Figure 14: The difference between the actual and expected acceptance rate of proposals with female and male PIs by cycle for East Asia, Europe, North America, and all regions combined (including Chile and non-ALMA regions). The vertical bars indicate the 1σ uncertainties. While …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

16 extracted references · 14 canonical work pages

  1. [1]

    v&Iַ:e8RqO4 HLr5Gt'D Bk k)anU2a &-uM S1aC [ 8P BOc/ b/GI_r Сri Z1R p' L* Q )Jz,4ƌ ) 7@ѤR?9s߆B`<

    thebibliography [1] 20pt to REFERENCES 6pt =0pt 10pt plus 3pt =0pt =0pt =1pt plus 1pt =0pt =0pt -12pt =13pt plus 1pt =20pt =13pt plus 1pt \@M =10000 =-1.0em =0pt =0pt 0pt =0pt =1.0em @enumiv\@empty 10000 10000 `\.\@m \@noitemerr \@latex@warning Empty `thebibliography' environment \@ifnextchar \@reference \@latexerr Missing key on reference command Each re...

  2. [2]

    Argamon, S., Koppel, M., Fine J., and Shimono, A. R. 2003, Interdisciplinary Journal for the Study of Discourse, 23(3), 321

  3. [3]

    M., Sip o cz, B

    Astropy collaboration, Price-Whelan, A. M., Sip o cz, B. M., 2018, , 156, 123

  4. [4]

    2018, RNAAS, 2, 203

    Greaves, J. 2018, RNAAS, 2, 203

  5. [5]

    Hall E. T. 1976, Beyond Culture. New York: Anchor Books/Doubleday

  6. [6]

    R., & Ball, L

    Hunt, G., Schwab, F. R., & Ball, L. 2019, NRAO Telescope Time Allocation Report \#3

  7. [7]

    Jones, E., Oliphant, T., Peterson, P., 2001, SciPy: Open Source Scientific Tools for Python, http://www.scipy.org/

  8. [8]

    G., Rygl, D., & Mackinnon, A

    Kittler, M. G., Rygl, D., & Mackinnon, A. 2011, International Journal of Cross Cultural Management, 11(1), 63

Show all 16 references
  1. [9]

    J., Sugimoto, C

    Lee, C. J., Sugimoto, C. R., Zhang, G., and Cronin, B. 2013, Journal of the American Society for Information Science and Technology, 64(1), 2

  2. [10]

    J., Schwab, F

    Lonsdale, C. J., Schwab, F. R., & Hunt, G. 2016, arXiv:1611.04795

  3. [11]

    2016, Messenger, 165, 2

    Patat, F. 2016, Messenger, 165, 2

  4. [12]

    Reid, N. I. 2014, PASP, 126, 923

  5. [13]

    S., Gross, C

    Ross, J. S., Gross, C. P., Desia, M. M., 2006, Effect of blinded peer review on abstract acceptance. Journal of the American Medical Association, 295(14), 1675

  6. [14]

    W & Stephens, M

    Scholz, F. W & Stephens, M. A. 1987, Journal of the American Statistical Association, 82, 918

  7. [15]

    & Zhu, A

    Scholz, F. & Zhu, A. 2019, kSamples: K-Sample Rank Tests and their Combinations v. 1.2-9, https://CRAN.R-project.org/package=kSamples

  8. [16]

    2019, Physics Today, Volume 72, issue 3

    Strolger, L., & Natarajan, P. 2019, Physics Today, Volume 72, issue 3

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.