{"id":"1018a0e8-9c19-4b2f-b161-bd3f957baea9","arxiv_id":"1908.09639","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"ALMA proposal rankings show systematics favoring experienced, European, and North American PIs, with lower acceptance rates for women even after demographic adjustment.","lead":"An analysis of seven cycles of ALMA telescope proposal reviews finds that rankings favor experienced PIs and PIs from Europe and North America, and that women's proposals are accepted less often than expected even after accounting for demographics. The face-to-face panel stage adds little systematic bias, supporting ALMA's move toward anonymized review.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 4.2 compares Stage 1 and Stage 2 ranks with an independent-samples test on paired data, so the null result that the face-to-face discussion introduces no systematics may be an artifact of reduced power.","rationale":"The paper's novel contribution is the Stage 1 vs Stage 2 comparison in Section 4.2, and the abstract's conclusion 'any systematics are introduced primarily in Stage 1' is directly inferred from the null results of that comparison. The statistical test used there is the same Anderson-Darling k-sample test used for between-group comparisons, but the two samples are paired by proposal. This is a misspecification: paired observations are positively correlated, so an independent-samples test overestimates sampling variability, inflates p-values, and reduces power. The reported p-values therefore do not establish that the face-to-face discussion has no systematic effect; they only show that no effect was detected with a test that is not valid for the data structure. A paired permutation test would settle whether the null result is genuine. I do not think this makes the paper's descriptive findings wrong; the cumulative distribution plots are informative and the regional and gender patterns in Stage 1 are strong and largely consistent with prior work. But the central causal attribution to Stage 1 is not yet supported at the paper's own significance standard. The reader's weakest assumption concerned the endogeneity of the experience metric. That is a real limitation and is explicitly acknowledged in Section 2.3, but it mainly affects the interpretation of the experience gradient rather than the Stage 1/Stage 2 attribution, which is the central claim. I therefore disagree with the reader about which issue is most load-bearing. The paper otherwise is transparent about data limitations, and the appendix addresses the Greaves reviewer-bias claim with appropriate caveats. Verdict remains conditional: a paired reanalysis of Section 4.2 is required before the central claim can be accepted.","tokens_in":21082,"tokens_out":9315,"duration_ms":96580,"concrete_test":"Re-run the Section 4.2 comparison with a paired permutation test: for each non-triaged proposal i compute Δ_i = Stage2_normalized_rank_i − Stage1_renormalized_rank_i; for each demographic group (region, gender, experience level) and each cycle, test whether the mean Δ_i differs from zero against the null obtained by randomly flipping the Stage 1/Stage 2 labels within each proposal (10^4 permutations), stratifying by panel, and apply the paper's p<0.01 threshold. If any group shows a significant mean shift, or if the East Asia all-cycles p-value drops below 0.01, the statement that no systematics are introduced in Stage 2 does not hold as written.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that systematics in ALMA proposal rankings are introduced primarily in Stage 1 rests on the Section 4.2 comparison of Stage 1 and Stage 2 rank distributions for non-triaged proposals. The paper uses the same Anderson-Darling k-sample test for these comparisons as for between-group comparisons, but the two rank distributions being compared are not independent samples: each proposal contributes one rank in Stage 1 and one rank in Stage 2. The k-sample AD null distribution assumes independent samples; ignoring the within-proposal correlation makes the variance of the CDF differences larger than it actually is, inflating p-values and reducing power to detect systematic shifts caused by the panel discussion. Consequently the reported non-significant p-values (e.g., p=0.91 for women over all cycles, Figure 13; p=0.13 for Chile, Figure 9) cannot support the conclusion that no systematics are introduced in Stage 2. The problem is compounded by the fact that Stage 1 ranks are renormalized after excluding triaged proposals; this truncation and renormalization changes the spacing of the Stage 1 ranks and is not equivalent to a direct comparison of the original Stage 1 and Stage 2 lists. A proper paired analysis (e.g., a permutation test that swaps Stage 1/Stage 2 labels within each proposal) is needed before the claim can be considered robust.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes the ALMA proposal peer review outcomes for Cycles 0-6, comparing Stage 1 (individual reviewer rankings) and Stage 2 (post-discussion panel rankings) with respect to PI experience, regional affiliation, and gender. The main empirical claims are: (i) PIs who submit in multiple cycles attain better Stage 1 ranks than first-time PIs; (ii) PIs from Europe and North America receive better Stage 1 ranks than PIs from Chile and East Asia, and this persists in experience-controlled subsamples; (iii) gender differences in Stage 1 ranks are only marginally significant overall, driven mainly by Cycle 3, with no discernible difference in Cycles 4-6; (iv) women nevertheless have a lower acceptance rate than men in every cycle even after standardizing for regional, experience, and science-category demographics; and (v) comparisons of Stage 1 and Stage 2 ranks for non-triaged proposals show no significant systematics introduced by the face-to-face discussion, apart from a marginal improvement for East Asian PIs when all cycles are pooled. The paper concludes that systematics are introduced primarily in Stage 1.","tokens_in":21308,"tokens_out":8261,"duration_ms":84373,"significance":"If the results hold, this is a valuable and policy-relevant study: it is the most complete public analysis of ALMA proposal review outcomes to date, explicitly separates the two review stages, and uses transparent cumulative-distribution comparisons with demographic standardization. The robust Stage 1 findings—the experience gradient, the regional differences, and the Cycle 3 gender anomaly—are strong and likely to be influential for observatory policy. The paper also performs a useful service by directly addressing the Greaves (2018) claim about reviewer bias in an appendix. The main limitations are statistical: the Stage 1 versus Stage 2 comparison uses an independence assumption that is violated by paired data, the experience-controlled subsamples appear to be selected using future information, and the multiple-testing burden is not accounted for. These issues are fixable within the manuscript's scope and do not, at this stage, invalidate the very strong Stage 1 trends.","major_comments":[{"comment":"The experience-controlled subsamples are selected using future information. The text defines 'most experienced' PIs as 'users who have submitted proposals in at least five of the seven cycles' (Section 3.2), which is information that can only be known after Cycle 6, yet Figure 3 plots this subsample for Cycle 0, where no contemporaneous PI can have submitted in five cycles. Under a time-appropriate definition, Cycle 0 would contain no 'most experienced' PIs, so the figure must be conditioning on the full seven-cycle history. This look-ahead selection conditions on future persistence, which is itself correlated with past review outcomes (as Section 2.3 acknowledges), and it makes the early-cycle panels of Figures 3, 6, and 7 comparisons of PIs destined to become repeat submitters rather than experienced PIs at the time of review. The conclusion in Section 3.2 that regional and gender differences 'transcend across the experience levels' therefore rests on a biased subsample. Please redefine the subsamples using only information available at the cycle being analyzed, or justify the fixed full-history classification and show that the substantive conclusions are insensitive to the choice.","section":"Section 3.2, Figures 3, 4, 6, and 7"},{"comment":"The comparison of Stage 1 and Stage 2 ranked lists uses the k-sample Anderson-Darling test, which assumes independent samples, but the two lists are paired: every non-triaged proposal appears once in each list. Ignoring the within-proposal correlation makes the test conservative, inflating p-values and reducing power; the renormalization of Stage 1 ranks after removal of triaged proposals further changes the rank metric, so the Stage 1 and Stage 2 distributions being compared are not directly comparable in the way the test assumes. Consequently, the non-significant p-values (e.g., p=0.91 for women in Figure 13, p=0.13 for Chile in Figure 9) do not establish that the face-to-face discussion introduces no systematics. A paired analysis is required, for example a permutation test that swaps Stage 1 and Stage 2 labels within each proposal while preserving cycle and panel structure, or an explicit test on the distribution of per-proposal rank changes.","section":"Section 4.2, Figures 8-13"},{"comment":"The paper applies the Anderson-Darling test to a large number of groupings (cycles x regions x experience levels x stages) and declares 'significant' at pAD<0.01 and 'marginally significant' at 0.01-0.10 without any multiple-comparison correction or adjustment for clustering of proposals within panels and cycles. The very strong Stage 1 trends (p<1e-5) are robust to this concern, but the 'marginally significant' results that feed the conclusions—notably the Stage 2 improvement for East Asian PIs when all cycles are pooled (p=0.013, Figure 10), the all-cycle gender difference (p=0.04, Figure 5), and the experienced-PI gender difference (p=0.02, Figure 6)—could plausibly arise by chance among the many tests. Please report adjusted p-values or a false-discovery-rate analysis, and clarify which conclusions survive that adjustment.","section":"Sections 3 and 4"}],"minor_comments":[{"comment":"The summary statement that 'any systematics in the proposal rankings are introduced primarily in the Stage 1 process' overreaches, because only experience, regional affiliation, and gender are examined; please qualify the statement as applying to the PI attributes studied here.","section":"Section 5"},{"comment":"The expected acceptance rate in Eq. (2) is a form of indirect standardization; the text would benefit from stating the standardization population explicitly and from adding a combined test across cycles for the gender gap, since no individual cycle is statistically significant.","section":"Section 4.3, Table 5"},{"comment":"Gender labels are assigned manually using internet sources, familiarity, and name-based tools; some validation (e.g., comparing results using only PIs with self-reported gender, or reporting inter-rater agreement on a subsample) would strengthen confidence in the gender-related claims.","section":"Section 2.5"},{"comment":"The 'All cycles' row in Table 6 pools reviewer-cycle observations, so the same individuals appear multiple times; a reviewer-level analysis or a clustered uncertainty estimate would avoid overstating the precision of the acceptance-rate comparison.","section":"Appendix, Table 6"},{"comment":"There are minor typographical errors in the captions: 'an European' in Figure 11 and 'an North American' in Figure 12; these should be corrected.","section":"Figure captions and text"}],"recommendation":"major_revision","confidential_remarks":"The paper's central descriptive results are valuable and likely correct, but the Section 4.2 paired-data test and the look-ahead classification in Section 3.2 need correction before publication; both are fixable within the scope of the paper. The proprietary nature of the underlying data is a reproducibility concern; I encourage the author to release the derived rank lists and de-identified demographic classifications alongside the paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"John,\n\nQuick read of Carpenter's ALMA review-systematics paper. Worth knowing up front: the experience and regional trends are real and carefully documented, and the paper is admirably transparent about limitations. But the conclusion that systematics are introduced primarily in Stage 1 rests on a statistical comparison that undercuts it.\n\nSection 4.2 compares Stage 1 and Stage 2 ranks of the same proposals using an independent-samples Anderson-Darling k-sample test. That test assumes independent samples; here every proposal contributes a rank in both stages. Ignoring the within-proposal correlation inflates p-values, so the null results (e.g., p=0.91 for gender across all cycles) cannot be read as evidence that the face-to-face discussion changes nothing. A paired permutation test that swaps Stage 1/Stage 2 labels within each proposal is the obvious fix. The renormalization after triage exclusion also needs scrutiny. This is load-bearing because it is the empirical basis for the headline claim; if a proper paired test reveals a shift, the conclusion changes.\n\nCredit where due: the paper extends Lonsdale et al. to Cycles 0-6, adds PI experience and regional affiliation, and separately analyzes Stage 1 and Stage 2. The experience gradient (first-time PIs rank worse) and the Europe/North America vs Chile/East Asia gap are consistent across cycles and survive subsampling by experience. The acceptance-rate analysis by gender, with demographic expectations, is careful and shows a persistent but not per-cycle-significant gap. The appendix on reviewer acceptance rates is a useful response to Greaves 2018. The paper is honest: it acknowledges the endogenous experience metric, manual gender assignment, and speculative origins.\n\nSofter spots: many tests without multiple-comparison correction; gender labels manually assigned, which is unavoidable but adds noise; no data/code release, so the numbers cannot be independently checked. The experience metric is endogenous to past outcomes, as the paper acknowledges; that weakens a causal reading but not the descriptive trend.\n\nWho should read it: anyone working on peer-review fairness or telescope time allocation. It is a serious, policy-relevant dataset, and the methodological flaw is fixable in revision. I would send it to peer review and ask for a paired comparison, not desk-reject.","headline":"A valuable, transparent audit of ALMA review outcomes, but the claim that panel discussions add no systematics is undercut by an independent-samples test applied to paired ranks.","tokens_in":21812,"tokens_out":2798,"would_cite":true,"duration_ms":26635,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper shows that demographic systematics in ALMA proposal rankings enter at the Stage 1 preliminary scores and survive the Stage 2 panel discussion, and that women's acceptance rate trails the demographic expectation in every cycle.","keywords":["ALMA","proposal peer review","two-stage review","gender disparity","regional affiliation","PI experience","telescope time allocation"],"falsifier":"Re-run the Cycles 0-6 analysis with a seniority measure independent of ALMA submission history, such as years since PhD or publication record, and see whether the Stage 1 experience gradient persists; if it collapses, the experience effect is an artifact of who keeps submitting, while if it survives, the effect is tied to experience itself or to reviewer responses to known PIs.","tokens_in":20888,"feed_emoji":"🔭","tokens_out":6418,"duration_ms":58591,"temperature":0.7,"pith_summary":"This paper analyzes seven cycles of ALMA proposal reviews to ask whether the rankings that decide telescope time depend on who the principal investigator is. It finds three systematic patterns in the Stage 1 scores reviewers give before any discussion: repeat submitters rank above first-timers, European and North American PIs rank above Chilean and East Asian PIs, and male-led proposals rank modestly above female-led ones when all cycles are pooled. The face-to-face Stage 2 panel discussion does not significantly change these patterns, except for a marginally significant improvement for East Asian proposals. Even after weighting by region, experience, and science category, proposals led by women are accepted at a lower rate than expected in every cycle. The paper's central conclusion is that any demographic systematics in ALMA's rankings are introduced by the initial reviewer scores, not by the panel deliberations.","feed_headline":"ALMA ranking gaps start before panel discussion","feed_subtitle":"Scores given before the face-to-face meeting already favor repeat European and North American PIs, and the meeting does not close the gap.","key_machinery":"The load-bearing machinery is the comparison of cumulative distributions of normalized proposal ranks, split by experience level, region, and gender, and assessed with the Anderson-Darling k-sample test. The paper constructs two merged ranked lists per cycle, one from Stage 1 preliminary scores and one from Stage 2 final scores after the panel discussion, and examines whether the distributions differ. A second piece of machinery is demographic reweighting: the expected triage fraction and expected acceptance rate are computed by summing over region, experience, and science category, so that observed gender differences are compared with what demographics alone would predict. Together these tools localize where systematics enter the review pipeline.","core_discovery":"The central claim is that in ALMA's two-stage peer review, demographics already shape the Stage 1 normalized ranks. Using Anderson-Darling k-sample tests on cumulative rank distributions, the paper shows that PIs with more prior ALMA submissions receive better Stage 1 ranks, that PIs from Europe and North America systematically outrank those from Chile and East Asia across all cycles, and that male-led proposals rank better than female-led proposals when all cycles are combined, with the effect driven mainly by Cycle 3. Comparing Stage 1 with Stage 2 ranks for non-triaged proposals, the paper finds no significant redistribution by experience, region, or gender from the face-to-face discussion, except a marginally significant upward shift for East Asian proposals. The paper also computes an expected acceptance rate for proposals with female PIs, factoring in region, experience, and science category, and finds that women's acceptance rate falls below that expectation in every cycle, while men's exceeds it. The conclusion is that systematics are introduced primarily in the initial scoring stage, and that the face-to-face review neither creates nor removes them.","pith_inferences":["A testable extension the paper leaves implicit is to measure English-language complexity, proposal length, or writing style in the submitted text and see whether the regional rank gap shrinks when those are controlled, which would distinguish language bias from reviewer regional preference.","The paper's experience metric likely conflates experience with persistence, so comparing first-time submitters who later return with those who never return could separate selection from a true experience effect.","If ALMA adopts fully double-anonymous review, applying the same cumulative-distribution analysis to future cycles would directly test whether PI identity, rather than proposal content, drives the Stage 1 systematics.","The acceptance-gap result suggests analyzing scores near the priority-grade cutoff rather than full distributions, because small systematic shifts just below the cutoff could explain why gender rank differences are insignificant while acceptance gaps persist."],"forward_implications":["Because the systematics are visible in Stage 1 ranks, mitigating them means changing how initial scores are produced, such as reviewer training, anonymized proposal text, or revised scoring rubrics, rather than relying on panel discussion.","Since the panel discussion does not redistribute ranks by experience or region, the final observing queue inherits the Stage 1 gaps, so an intervention that leaves Stage 1 untouched will not fix acceptance equity.","The persistent female acceptance deficit after demographic adjustment implies that removing the average gender difference in rank may still leave a gap if women's proposals are concentrated just below acceptance thresholds.","The marginal Stage 2 improvement for East Asian proposals suggests face-to-face discussion can partially offset one regional gap, but the effect is small and inconsistent across cycles."],"supporting_citations":[{"why":"Establishes the precedent of a gender acceptance gap at HST and motivates the search for systematics in ALMA reviews.","marker":"Reid (2014)"},{"why":"Provides the ESO seniority-based demographic interpretation that the paper uses to frame the residual gender acceptance gap.","marker":"Patat (2016)"},{"why":"Supplies the prior ALMA Cycles 2-4 gender-ranking analysis and the gender database that this study extends.","marker":"Lonsdale et al. (2016)"},{"why":"Defines the k-sample Anderson-Darling test used for every cumulative-distribution comparison in the paper.","marker":"Scholz et al. (1987)"},{"why":"Presents the reviewer acceptance-rate claim that the appendix tests using both submitted and accepted proposals.","marker":"Greaves (2018)"},{"why":"Reports the outcome of double-anonymous review at HST, providing context for ALMA's proposed mitigation steps.","marker":"Strolger & Natarajan (2019)"}],"fun_headline_variants":["ALMA review bias begins with initial scores, panel doesn't fix","Early ALMA scores favor repeat, N. American, male PIs","ALMA rankings skewed before panel, not after","Gender, region gaps in ALMA proposals predate discussion","ALMA review gaps trace to Stage 1 scores, not panel talks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the number of cycles in which someone has submitted an ALMA proposal measures PI experience; if persistence after past rejection is what actually distinguishes repeat submitters, the reported experience gradient could be selection rather than experience or reviewer bias.","fun_headline_variants_meta":{"raw":{"variants":["ALMA review bias begins with initial scores, panel doesn't fix","Early ALMA scores favor repeat, N. American, male PIs","ALMA rankings skewed before panel, not after","Gender, region gaps in ALMA proposals predate discussion","ALMA review gaps trace to Stage 1 scores, not panel talks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000192,"raw_usage":{"total_tokens":1422,"prompt_tokens":1099,"completion_tokens":323,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":715,"completion_tokens_details":{"reasoning_tokens":236}},"tokens_in":715,"tokens_out":323,"duration_ms":3426,"temperature":1.0,"reasoning_tokens":236,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:47:45.880595+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the Cycles 0-6 analysis with a seniority measure independent of ALMA submission history, such as years since PhD or publication record, and see whether the Stage 1 experience gradient persists; if it collapses, the experience effect is an artifact of who keeps submitting, while if it survives, the effect is tied to experience itself or to reviewer responses to known PIs.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the precedent of a gender acceptance gap at HST and motivates the search for systematics in ALMA reviews."},{"cited_title":"2016, Messenger, 165, 2","cited_arxiv_id":null,"evidence_quote":"Provides the ESO seniority-based demographic interpretation that the paper uses to frame the residual gender acceptance gap."},{"cited_title":"W & Stephens, M","cited_arxiv_id":null,"evidence_quote":"Defines the k-sample Anderson-Darling test used for every cumulative-distribution comparison in the paper."},{"cited_title":"2018, RNAAS, 2, 203","cited_arxiv_id":null,"evidence_quote":"Presents the reviewer acceptance-rate claim that the appendix tests using both submitted and accepted proposals."},{"cited_title":"2019, Physics Today, Volume 72, issue 3","cited_arxiv_id":null,"evidence_quote":"Reports the outcome of double-anonymous review at HST, providing context for ALMA's proposed mitigation steps."}],"review_version":1}