REVIEW 5 major objections 5 minor 15 references
Retrospective score tests versus prospective score tests for genetic association with case-control data
T0 review · 5 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Case-control gene tests should use retrospective scores, not prospective ones, when testing random genetic effects.
desk verdict The retrospective-vs-prospective intercept-shift result is a real contribution, but the simulation section has fixable errors that make the practical test non-reproducible as printed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the retrospective score statistic for the random effect: $$U_1(\alpha_*) = \sum_{i=1}^n \{D_i - \pi(\hat\$\alpha$ + \hat\$\beta$^\top x_i + \hat\gamma^\top y_i \mathbf{1})\}\{1 - 2\pi(\alpha_* + \hat\$\beta$^\top x_i + \hat\gamma^\top y_i \mathbf{1})\} y_i^\top y_i,$$ where $\hat\alpha, \hat\beta, \hat\gamma$ are obtained under the null of no genetic effect from the empirical-likelihood maximization of Qin and Zhang (1997), and $\alpha_*$ is an intercept evaluated at a prevalence guess. The prospective score is $U_1(\hat\alpha)$. The variance of $U_1(\alpha_*)$ depends on $\alpha_*$ through the weight $1 - 2\pi(\alpha_* + \cdots)$, which is why the choice of prevalence changes power. The combined test for overall genetic effect is $SS(\alpha_*) = \{\max(U_{1s}(\alpha_*), 0)\}^2 + U_{2s}^2$, where $U_{2s}$ is the standardized fixed-effect score; with unknown prevalence, SS-MAX takes the maximum over the grid $\alpha_*^{(i)} = \hat\alpha - \log(n_1/n_0) + (i-1)(b_2-b_1)/(m-1)$ with $m=4$ and $[b_1,b_2] = [-10,-0.5]$. Theorem 1 supplies the joint asymptotic normality of the standardized scores and consistent variance estimators.
What would settle it
Run the paper's simulation design with a true prevalence below $4.54\times 10^{-5}$ or above $0.38$, so the logit prevalence falls outside $[-10,-0.5]$, and compare SS-MAX with the oracle $SS(\alpha_p)$ and $SS(\hat\alpha)$; if SS-MAX tracks the oracle the grid implementation is robust, while if SS-MAX collapses toward $SS(\hat\alpha)$ the practical recommendation fails for the low-prevalence case-control studies the method is meant to serve.
Extended reading notes
Core claim
The paper's central claim is that the sampling design matters for score testing in genetic association studies. In case-control data, cases and controls are fixed by design, so the density of covariates conditional on disease status is the correct retrospective likelihood, while the prospective logistic likelihood is only a working device. Under an independent random effect model for variant effects, the retrospective score statistic for the random-effect parameter $\theta$ is $U_1(\alpha_p)$, where $\alpha_p$ is the intercept determined by the true disease prevalence $p$. The prospective score statistic used in tests such as SKAT and MiST is exactly the same formula evaluated at $\hat\alpha$, the intercept estimated from the observed case fraction. Because $\hat\alpha = \alpha_p + \log\{(1-p)n_1/(p n_0)\}$ differs from $\alpha_p$ unless $n_1/n = p$, the two statistics have different asymptotic variances, and simulations show the prospective version can lose substantial power. The paper concludes that disease prevalence, although uninformative for estimating odds ratios in logistic regression, is informative for testing genetic association, and proposes the retrospective-score tests RS, SS, and their max-over-grid versions.
Load-bearing premise
The practical power gains of the proposed MAX tests rest on the user-supplied interval for the disease-prevalence intercept being wide enough to contain the true value; the simulations demonstrate the oracle gains only for true prevalence values inside the chosen grid, and no sensitivity results are shown for prevalence outside that interval.
Editorial extensions
If this is right
- Prospective-likelihood tests such as SKAT, SKAT-O, and MiST can be substantially underpowered for case-control studies of genes with random effects, particularly when the case fraction is far from the disease prevalence.
- Disease prevalence, routinely available from registries, becomes a usable input for genetic association testing, not just for estimating odds ratios.
- When prevalence is unknown, the max-over-grid version SS-MAX approximates the oracle test that knows the true prevalence and remains more powerful than existing tests under the paper's simulation settings.
- The retrospective score test for random effects controls type I error under the independent random-effect model, whereas the MiST-Random test can have severely inflated type I errors when that model is violated.
- In the combined PanScan I/II pancreatic-cancer analysis, SS-MAX detects the gene IRX2 with p-value $7\times 10^{-6}$, while some existing tests do not reach Bonferroni significance on that gene.
Reading between the lines
- The practical success of the grid-based tests shifts the burden onto the width of the prevalence interval; a sensitivity analysis for prevalence outside the logit interval $[-10,-0.5]$ is needed before recommending SS-MAX as a default, since the oracle gains are shown only for true $\alpha_p$ values inside that grid.
- The same retrospective-versus-prospective distinction should apply to other outcome-dependent sampling designs, such as choice-based samples in econometrics, where the intercept correction can alter the power of score tests even if coefficient estimation is unaffected.
- The real-data qq-plots suggest that existing tests such as Burden, SKAT, SKAT-O, and MiST may have miscalibrated p-value distributions in gene-based scans, not merely lower power; this implicit claim could be checked with negative-control outcomes.
- A testable extension is that the power advantage of $SS(\alpha_p)$ over $SS(\hat\alpha)$ should grow as $|\logit(p) - \logit(n_1/n)|$ grows, so simulations at very low prevalence, say below $10^{-4}$, would reveal where the proposed advantage saturates or reverses.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies score tests for genetic association in case-control studies under an independent random-effect logistic model. The authors derive retrospective-likelihood score statistics for the random effect (U1(αp)) and fixed effect (U2), show that the prospective score for the random effect equals U1(α) with a different intercept, and propose combined tests SS(α*) and RS(α*) as well as adaptive maximum versions over a grid of prevalence-based α* values. Simulations and a pancreatic cancer GWAS application are used to argue that the retrospective tests, particularly with the true prevalence, outperform standard prospective tests such as SKAT, SKAT-O, MiST and burden under the assumed model.
Significance. If the main claims hold, the paper would provide a new theoretical justification for using retrospective likelihoods and disease-prevalence information in variance-component tests, where current practice relies on prospective likelihoods. The manuscript's score derivations in Section 2 are transparent and the proposed tests are benchmarked against external competitors. However, validation is currently incomplete due to several specific errors and gaps (Theorem 1 proof deferred, Table 2 panel swap, grid formula typo, variance estimator consistency), so the significance cannot be fully assessed from the submitted version.
major comments (5)
- [Section 3.1, Theorem 1] Theorem 1, which underpins all null distributions used in Section 4, is stated with its proof deferred to a supplementary file that is not included in the arXiv submission. In particular, the claimed asymptotic independence between U1 and U2 after nuisance-parameter estimation is critical for the SS/RS combination formulas; without the proof, the validity of the proposed tests cannot be verified. Please provide the proof or a detailed derivation in the revision.
- [Section 4] The grid formula for α∗i is printed as α∗i = ˆα − log(n1/n0) + (i−1)(b2−b1)/(m−1). This omits the lower endpoint b1 and therefore does not cover the intended interval [ˆα − log(n1/n0) + b1, ˆα − log(n1/n0) + b2]. With the formula as written, the grid in Example 1 is approximately {0.39, 3.55, 6.71, 9.89}, none of which is near the true αp = −1, so the reported SS-MAX powers (e.g., 38.85% at k=2 in C2) could not have been produced. Please correct the formula (likely adding + b1) and re-run the affected simulations or confirm that the implementation used the corrected version.
- [Table 2] The panels labeled C1 and C2 in Table 2 appear to be swapped. For C1, which has θ=0 and γ=−0.02k×1, the k=0 row is the null case and should show rejection rates near 5%; instead it shows about 92% for the SS tests. For C2, which has √θ=0.5 and γ=−0.02k×1, the k=0 row is an alternative and the random-effect component alone gives about 60–70% power in Table 1, yet the printed C2 row at k=0 is about 4%. The verbal description in Section 5.2 matches the reversed panels, so the table entries are mislabeled. Please correct the table and re-check the corresponding simulation statements.
- [Section 3.2] The claimed consistency of the variance estimator \hatσ11 is not supported. The estimator in Section 3.2 averages π(\tildeξ^T z)(1−π(\tildeξ^T z)) \tilde C1(x,y,α1)\tilde C1(x,y,α2), whereas the variance σ11 for retrospective data in Section 3.1 is defined with weights π(ξ0^T z)^2 under E0 and (1−π(ξ0^T z))^2 under E1. Under retrospective sampling, the empirical average of the former converges to a different mixture unless n1/n = p. The sentence 'One can straightforwardly verify' is insufficient; a proof or a corrected estimator is needed, because an inconsistent variance estimator would invalidate the null distributions of U1s and all tests based on it.
- [Section 5] The adaptive tests SS-MAX and RS-MAX are the practical implementations of the proposed idea, but all simulated scenarios use prevalence values inside the interval [−10, −0.5] (αp = −1 and αp = −2). No sensitivity analysis is provided for interval guesses that are wrong or for αp outside this range, and the text acknowledges that αp is not estimable from case-control data. Since the practical power gain is the central claim, the authors should report the performance of SS-MAX/RS-MAX when the interval misses the true logit(p) (e.g., rare disease with αp < −10 or common disease with αp > −0.5), or state the applicability limitation clearly.
minor comments (5)
- [Throughout] There are numerous typographical issues, including 'f or', 'prosp ective', 'efficiency', 'SKA T' (should be SKAT), 'samller', and 'principle components' (should be principal).
- [Section 5.4] The sentence 'the three SS tests ... have almost the same type I error and powers when testing the random effect of rare variants' is confusing because the preceding results refer to overall genetic effect; clarify whether RS or SS tests are meant.
- [Table 7] The table caption mentions 'SSMAX, SS(α^), burden, SKA T, SKA T-O and MiST' but the table also includes RS-MAX, RS(α^), and MiST-Random; align the caption with the columns.
- [Section 2.1] The notation ξ0 is used for both (α,β) and (αp,β) without a distinguishing superscript; this may cause confusion in Section 3.1.
- [Abstract] The abstract claims the proposed test is 'locally most powerful' for aggregation, but only RS(αp) is LMP for the random effect; the combined SS test is not LMP for the joint alternative. Suggest rewording.
Circularity Check
No significant circularity: the retrospective/prospective score comparison is a direct likelihood calculation and the proposed tests are benchmarked externally; the grid-formula and prevalence-interval issues are reproducibility and robustness concerns, not circularity.
full rationale
The paper's central result is a direct algebraic comparison of two score statistics derived from explicitly stated likelihoods. The retrospective score U1(alpha_p) in Section 2.2 and the prospective score U1(tilde alpha) in Section 2.4 are both calculated from the same underlying model, and Section 2.4 shows by substitution that they coincide except for the alpha argument; this is a derived identity, not an equivalence imposed by definition. The claimed power advantage is supported by simulations that compare the proposed tests with external benchmarks such as SKAT, SKAT-O, MiST, and burden tests, and the oracle SS(alpha_p) uses the true intercept only as an ideal simulation benchmark rather than as a fitted parameter renamed as a prediction. The adaptive SS-MAX and RS-MAX procedures use a user-supplied interval for the log-odds prevalence, not a value fitted to the simulated outcomes, so no fitted input is relabeled as a prediction. The citation of Qin and Zhang (1997) is to a published, externally checkable empirical-likelihood estimator; although one author is Jing Qin, the present derivation does not reduce to that citation, and no self-citation is used to forbid alternatives or justify an ansatz. The manuscript does contain a reproducibility issue in Section 4: the printed grid alpha*_i = hat alpha - log(n1/n0) + (i-1)*(b2-b1)/(m-1) omits b1, and no sensitivity analysis is given for prevalence outside [-10,-0.5]; however, these are correctness and robustness concerns, not circularity.
Assumptions & free parameters
free parameters (1)
- Prevalence grid for SS-MAX and RS-MAX =
b1=-10, b2=-0.5, m=4
assumptions (5)
- domain assumption Logistic density-ratio model: f1(x,y)=exp(alpha_r + beta'x + y'1 gamma) f0(x,y) under H0
- domain assumption Independent random effects v_i i.i.d. with mean 0 and variance 2
- domain assumption Known or bracketed disease prevalence p
- standard math Asymptotic regularity: finite moments, n1/n constant, root-n consistent estimators
- standard math Qin-Zhang (1997) empirical likelihood estimators are consistent under the density-ratio model
Cite this review
Pith. "Pith review of Retrospective score tests versus prospective score tests for genetic association with case-control data." pith.science (2026). https://pith.science/paper/IFY2CVIV
@misc{pith2026250719893,
author = {Pith},
title = {Pith review of: Retrospective score tests versus prospective score tests for genetic association with case-control data},
year = {2026},
howpublished = {\url{https://pith.science/paper/IFY2CVIV}},
note = {Machine review of arXiv:2507.19893}
}
read the original abstract
Since the seminal work by Prentice and Pyke (1979), the prospective logistic likelihood has become the standard method of analysis for retrospectively collected case-control data, in particular for testing the association between a single genetic marker and a disease outcome in genetic case-control studies. When studying multiple genetic markers with relatively small effects, especially those with rare variants, various aggregated approaches based on the same prospective likelihood have been developed to integrate subtle association evidence among all considered markers. In this paper we show that using the score statistic derived from a prospective likelihood is not optimal in the analysis of retrospectively sampled genetic data. We develop the locally most powerful genetic aggregation test derived through the retrospective likelihood under a random effect model assumption. In contrast to the fact that the disease prevalence information cannot be used to improve the efficiency for the estimation of odds ratio parameters in logistic regression models, we show that it can be utilized to enhance the testing power in genetic association studies. Extensive simulations demonstrate the advantages of the proposed method over the existing ones. One real genome-wide association study is analyzed for illustration.
Figures
Reference graph
Works this paper leans on
-
[1]
Amemiya, T. (1985). Advanced Econometrics. Harvard Univer sity Press. 24
work page 1985
-
[2]
Amundadottir, L., Kraft, P ., Stolzenberg-Solomon, R. Z., e t al. (2009). Genome-wide associ- ation study identifies variants in the ABO locus associated w ith susceptibility to pancreatic cancer. Nature Genetics 41, 986–990
work page 2009
-
[3]
Anderson, J. A. (1979). Multivariate logistic compounds. Biometrika 66, 17–26
work page 1979
-
[4]
Gamazon, E. R., Wheeler, H. E., Shah, K.P ., Mozaffari, S. V .,Aquino-Michaels, K., Carroll, R. J., Eyler, A. E., Denny, J. C., Nicolae, D. L., Cox, N. J. and Im , H. K. (2015). A gene-based association method for mapping traits using reference tran scriptome data. Nature Genetics 47, 1091–1098
work page 2015
-
[5]
Jiang, J. (2007). Linear and Generalized Linear Mixed Model s and Their Applications. New Y ork: Springer
work page 2007
-
[6]
Ke, C. and Wang, Y . (2001). Semiparametric nonlinear mixed- effects models and their appli- cations (with discussion). Journal of the American Statistical Association 96, 1272–1298
work page 2001
-
[7]
Lee, S., Wu, M. C. and Lin, X. (2012) Optimal tests for rare var iant effects in sequencing association studies. Biostatistics 13, 762–775. Lin D. Y . and Tang, Z. Z. (2011). A general framework for detec ting disease associations with rare variants in sequencing studies. American Journal of Human Genetics 89, 354–367
work page 2012
-
[8]
Lin, X. (1997). V ariance component testing in generalized l inear models with random effects. Biometrika 84, 309–326
work page 1997
Show all 15 references
-
[9]
M., Amundadottir, L., Fuchs, C
Petersen, G. M., Amundadottir, L., Fuchs, C. S., et al. (2010 ). A genome-wide association study identifies pancreatic cancer susceptibility loci on c hromosomes 13q22.1, 1q32.1 and 5p15.33. Nature Genetics 42, 224–228
2010
-
[10]
Prentice, R. L. and Pyke, R. (1979). Logistic disease incidence models and case-control studies. Biometrika 66, 403–411
1979
-
[11]
K., Stephens, M
Pritchard, J. K., Stephens, M. and Donnelly, P . (2000). Infe rence of population structure using multilocus genotype data. Genetics 155, 945–959
2000
-
[12]
and Zhang, B
Qin, J. and Zhang, B. (1997). A goodness-of-fit test for logis tic regression models based on case-control data. Biometrika 84, 609–618
1997
-
[13]
and Hsu, L
Sun, J., Zheng, Y . and Hsu, L. (2013). A unified mixed-effects model for rare-variant associa- tion in sequencing studies. Genetic Epidemiology 37, 334–344
2013
-
[14]
Wang, Y . (1998). Mixed-effects smoothing spline analysis o f variance. Journal of the Royal Statistical Society, Series B 60, 159–174
1998
-
[15]
C., Lee, S., Cai, T., Li, Y ., Boehnke, M
Wu, M. C., Lee, S., Cai, T., Li, Y ., Boehnke, M. and Lin, X. (201 1). Rare-variant association testing for sequencing data with the sequence kernel associ ation test. American Journal of Human Genetics 89, 82–93. 25 V erbeke, G. and Lesaffre, E. (1996). A linear mixed-effects...
1996
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.