Pith. sign in

REVIEW 5 major objections 5 minor 15 references

Retrospective score tests versus prospective score tests for genetic association with case-control data

T0 review · 5 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Case-control gene tests should use retrospective scores, not prospective ones, when testing random genetic effects.

desk verdict The retrospective-vs-prospective intercept-shift result is a real contribution, but the simulation section has fixable errors that make the practical test non-reproducible as printed. read the letter →

arxiv 2507.19893 v1 pith:IFY2CVIV submitted 2025-07-26 stat.ME

classification stat.ME MSC 62F0362J1262P10
keywords case-controlstudyretrospectivelikelihoodprospectivescoretestrandomeffectmodelgeneticassociationdiseaseprevalencerarevariant
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Case-control studies fix the number of cases and controls, so the case fraction is a design choice, not a disease frequency. This paper argues that score tests built from the prospective logistic likelihood—the standard approach behind tests such as SKAT, SKAT-O, and MiST—ignore this distinction and can lose substantial power when testing for random genetic effects. Using the retrospective likelihood with disease prevalence plugged in yields a score statistic that is locally most powerful for the random-effect parameter, and the only difference from the prospective score is which intercept value enters the weight function. Simulations show large power gains when true prevalence differs from the case fraction, and a combined fixed-plus-random-effect test built on retrospective scores also finds a different pancreatic-cancer gene than existing tests in the real data analysis.

What carries the argument

The load-bearing object is the retrospective score statistic for the random effect: $$U_1(\alpha_*) = \sum_{i=1}^n \{D_i - \pi(\hat\$\alpha$ + \hat\$\beta$^\top x_i + \hat\gamma^\top y_i \mathbf{1})\}\{1 - 2\pi(\alpha_* + \hat\$\beta$^\top x_i + \hat\gamma^\top y_i \mathbf{1})\} y_i^\top y_i,$$ where $\hat\alpha, \hat\beta, \hat\gamma$ are obtained under the null of no genetic effect from the empirical-likelihood maximization of Qin and Zhang (1997), and $\alpha_*$ is an intercept evaluated at a prevalence guess. The prospective score is $U_1(\hat\alpha)$. The variance of $U_1(\alpha_*)$ depends on $\alpha_*$ through the weight $1 - 2\pi(\alpha_* + \cdots)$, which is why the choice of prevalence changes power. The combined test for overall genetic effect is $SS(\alpha_*) = \{\max(U_{1s}(\alpha_*), 0)\}^2 + U_{2s}^2$, where $U_{2s}$ is the standardized fixed-effect score; with unknown prevalence, SS-MAX takes the maximum over the grid $\alpha_*^{(i)} = \hat\alpha - \log(n_1/n_0) + (i-1)(b_2-b_1)/(m-1)$ with $m=4$ and $[b_1,b_2] = [-10,-0.5]$. Theorem 1 supplies the joint asymptotic normality of the standardized scores and consistent variance estimators.

What would settle it

Run the paper's simulation design with a true prevalence below $4.54\times 10^{-5}$ or above $0.38$, so the logit prevalence falls outside $[-10,-0.5]$, and compare SS-MAX with the oracle $SS(\alpha_p)$ and $SS(\hat\alpha)$; if SS-MAX tracks the oracle the grid implementation is robust, while if SS-MAX collapses toward $SS(\hat\alpha)$ the practical recommendation fails for the low-prevalence case-control studies the method is meant to serve.

Watch

Extended reading notes

Core claim

The paper's central claim is that the sampling design matters for score testing in genetic association studies. In case-control data, cases and controls are fixed by design, so the density of covariates conditional on disease status is the correct retrospective likelihood, while the prospective logistic likelihood is only a working device. Under an independent random effect model for variant effects, the retrospective score statistic for the random-effect parameter $\theta$ is $U_1(\alpha_p)$, where $\alpha_p$ is the intercept determined by the true disease prevalence $p$. The prospective score statistic used in tests such as SKAT and MiST is exactly the same formula evaluated at $\hat\alpha$, the intercept estimated from the observed case fraction. Because $\hat\alpha = \alpha_p + \log\{(1-p)n_1/(p n_0)\}$ differs from $\alpha_p$ unless $n_1/n = p$, the two statistics have different asymptotic variances, and simulations show the prospective version can lose substantial power. The paper concludes that disease prevalence, although uninformative for estimating odds ratios in logistic regression, is informative for testing genetic association, and proposes the retrospective-score tests RS, SS, and their max-over-grid versions.

Load-bearing premise

The practical power gains of the proposed MAX tests rest on the user-supplied interval for the disease-prevalence intercept being wide enough to contain the true value; the simulations demonstrate the oracle gains only for true prevalence values inside the chosen grid, and no sensitivity results are shown for prevalence outside that interval.

Editorial extensions

If this is right

  • Prospective-likelihood tests such as SKAT, SKAT-O, and MiST can be substantially underpowered for case-control studies of genes with random effects, particularly when the case fraction is far from the disease prevalence.
  • Disease prevalence, routinely available from registries, becomes a usable input for genetic association testing, not just for estimating odds ratios.
  • When prevalence is unknown, the max-over-grid version SS-MAX approximates the oracle test that knows the true prevalence and remains more powerful than existing tests under the paper's simulation settings.
  • The retrospective score test for random effects controls type I error under the independent random-effect model, whereas the MiST-Random test can have severely inflated type I errors when that model is violated.
  • In the combined PanScan I/II pancreatic-cancer analysis, SS-MAX detects the gene IRX2 with p-value $7\times 10^{-6}$, while some existing tests do not reach Bonferroni significance on that gene.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The practical success of the grid-based tests shifts the burden onto the width of the prevalence interval; a sensitivity analysis for prevalence outside the logit interval $[-10,-0.5]$ is needed before recommending SS-MAX as a default, since the oracle gains are shown only for true $\alpha_p$ values inside that grid.
  • The same retrospective-versus-prospective distinction should apply to other outcome-dependent sampling designs, such as choice-based samples in econometrics, where the intercept correction can alter the power of score tests even if coefficient estimation is unaffected.
  • The real-data qq-plots suggest that existing tests such as Burden, SKAT, SKAT-O, and MiST may have miscalibrated p-value distributions in gene-based scans, not merely lower power; this implicit claim could be checked with negative-control outcomes.
  • A testable extension is that the power advantage of $SS(\alpha_p)$ over $SS(\hat\alpha)$ should grow as $|\logit(p) - \logit(n_1/n)|$ grows, so simulations at very low prevalence, say below $10^{-4}$, would reveal where the proposed advantage saturates or reverses.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. This paper studies score tests for genetic association in case-control studies under an independent random-effect logistic model. The authors derive retrospective-likelihood score statistics for the random effect (U1(αp)) and fixed effect (U2), show that the prospective score for the random effect equals U1(α) with a different intercept, and propose combined tests SS(α*) and RS(α*) as well as adaptive maximum versions over a grid of prevalence-based α* values. Simulations and a pancreatic cancer GWAS application are used to argue that the retrospective tests, particularly with the true prevalence, outperform standard prospective tests such as SKAT, SKAT-O, MiST and burden under the assumed model.

Significance. If the main claims hold, the paper would provide a new theoretical justification for using retrospective likelihoods and disease-prevalence information in variance-component tests, where current practice relies on prospective likelihoods. The manuscript's score derivations in Section 2 are transparent and the proposed tests are benchmarked against external competitors. However, validation is currently incomplete due to several specific errors and gaps (Theorem 1 proof deferred, Table 2 panel swap, grid formula typo, variance estimator consistency), so the significance cannot be fully assessed from the submitted version.

major comments (5)
  1. [Section 3.1, Theorem 1] Theorem 1, which underpins all null distributions used in Section 4, is stated with its proof deferred to a supplementary file that is not included in the arXiv submission. In particular, the claimed asymptotic independence between U1 and U2 after nuisance-parameter estimation is critical for the SS/RS combination formulas; without the proof, the validity of the proposed tests cannot be verified. Please provide the proof or a detailed derivation in the revision.
  2. [Section 4] The grid formula for α∗i is printed as α∗i = ˆα − log(n1/n0) + (i−1)(b2−b1)/(m−1). This omits the lower endpoint b1 and therefore does not cover the intended interval [ˆα − log(n1/n0) + b1, ˆα − log(n1/n0) + b2]. With the formula as written, the grid in Example 1 is approximately {0.39, 3.55, 6.71, 9.89}, none of which is near the true αp = −1, so the reported SS-MAX powers (e.g., 38.85% at k=2 in C2) could not have been produced. Please correct the formula (likely adding + b1) and re-run the affected simulations or confirm that the implementation used the corrected version.
  3. [Table 2] The panels labeled C1 and C2 in Table 2 appear to be swapped. For C1, which has θ=0 and γ=−0.02k×1, the k=0 row is the null case and should show rejection rates near 5%; instead it shows about 92% for the SS tests. For C2, which has √θ=0.5 and γ=−0.02k×1, the k=0 row is an alternative and the random-effect component alone gives about 60–70% power in Table 1, yet the printed C2 row at k=0 is about 4%. The verbal description in Section 5.2 matches the reversed panels, so the table entries are mislabeled. Please correct the table and re-check the corresponding simulation statements.
  4. [Section 3.2] The claimed consistency of the variance estimator \hatσ11 is not supported. The estimator in Section 3.2 averages π(\tildeξ^T z)(1−π(\tildeξ^T z)) \tilde C1(x,y,α1)\tilde C1(x,y,α2), whereas the variance σ11 for retrospective data in Section 3.1 is defined with weights π(ξ0^T z)^2 under E0 and (1−π(ξ0^T z))^2 under E1. Under retrospective sampling, the empirical average of the former converges to a different mixture unless n1/n = p. The sentence 'One can straightforwardly verify' is insufficient; a proof or a corrected estimator is needed, because an inconsistent variance estimator would invalidate the null distributions of U1s and all tests based on it.
  5. [Section 5] The adaptive tests SS-MAX and RS-MAX are the practical implementations of the proposed idea, but all simulated scenarios use prevalence values inside the interval [−10, −0.5] (αp = −1 and αp = −2). No sensitivity analysis is provided for interval guesses that are wrong or for αp outside this range, and the text acknowledges that αp is not estimable from case-control data. Since the practical power gain is the central claim, the authors should report the performance of SS-MAX/RS-MAX when the interval misses the true logit(p) (e.g., rare disease with αp < −10 or common disease with αp > −0.5), or state the applicability limitation clearly.
minor comments (5)
  1. [Throughout] There are numerous typographical issues, including 'f or', 'prosp ective', 'efficiency', 'SKA T' (should be SKAT), 'samller', and 'principle components' (should be principal).
  2. [Section 5.4] The sentence 'the three SS tests ... have almost the same type I error and powers when testing the random effect of rare variants' is confusing because the preceding results refer to overall genetic effect; clarify whether RS or SS tests are meant.
  3. [Table 7] The table caption mentions 'SSMAX, SS(α^), burden, SKA T, SKA T-O and MiST' but the table also includes RS-MAX, RS(α^), and MiST-Random; align the caption with the columns.
  4. [Section 2.1] The notation ξ0 is used for both (α,β) and (αp,β) without a distinguishing superscript; this may cause confusion in Section 3.1.
  5. [Abstract] The abstract claims the proposed test is 'locally most powerful' for aggregation, but only RS(αp) is LMP for the random effect; the combined SS test is not LMP for the joint alternative. Suggest rewording.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the retrospective/prospective score comparison is a direct likelihood calculation and the proposed tests are benchmarked externally; the grid-formula and prevalence-interval issues are reproducibility and robustness concerns, not circularity.

full rationale

The paper's central result is a direct algebraic comparison of two score statistics derived from explicitly stated likelihoods. The retrospective score U1(alpha_p) in Section 2.2 and the prospective score U1(tilde alpha) in Section 2.4 are both calculated from the same underlying model, and Section 2.4 shows by substitution that they coincide except for the alpha argument; this is a derived identity, not an equivalence imposed by definition. The claimed power advantage is supported by simulations that compare the proposed tests with external benchmarks such as SKAT, SKAT-O, MiST, and burden tests, and the oracle SS(alpha_p) uses the true intercept only as an ideal simulation benchmark rather than as a fitted parameter renamed as a prediction. The adaptive SS-MAX and RS-MAX procedures use a user-supplied interval for the log-odds prevalence, not a value fitted to the simulated outcomes, so no fitted input is relabeled as a prediction. The citation of Qin and Zhang (1997) is to a published, externally checkable empirical-likelihood estimator; although one author is Jing Qin, the present derivation does not reduce to that citation, and no self-citation is used to forbid alternatives or justify an ansatz. The manuscript does contain a reproducibility issue in Section 4: the printed grid alpha*_i = hat alpha - log(n1/n0) + (i-1)*(b2-b1)/(m-1) omits b1, and no sensitivity analysis is given for prevalence outside [-10,-0.5]; however, these are correctness and robustness concerns, not circularity.

Assumptions & free parameters 1 free parameters · 5 assumptions · 0 invented entities

The central claim relies on the density-ratio form of logistic regression, independent random effects with a known variance scale, and the availability or bracketing of disease prevalence. The grid size and prevalence interval are user choices rather than fitted values. No new physical or statistical entities are invented beyond the conventional random effect.

free parameters (1)
  • Prevalence grid for SS-MAX and RS-MAX = b1=-10, b2=-0.5, m=4
    This user input determines the alpha* values used in the adaptive tests. It is chosen in the simulation study rather than fitted to data, and performance outside the interval is not assessed.
assumptions (5)
  • domain assumption Logistic density-ratio model: f1(x,y)=exp(alpha_r + beta'x + y'1 gamma) f0(x,y) under H0
    Used in Sections 2.2-2.3 to estimate alpha_r and beta via Qin-Zhang empirical likelihood. If the logistic model is misspecified, the intercept shift in U1 may not behave as derived.
  • domain assumption Independent random effects v_i i.i.d. with mean 0 and variance 2
    Section 2.1 model (1). This is the conventional random-effect assumption, in contrast to the common-random-effect assumption in SKAT and MiST, and it is load-bearing for the retrospective likelihood factorization.
  • domain assumption Known or bracketed disease prevalence p
    The retrospective score U1(alpha_p) and the practical RS-MAX and SS-MAX tests require alpha_p or an interval [b1,b2] for logit(p). Since p is not estimable from case-control data, this user input is load-bearing for the claimed power gains.
  • standard math Asymptotic regularity: finite moments, n1/n constant, root-n consistent estimators
    The conditions of Theorem 1 in Section 3.1. The proof is deferred to an unavailable supplementary file.
  • standard math Qin-Zhang (1997) empirical likelihood estimators are consistent under the density-ratio model
    Used in Sections 2.2-2.3 to plug estimators into the score equations. Treated as an established external result.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Retrospective score tests versus prospective score tests for genetic association with case-control data." pith.science (2026). https://pith.science/paper/IFY2CVIV

@misc{pith2026250719893,
  author       = {Pith},
  title        = {Pith review of: Retrospective score tests versus prospective score tests for genetic association with case-control data},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IFY2CVIV}},
  note         = {Machine review of arXiv:2507.19893}
}
read the original abstract

Since the seminal work by Prentice and Pyke (1979), the prospective logistic likelihood has become the standard method of analysis for retrospectively collected case-control data, in particular for testing the association between a single genetic marker and a disease outcome in genetic case-control studies. When studying multiple genetic markers with relatively small effects, especially those with rare variants, various aggregated approaches based on the same prospective likelihood have been developed to integrate subtle association evidence among all considered markers. In this paper we show that using the score statistic derived from a prospective likelihood is not optimal in the analysis of retrospectively sampled genetic data. We develop the locally most powerful genetic aggregation test derived through the retrospective likelihood under a random effect model assumption. In contrast to the fact that the disease prevalence information cannot be used to improve the efficiency for the estimation of odds ratio parameters in logistic regression models, we show that it can be utilized to enhance the testing power in genetic association studies. Extensive simulations demonstrate the advantages of the proposed method over the existing ones. One real genome-wide association study is analyzed for illustration.

Figures

Figures reproduced from arXiv: 2507.19893 by the authors.

Figure 1
Figure 1. qq-plots of minus logarithm of p-values. [PITH_FULL_IMAGE:figures/full_fig_p024_1.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

15 extracted references · 15 canonical work pages

  1. [1]

    Amemiya, T. (1985). Advanced Econometrics. Harvard Univer sity Press. 24

  2. [2]

    Z., e t al

    Amundadottir, L., Kraft, P ., Stolzenberg-Solomon, R. Z., e t al. (2009). Genome-wide associ- ation study identifies variants in the ABO locus associated w ith susceptibility to pancreatic cancer. Nature Genetics 41, 986–990

  3. [3]

    Anderson, J. A. (1979). Multivariate logistic compounds. Biometrika 66, 17–26

  4. [4]

    R., Wheeler, H

    Gamazon, E. R., Wheeler, H. E., Shah, K.P ., Mozaffari, S. V .,Aquino-Michaels, K., Carroll, R. J., Eyler, A. E., Denny, J. C., Nicolae, D. L., Cox, N. J. and Im , H. K. (2015). A gene-based association method for mapping traits using reference tran scriptome data. Nature Genetics 47, 1091–1098

  5. [5]

    Jiang, J. (2007). Linear and Generalized Linear Mixed Model s and Their Applications. New Y ork: Springer

  6. [6]

    and Wang, Y

    Ke, C. and Wang, Y . (2001). Semiparametric nonlinear mixed- effects models and their appli- cations (with discussion). Journal of the American Statistical Association 96, 1272–1298

  7. [7]

    Lee, S., Wu, M. C. and Lin, X. (2012) Optimal tests for rare var iant effects in sequencing association studies. Biostatistics 13, 762–775. Lin D. Y . and Tang, Z. Z. (2011). A general framework for detec ting disease associations with rare variants in sequencing studies. American Journal of Human Genetics 89, 354–367

  8. [8]

    Lin, X. (1997). V ariance component testing in generalized l inear models with random effects. Biometrika 84, 309–326

Show all 15 references
  1. [9]

    M., Amundadottir, L., Fuchs, C

    Petersen, G. M., Amundadottir, L., Fuchs, C. S., et al. (2010 ). A genome-wide association study identifies pancreatic cancer susceptibility loci on c hromosomes 13q22.1, 1q32.1 and 5p15.33. Nature Genetics 42, 224–228

  2. [10]

    Prentice, R. L. and Pyke, R. (1979). Logistic disease incidence models and case-control studies. Biometrika 66, 403–411

  3. [11]

    K., Stephens, M

    Pritchard, J. K., Stephens, M. and Donnelly, P . (2000). Infe rence of population structure using multilocus genotype data. Genetics 155, 945–959

  4. [12]

    and Zhang, B

    Qin, J. and Zhang, B. (1997). A goodness-of-fit test for logis tic regression models based on case-control data. Biometrika 84, 609–618

  5. [13]

    and Hsu, L

    Sun, J., Zheng, Y . and Hsu, L. (2013). A unified mixed-effects model for rare-variant associa- tion in sequencing studies. Genetic Epidemiology 37, 334–344

  6. [14]

    Wang, Y . (1998). Mixed-effects smoothing spline analysis o f variance. Journal of the Royal Statistical Society, Series B 60, 159–174

  7. [15]

    C., Lee, S., Cai, T., Li, Y ., Boehnke, M

    Wu, M. C., Lee, S., Cai, T., Li, Y ., Boehnke, M. and Lin, X. (201 1). Rare-variant association testing for sequencing data with the sequence kernel associ ation test. American Journal of Human Genetics 89, 82–93. 25 V erbeke, G. and Lesaffre, E. (1996). A linear mixed-effects...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.