Pith. sign in

REVIEW 3 major objections 4 minor 31 references

Effect of Interim Adaptations in Group Sequential Designs

T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Interim stopping changes maximum likelihood estimates into mixtures of truncated distributions, and conditioning on the stopping stage loses exactly the Fisher information carried by the stopping rule.

desk verdict Useful message about non-normal MLE mixtures and information loss, but the asymptotic section has a fixable shift error; the UMP claim survives its flawed proof. read the letter →

arxiv 1908.01411 v1 pith:32RFK6CO submitted 2019-08-04 stat.ME

classification stat.ME MSC 62L0562L1062F12
keywords groupsequentialdesignsadaptivemaximumlikelihoodestimationasymptoticdistributiontheoryinterimanalyseslocalalternativehypothesesinformationlosstruncateddistributions
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Group sequential designs let a trial stop early, and this paper argues that the stopping rule is part of the statistical model: it cuts the sample space down, so maximum likelihood estimators no longer have their usual distributions. The paper shows that the MLE is a mixture of truncated distributions, one component for each stage at which the trial can stop, and that under local alternatives the asymptotic distribution is a non-normal mixture of truncated normals. It also shows that conditioning on the observed stopping stage discards exactly the Fisher information contained in the stopping rule, so standard information fractions computed from fixed-sample information overstate what an interim look provides. As a remedy, it proposes group sequential designs in which stage-specific sample sizes and critical values are chosen to guarantee a preset power against a sequence of ordered alternatives, avoiding the need to compute adapted information fractions. The practical point is that trials that can stop early need estimators, confidence statements, and design calculations reflecting the adaptation rule actually used.

What carries the argument

The central object is the sub-density form of the likelihood under a group sequential design, $f_{X_{(D)}}(x_{(D)}|\theta)=\prod_{k=1}^K \left[f_{X_{(k)}}(x_{(k)}|\theta)\right]^{I(D=k)}$, in which each stopping event $D=k$ contributes a density supported only on the region of the sample space that the stopping rule leaves reachable. Conditioning on stopping divides that sub-density by $\Pr_\theta(D=k)$, which inserts the term $\partial^2\log\Pr_\theta(D=k)/\partial\theta^2$ into the observed information and is the source of the information-loss identity. The proposed design machinery is a recursive system that fixes stage-specific type I error and powers each stage against a different ordered alternative. The asymptotic machinery is the limiting random variable $r(D)=\sum_{k=1}^K I(D=k)\lim n_{(k)}/n_k$, which determines how stagewise normal components combine into the mixture of truncated normals.

What would settle it

Simulate the two-stage Pocock design with $X_i\sim N(h/\sqrt{n_1},1)$, $n_1=n_2=100$, and $c_1=2.18$ for several $h$, and compare the empirical CDF of $V(D)$ with equation (19); if the mixture-of-truncated-normals limit is not within Monte Carlo error, the paper's asymptotic claim fails. For the most-powerful claim, evaluate the conditional likelihood ratio $LR^c(t)$ at a fixed $k$: any non-monotonicity in $t$ disproves the theorem.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes that when the decision to stop at stage $D$ depends on the parameter of interest, the maximum likelihood estimator is $\hat\theta(D)=\sum_{k=1}^K I(D=k)\hat\theta_{(k)}$, a mixture whose components are distributions truncated to the region of support left reachable by the stopping rule. Under local alternatives $\theta=h/\sqrt{n_1}$, the standardized estimator $V(D)$ converges to a mixture of truncated normal distributions: equation (19) for two stages and equation (22) for $K$ stages, with the mixing weights driven by the limiting ratios $r_{(k)}=\lim n_{(k)}/n_k$; under fixed alternatives the test statistic degenerates to a point mass. The information result is the identity $I-I^c=E_D[-\partial^2\log\Pr_\theta(D)/\partial\theta^2]$, the Fisher information about $\theta$ in the stopping rule itself. The design contribution is a group sequential scheme in which a sequence of ordered alternative hypotheses $\theta_1>\theta_2>\cdots>\theta_K>0$, each powered at $1-\beta$, determines stage-specific sample sizes and critical values, so adapted information fractions never have to be computed.

Load-bearing premise

The claim that each stage's test is most powerful assumes that conditioning on the trial having reached that stage does not change the relative likelihood of different parameter values; because the probability of stopping at each stage depends on the parameter itself, that assumption can fail, and with it the optimality claim.

Editorial extensions

If this is right

  • Trialists should not quote the usual fixed-sample information fraction $n_{(k)}/n_{(K)}$ when planning interim looks; the adapted fraction $I_{(k)}/I_{(K)}$ is the relevant quantity and is generally smaller.
  • Post-hoc analyses that condition on the observed stopping stage pay a measurable precision penalty equal to the Fisher information about the treatment effect in the stopping rule, and near a stopping boundary the conditional MLE can be badly biased.
  • Under local alternatives, normal approximations for interim test statistics give way to mixtures of truncated normals, so repeated confidence intervals built on normality can be miscalibrated.
  • A design powered separately for each ordered alternative gives each interim look a clinically meaningful power guarantee, and its operating characteristics can be read from stagewise rejection probabilities and expected sample size without computing information fractions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A simulation check the paper does not run is whether equation (19) still tracks the empirical distribution of $V(D)$ when stagewise estimators come from binary or time-to-event endpoints; the regularity setup suggests it should, and this would extend the result beyond normal and survival-model examples.
  • The identity for $I-I^c$ supplies a practical diagnostic: estimate the stopping probabilities under a working model and compare conditional with unconditional information; designs whose stopping probabilities depend steeply on the parameter will show the largest penalties.
  • Treating unreached patient data as structural zeros rather than missing data is a convention the paper chooses explicitly, so likelihood-based model selection and multiple-imputation approaches would systematically disagree with its information calculations; that disagreement is a conceptual choice, not a numerical error.
  • The ordered-alternative design can be generalized to unequal stagewise error spending, trading early-stage power for smaller expected sample size while keeping the no-information-fraction feature.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The manuscript studies maximum likelihood estimation and information measures in group sequential designs (GSDs) whose stopping rules depend on the parameter of interest. It argues that early stopping changes the support of the data distribution, so MLEs have distributions that are mixtures of truncated distributions; that unconditioned MLEs carry more Fisher information than MLEs conditioned on the stopping stage; that the usual 'information fraction' based on fixed-sample information is not generally correct; and that a design can instead be specified directly by ordered alternative hypotheses and desired stage-wise power. Section 3 proposes such a design and claims, via a monotone likelihood ratio argument, that the resulting tests are most powerful for a fixed alpha-spending function. Section 4 gives local asymptotic distributional formulas for the standardized MLE under two-stage and K-stage GSDs, claiming non-normal mixtures of truncated normals. A Cox regression example illustrates the design. The paper's central qualitative claims are plausible, but the local asymptotic formulas in Section 4 omit the noncentrality parameter and the corresponding truncation shifts, so the quantitative content of that section is not established as written.

Significance. If corrected, the paper would make a useful contribution. Its finite-sample representation of GSD MLEs as mixtures of truncated distributions is clearly illustrated by simulation, and the information decomposition in Eq. (10) — that conditioning on stopping loses exactly the Fisher information in the stopping rule — is clean and potentially valuable for practitioners. The proposed design driven by ordered alternatives is a practical alternative to information-fraction-based designs, and the Cox example demonstrates its applicability. The paper also correctly emphasizes that the canonical multivariate normality assumption fails after early stopping. The main reservation is the local asymptotic theory: as printed, the formulas in Section 4 are the null-limits (h=0) and do not describe the claimed local alternative limits. Because those formulas underlie the paper's principal asymptotic claims, the manuscript needs substantive revision before the results can be relied upon.

major comments (3)
  1. [Section 4.1, Eq. (19)] Equation (19) omits the local noncentrality parameter. Under theta = h/sqrt(n1), the stage-1 statistic satisfies Z1 ~ N(h,1), so the event D=2 is {Z1 <= c1} = {xi1 <= c1 - h} and the conditional density of Z1 given D=2 is phi(xi1)/Phi(c1-h) with xi1 = Z1 - h, not phi(y)/Pr(y<=c1). The limiting stage-1 stopping probability is p1 = 1 - Phi(c1 - h), not the null value alpha/2. The integral therefore needs a shifted argument and a shifted upper truncation limit; as printed, Eq. (19) gives the h=0 limit. The quantitative local asymptotic claim of Section 4 is not established.
  2. [Section 4.3, Eq. (22)] The K-stage formula (22) has the same defect as Eq. (19). The event r(D)=r(k) corresponds to D=k, which imposes joint constraints on the stage-wise statistics, e.g. xi1 <= c1 - h, xi2 <= c2 - h, and so on, and the stage-wise statistics have non-zero means under local alternatives. None of these shifts or truncation constraints appear in (22), so the displayed expression is not the limiting distribution under theta = h/sqrt(n1). The qualitative claim that the limit is a non-normal mixture of truncated normals may survive a correction, but the formulas as given are incorrect.
  3. [Section 2.2, Eq. (11)] The inequality I(k) <= Ifix(k) at theta = hattheta does not follow from the sign of the second derivative of log fX(k). Comparing (11) with (9) involves different normalizing constants and truncated supports; the integrand's sign alone cannot establish the inequality. In the normal example of Section 2.3, Iobs(k) = n(k) is constant, so I(k) = Ifix(k) rather than a strict inequality. The dashed and dotted curves in Figure 4 and the surrounding text need to be reconciled with this observation.
minor comments (4)
  1. [Section 3.3, Eqs. (14)-(15)] The statement that 'For every D=k, LRc(t)=LR(t)' is not correct: from the displayed formulas, LRc(t) equals LR(t) multiplied by Prtheta0(D=k)/Prtheta(D=k). Because this factor is constant in t, the monotone likelihood ratio conclusion still holds, but the displayed equality should be corrected and the proof rephrased accordingly.
  2. [Section 4.1] The notation p1 = lim_{n1->infinity} Prtheta(D=1) is ambiguous because theta changes with n1 in the local alternative setting; the limiting probability should be written explicitly as a function of h, e.g. 1 - Phi(c1 - h) in the example.
  3. [Section 1] There is a typo: 'the probability to reject the converges to 1' should read 'the probability to reject the null hypothesis converges to 1'.
  4. [Figure 4 caption] The caption's phrases 'thick right is for stopping at stage 1 and a thick left for stage 2' are unclear; please indicate which curve corresponds to I(1), I(2), Ic(1), and Ic(2) explicitly.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: information identities are derived from definitions, and the self-cited asymptotic theorem is used as independent stated-assumption support.

full rationale

The paper's central derivations do not reduce to their inputs by construction. The information loss identity in equation (10) follows algebraically from the definitions of I, Ic, and the conditional likelihood in (4), with no fitted parameters. The finite-sample distributions in Section 2 are derived from the stated sub-density likelihood (1), not assumed into the conclusion. The new design in Section 3 solves a system of power and error constraints for sample sizes and critical values, so operational characteristics are inputs, not outputs relabeled as predictions. The only self-citation is in Section 4, where the authors import Theorem 1 from Tarima and Flournoy (2019) for the local asymptotic behavior; this is a parameter-free theorem with stated regularity assumptions and does not smuggle the present conclusion into the citation. The UMP argument in Section 3.3 contains a slip in writing LRc(t)=LR(t) without the ratio Pr_theta0(D=k)/Pr_theta(D=k), but that ratio is constant in the sufficient statistic t for each fixed k, so the monotone likelihood-ratio conclusion does not circularly depend on the conclusion. Any concern that equation (19) omits the local shift h is a correctness or derivational issue, not an equivalence-by-construction between input and output, and therefore does not affect the circularity score.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new physical entities. Its theoretical claims use no fitted constants; the only user-specified inputs are the ordered alternative effect sizes in the example. The main unstated burdens are the conditional-MLR equality and the solvability of the design equations.

free parameters (1)
  • Ordered alternative effect sizes for the example design = theta1=0.3, theta2=0.2, theta3=0.1
    User-chosen inputs in Section 3.2.1, not fitted to data. The proposed design method requires the analyst to specify a decreasing sequence, and the resulting sample sizes and critical values depend on it.
assumptions (4)
  • domain assumption Every stage is reachable with positive probability (Prθ(D=k)>0 for all k).
    Section 2 uses this to define conditional densities and information measures.
  • domain assumption Stage-specific MLEs satisfy the asymptotic normality regularity conditions in Assumption (16).
    Section 4 imports this from Tarima and Flournoy (2019) for the local asymptotic mixture results.
  • ad hoc to paper The conditional likelihood ratio LRc(t) equals the unconditional likelihood ratio LR(t) in exponential families with early stopping.
    Section 3.3 claims this equality, but the factor Prθ0(D=k)/Prθ(D=k) does not cancel, so this assumption is load-bearing for Theorem 1.
  • ad hoc to paper The system of equations (13) has a solution with positive sample sizes and critical values.
    Section 3.2.1 solves one numerical instance; no existence or uniqueness proof is given for the proposed design.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Effect of Interim Adaptations in Group Sequential Designs." pith.science (2026). https://pith.science/paper/32RFK6CO

@misc{pith2026190801411,
  author       = {Pith},
  title        = {Pith review of: Effect of Interim Adaptations in Group Sequential Designs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/32RFK6CO}},
  note         = {Machine review of arXiv:1908.01411}
}
read the original abstract

This manuscript investigates unconditional and conditional-on-stopping maximum likelihood estimators (MLEs), information measures and information loss associated with conditioning in group sequential designs (GSDs). The possibility of early stopping brings truncation to the distributional form of MLEs; sequentially, GSD decisions eliminate some events from the sample space. Multiple testing induces mixtures on the adapted sample space. Distributions of MLEs are mixtures of truncated distributions. Test statistics that are asymptotically normal without GSD, have asymptotic distributions, under GSD, that are non-normal mixtures of truncated normal distributions under local alternatives; under fixed alternatives, asymptotic distributions of test statistics are degenerate. Estimation of various statistical quantities such as information, information fractions, and confidence intervals should account for the effect of planned adaptations. Calculation of adapted information fractions requires substantial computational effort. Therefore, a new GSD is proposed in which stage-specific sample sizes are fully determined by desired operational characteristics, and calculation of information fractions is not needed.

Figures

Figures reproduced from arXiv: 1908.01411 by the authors.

Figure 1
Figure 1. Support associated with different stopping decisions when [PITH_FULL_IMAGE:figures/full_fig_p009_1.png] view at source ↗
Figure 2
Figure 2. Histograms of MLEs for the Pocock example with critical value [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗
Figure 3
Figure 3. Log-likelihood plots as a functions of θ given selected realizations of Z1 = √ n1X¯ 1; n1 = n2 = 100, X ∼ N(θ, 1); dotted lines trace maximums. The critical value for stopping is c1 = 2.18 [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Information Measures as a Function of θ analyses. In two-stage designs, the information fraction is often estimated as I f ix (1) /I f ix (2) , which in the Pocock example is equal to 1/2 since I f ix (1) = 100 and I f ix (2) = 200. This information fraction argument s…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

31 extracted references · 31 canonical work pages

  1. [1]

    McPherson, and B

    Armitage, P., C. McPherson, and B. Rowe (1969). Repeated significance tests on accumu- lating data. Journal of the Royal Statistical Society: Series A 132 , 235–244

  2. [2]

    Henderson, H

    Asendorf, T., R. Henderson, H. Schmidli, and T. Friede (2019). Sample size re-estimation for clinical trials with longitudinal negative binomial counts including time trends: Sample size re-estimation with longitudinal negative binomial counts. Statistics in Medicine 38 (9), 1503–1528

  3. [3]

    Koenig, and P

    Brannath, W., F. Koenig, and P. Bauer (2006). Estimation in flexible two stage designs. Statistics in Medicine 25 (19), 3366–3381

  4. [4]

    Crowder, M. J. (1976). Maximum likelihood estimation for dependent observations.Journal of the Royal Statistical Society. Series B (Methodological) 38 (1), 45–53

  5. [5]

    Demets, D. L. and K. K. G. Lan (1994). Interim analysis: The alpha spending function approach. Statistics in Medicine 13 (13–14), 1341–1352

  6. [6]

    Dispenzieri, A., J. A. Katzmann, R. A. Kyle, D. R. Larson, T. M. Therneau, C. L. Colby, R. J. Clark, G. P. Mead, S. Kumar, L. J. Melton, et al. (2012). Use of nonclonal serum immunoglobulin free light chains to predict overall survival in the general population. In Mayo Clinic Proceedings, Volume 87, pp. 517–523. Elsevier. 27

  7. [7]

    Ferguson, T. (1996). A Course in Large Sample Theory . New York: Routledge

  8. [8]

    Fleming, T. R. and D. P. Harrington (1991). Counting Processes and Survival Analysis . John Wiley & Sons

Show all 31 references
  1. [9]

    Liu, and C

    Gao, P., L. Liu, and C. Mehta (2013). Exact inference for adaptive group sequential designs. Statistics in medicine 32 (23), 3991–4005

  2. [10]

    Graf, A. C., G. Gutjahr, and W. Brannath (2016). Precision of maximum likelihood estimation in adaptive designs. Statistics in Medicine 35 (6), 922–941. sim.6761

  3. [11]

    Ivanova, A. and N. Flournoy (2001).A birth and death urn for ternary outcomes: stochastic processes applied to urn models . Chapman and Hall/CRC, Boca Raton

  4. [12]

    Ivanova, A., W. F. Rosenberger, S. D. Durham, and N. Flournoy (2000). A birth and death urn for randomized clinical trials: Asymptotic methods. Sankhy, Series B 62 (1), 104–118

  5. [13]

    Jennison, C. and B. Turnbull (1999). Group Sequential Methods with Applications to Clin- ical Trials. Chapman & Hall/CRC Interdisciplinary Statistics. CRC Press

  6. [14]

    Koopmeiners, J. S., Z. Feng, and M. S. Pepe (2012). Conditional estimation after a two- stage diagnostic biomarker study that allows early termination for futility. Statistics in Medicine 31 (5), 420–435

  7. [15]

    Lane, A. and N. Flournoy (2012). Two-stage adaptive optimal design with fixed first-stage sample size. Journal of Probability and Statistics 2012 , n/a–n/a

  8. [16]

    Li, G., W. J. Shih, T. Xie, and J. Lu (2002). A sample size adjustment procedure for clinical trials based on conditional power. Biostatistics 3 (2), 277–287. 28

  9. [17]

    Liu, A. and W. Hall (1999). Unbiased estimation following a group sequential test. Biometrika 86 (1), 71–78

  10. [18]

    Liu, A., W. Hall, K. F. Yu, and C. Wu (2006). Estimation following a group sequential test for distributions in the one-parameter exponential family. Statistica Sinica 16 (1), 165–181

  11. [19]

    Marschner, I. C. and I. M. Schou (2018). Underestimation of treatment effects in sequen- tially monitored clinical trials that did not stop early for benefit. Statistical Methods in Medical Research 0 (0), 0962280218795320

  12. [20]

    Martens, M. J. and B. R. Logan (2018). A group sequential test for treatment effect based on the fine–gray model. Biometrics 74 (3), 1006–1013

  13. [21]

    May, C. and N. Flournoy (2009). Asymptotics in response-adaptive designs generated by a two-color, randomly reinforced urn. The Annals of Statistics 37 (2), 1058–1078

  14. [22]

    Mehta, C. R., P. Bauer, M. Posch, and W. Brannath (2007). Repeated confidence intervals for adaptive group sequential trials. Statistics in Medicine 26 (30), 5422–5433

  15. [23]

    Molenberghs, A

    Milanzi, E., G. Molenberghs, A. Alonso, M. G. Kenward, A. A. Tsiatis, M. Davidian, and G. Verbeke (2015). Estimation after a group sequential trial. Statistics in bio- sciences 7 (2), 187–205

  16. [24]

    Pampallona, S., A. A. Tsiatis, and K. Kim (2001). Interim monitoring of group sequen- tial trials using spending functions for the type i and type ii error probabilities. Drug Information Journal 35 (4), 1113–1121. 29

  17. [25]

    Philippou, A. N., G. Roussas, et al. (1973). Asymptotic distribution of the likelihood func- tion in the independent not identically distributed case. The Annals of Statistics 1 (3), 454–471

  18. [26]

    Proschan, M. A., K. K. G. Lan, and J. T. Wittes (2006). Statistical Monitoring of Clinical Trials: A Unified Approach . Springer

  19. [27]

    Schou, I. M. and I. C. Marschner (2013). Meta-analysis of clinical trials with early stopping: an investigation of potential bias. Statistics in medicine 32 (28), 4859–4874

  20. [28]

    Gosho, and A

    Shimura, M., M. Gosho, and A. Hirakawa (2017). Comparison of conditional bias- adjusted estimators for interim analysis in clinical trials with survival data. Statistics in medicine 36 (13), 2067–2080

  21. [29]

    Tarima, S. and N. Flournoy (2019). Asymptotic properties of maximum likelihood estima- tors with sample size recalculation. Statistical Papers 60 , 23–44

  22. [30]

    Flournoy, and E

    Wang, H., N. Flournoy, and E. Kpamegan (2014). A new bounded log-linear regression model. Metrika 77 (5), 695–720

  23. [31]

    Whitehead, J. (1986). On the bias of maximum likelihood estimation following a sequential test. Biometrika 73 , 573–581. 30

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.