REVIEW 3 major objections 4 minor 31 references
Effect of Interim Adaptations in Group Sequential Designs
T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Interim stopping changes maximum likelihood estimates into mixtures of truncated distributions, and conditioning on the stopping stage loses exactly the Fisher information carried by the stopping rule.
desk verdict Useful message about non-normal MLE mixtures and information loss, but the asymptotic section has a fixable shift error; the UMP claim survives its flawed proof. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the sub-density form of the likelihood under a group sequential design, $f_{X_{(D)}}(x_{(D)}|\theta)=\prod_{k=1}^K \left[f_{X_{(k)}}(x_{(k)}|\theta)\right]^{I(D=k)}$, in which each stopping event $D=k$ contributes a density supported only on the region of the sample space that the stopping rule leaves reachable. Conditioning on stopping divides that sub-density by $\Pr_\theta(D=k)$, which inserts the term $\partial^2\log\Pr_\theta(D=k)/\partial\theta^2$ into the observed information and is the source of the information-loss identity. The proposed design machinery is a recursive system that fixes stage-specific type I error and powers each stage against a different ordered alternative. The asymptotic machinery is the limiting random variable $r(D)=\sum_{k=1}^K I(D=k)\lim n_{(k)}/n_k$, which determines how stagewise normal components combine into the mixture of truncated normals.
What would settle it
Simulate the two-stage Pocock design with $X_i\sim N(h/\sqrt{n_1},1)$, $n_1=n_2=100$, and $c_1=2.18$ for several $h$, and compare the empirical CDF of $V(D)$ with equation (19); if the mixture-of-truncated-normals limit is not within Monte Carlo error, the paper's asymptotic claim fails. For the most-powerful claim, evaluate the conditional likelihood ratio $LR^c(t)$ at a fixed $k$: any non-monotonicity in $t$ disproves the theorem.
Extended reading notes
Core claim
On its own terms, the paper establishes that when the decision to stop at stage $D$ depends on the parameter of interest, the maximum likelihood estimator is $\hat\theta(D)=\sum_{k=1}^K I(D=k)\hat\theta_{(k)}$, a mixture whose components are distributions truncated to the region of support left reachable by the stopping rule. Under local alternatives $\theta=h/\sqrt{n_1}$, the standardized estimator $V(D)$ converges to a mixture of truncated normal distributions: equation (19) for two stages and equation (22) for $K$ stages, with the mixing weights driven by the limiting ratios $r_{(k)}=\lim n_{(k)}/n_k$; under fixed alternatives the test statistic degenerates to a point mass. The information result is the identity $I-I^c=E_D[-\partial^2\log\Pr_\theta(D)/\partial\theta^2]$, the Fisher information about $\theta$ in the stopping rule itself. The design contribution is a group sequential scheme in which a sequence of ordered alternative hypotheses $\theta_1>\theta_2>\cdots>\theta_K>0$, each powered at $1-\beta$, determines stage-specific sample sizes and critical values, so adapted information fractions never have to be computed.
Load-bearing premise
The claim that each stage's test is most powerful assumes that conditioning on the trial having reached that stage does not change the relative likelihood of different parameter values; because the probability of stopping at each stage depends on the parameter itself, that assumption can fail, and with it the optimality claim.
Editorial extensions
If this is right
- Trialists should not quote the usual fixed-sample information fraction $n_{(k)}/n_{(K)}$ when planning interim looks; the adapted fraction $I_{(k)}/I_{(K)}$ is the relevant quantity and is generally smaller.
- Post-hoc analyses that condition on the observed stopping stage pay a measurable precision penalty equal to the Fisher information about the treatment effect in the stopping rule, and near a stopping boundary the conditional MLE can be badly biased.
- Under local alternatives, normal approximations for interim test statistics give way to mixtures of truncated normals, so repeated confidence intervals built on normality can be miscalibrated.
- A design powered separately for each ordered alternative gives each interim look a clinically meaningful power guarantee, and its operating characteristics can be read from stagewise rejection probabilities and expected sample size without computing information fractions.
Reading between the lines
- A simulation check the paper does not run is whether equation (19) still tracks the empirical distribution of $V(D)$ when stagewise estimators come from binary or time-to-event endpoints; the regularity setup suggests it should, and this would extend the result beyond normal and survival-model examples.
- The identity for $I-I^c$ supplies a practical diagnostic: estimate the stopping probabilities under a working model and compare conditional with unconditional information; designs whose stopping probabilities depend steeply on the parameter will show the largest penalties.
- Treating unreached patient data as structural zeros rather than missing data is a convention the paper chooses explicitly, so likelihood-based model selection and multiple-imputation approaches would systematically disagree with its information calculations; that disagreement is a conceptual choice, not a numerical error.
- The ordered-alternative design can be generalized to unequal stagewise error spending, trading early-stage power for smaller expected sample size while keeping the no-information-fraction feature.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript studies maximum likelihood estimation and information measures in group sequential designs (GSDs) whose stopping rules depend on the parameter of interest. It argues that early stopping changes the support of the data distribution, so MLEs have distributions that are mixtures of truncated distributions; that unconditioned MLEs carry more Fisher information than MLEs conditioned on the stopping stage; that the usual 'information fraction' based on fixed-sample information is not generally correct; and that a design can instead be specified directly by ordered alternative hypotheses and desired stage-wise power. Section 3 proposes such a design and claims, via a monotone likelihood ratio argument, that the resulting tests are most powerful for a fixed alpha-spending function. Section 4 gives local asymptotic distributional formulas for the standardized MLE under two-stage and K-stage GSDs, claiming non-normal mixtures of truncated normals. A Cox regression example illustrates the design. The paper's central qualitative claims are plausible, but the local asymptotic formulas in Section 4 omit the noncentrality parameter and the corresponding truncation shifts, so the quantitative content of that section is not established as written.
Significance. If corrected, the paper would make a useful contribution. Its finite-sample representation of GSD MLEs as mixtures of truncated distributions is clearly illustrated by simulation, and the information decomposition in Eq. (10) — that conditioning on stopping loses exactly the Fisher information in the stopping rule — is clean and potentially valuable for practitioners. The proposed design driven by ordered alternatives is a practical alternative to information-fraction-based designs, and the Cox example demonstrates its applicability. The paper also correctly emphasizes that the canonical multivariate normality assumption fails after early stopping. The main reservation is the local asymptotic theory: as printed, the formulas in Section 4 are the null-limits (h=0) and do not describe the claimed local alternative limits. Because those formulas underlie the paper's principal asymptotic claims, the manuscript needs substantive revision before the results can be relied upon.
major comments (3)
- [Section 4.1, Eq. (19)] Equation (19) omits the local noncentrality parameter. Under theta = h/sqrt(n1), the stage-1 statistic satisfies Z1 ~ N(h,1), so the event D=2 is {Z1 <= c1} = {xi1 <= c1 - h} and the conditional density of Z1 given D=2 is phi(xi1)/Phi(c1-h) with xi1 = Z1 - h, not phi(y)/Pr(y<=c1). The limiting stage-1 stopping probability is p1 = 1 - Phi(c1 - h), not the null value alpha/2. The integral therefore needs a shifted argument and a shifted upper truncation limit; as printed, Eq. (19) gives the h=0 limit. The quantitative local asymptotic claim of Section 4 is not established.
- [Section 4.3, Eq. (22)] The K-stage formula (22) has the same defect as Eq. (19). The event r(D)=r(k) corresponds to D=k, which imposes joint constraints on the stage-wise statistics, e.g. xi1 <= c1 - h, xi2 <= c2 - h, and so on, and the stage-wise statistics have non-zero means under local alternatives. None of these shifts or truncation constraints appear in (22), so the displayed expression is not the limiting distribution under theta = h/sqrt(n1). The qualitative claim that the limit is a non-normal mixture of truncated normals may survive a correction, but the formulas as given are incorrect.
- [Section 2.2, Eq. (11)] The inequality I(k) <= Ifix(k) at theta = hattheta does not follow from the sign of the second derivative of log fX(k). Comparing (11) with (9) involves different normalizing constants and truncated supports; the integrand's sign alone cannot establish the inequality. In the normal example of Section 2.3, Iobs(k) = n(k) is constant, so I(k) = Ifix(k) rather than a strict inequality. The dashed and dotted curves in Figure 4 and the surrounding text need to be reconciled with this observation.
minor comments (4)
- [Section 3.3, Eqs. (14)-(15)] The statement that 'For every D=k, LRc(t)=LR(t)' is not correct: from the displayed formulas, LRc(t) equals LR(t) multiplied by Prtheta0(D=k)/Prtheta(D=k). Because this factor is constant in t, the monotone likelihood ratio conclusion still holds, but the displayed equality should be corrected and the proof rephrased accordingly.
- [Section 4.1] The notation p1 = lim_{n1->infinity} Prtheta(D=1) is ambiguous because theta changes with n1 in the local alternative setting; the limiting probability should be written explicitly as a function of h, e.g. 1 - Phi(c1 - h) in the example.
- [Section 1] There is a typo: 'the probability to reject the converges to 1' should read 'the probability to reject the null hypothesis converges to 1'.
- [Figure 4 caption] The caption's phrases 'thick right is for stopping at stage 1 and a thick left for stage 2' are unclear; please indicate which curve corresponds to I(1), I(2), Ic(1), and Ic(2) explicitly.
Circularity Check
No significant circularity: information identities are derived from definitions, and the self-cited asymptotic theorem is used as independent stated-assumption support.
full rationale
The paper's central derivations do not reduce to their inputs by construction. The information loss identity in equation (10) follows algebraically from the definitions of I, Ic, and the conditional likelihood in (4), with no fitted parameters. The finite-sample distributions in Section 2 are derived from the stated sub-density likelihood (1), not assumed into the conclusion. The new design in Section 3 solves a system of power and error constraints for sample sizes and critical values, so operational characteristics are inputs, not outputs relabeled as predictions. The only self-citation is in Section 4, where the authors import Theorem 1 from Tarima and Flournoy (2019) for the local asymptotic behavior; this is a parameter-free theorem with stated regularity assumptions and does not smuggle the present conclusion into the citation. The UMP argument in Section 3.3 contains a slip in writing LRc(t)=LR(t) without the ratio Pr_theta0(D=k)/Pr_theta(D=k), but that ratio is constant in the sufficient statistic t for each fixed k, so the monotone likelihood-ratio conclusion does not circularly depend on the conclusion. Any concern that equation (19) omits the local shift h is a correctness or derivational issue, not an equivalence-by-construction between input and output, and therefore does not affect the circularity score.
Assumptions & free parameters
free parameters (1)
- Ordered alternative effect sizes for the example design =
theta1=0.3, theta2=0.2, theta3=0.1
assumptions (4)
- domain assumption Every stage is reachable with positive probability (Prθ(D=k)>0 for all k).
- domain assumption Stage-specific MLEs satisfy the asymptotic normality regularity conditions in Assumption (16).
- ad hoc to paper The conditional likelihood ratio LRc(t) equals the unconditional likelihood ratio LR(t) in exponential families with early stopping.
- ad hoc to paper The system of equations (13) has a solution with positive sample sizes and critical values.
Cite this review
Pith. "Pith review of Effect of Interim Adaptations in Group Sequential Designs." pith.science (2026). https://pith.science/paper/32RFK6CO
@misc{pith2026190801411,
author = {Pith},
title = {Pith review of: Effect of Interim Adaptations in Group Sequential Designs},
year = {2026},
howpublished = {\url{https://pith.science/paper/32RFK6CO}},
note = {Machine review of arXiv:1908.01411}
}
read the original abstract
This manuscript investigates unconditional and conditional-on-stopping maximum likelihood estimators (MLEs), information measures and information loss associated with conditioning in group sequential designs (GSDs). The possibility of early stopping brings truncation to the distributional form of MLEs; sequentially, GSD decisions eliminate some events from the sample space. Multiple testing induces mixtures on the adapted sample space. Distributions of MLEs are mixtures of truncated distributions. Test statistics that are asymptotically normal without GSD, have asymptotic distributions, under GSD, that are non-normal mixtures of truncated normal distributions under local alternatives; under fixed alternatives, asymptotic distributions of test statistics are degenerate. Estimation of various statistical quantities such as information, information fractions, and confidence intervals should account for the effect of planned adaptations. Calculation of adapted information fractions requires substantial computational effort. Therefore, a new GSD is proposed in which stage-specific sample sizes are fully determined by desired operational characteristics, and calculation of information fractions is not needed.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Armitage, P., C. McPherson, and B. Rowe (1969). Repeated significance tests on accumu- lating data. Journal of the Royal Statistical Society: Series A 132 , 235–244
work page 1969
-
[2]
Asendorf, T., R. Henderson, H. Schmidli, and T. Friede (2019). Sample size re-estimation for clinical trials with longitudinal negative binomial counts including time trends: Sample size re-estimation with longitudinal negative binomial counts. Statistics in Medicine 38 (9), 1503–1528
work page 2019
-
[3]
Brannath, W., F. Koenig, and P. Bauer (2006). Estimation in flexible two stage designs. Statistics in Medicine 25 (19), 3366–3381
work page 2006
-
[4]
Crowder, M. J. (1976). Maximum likelihood estimation for dependent observations.Journal of the Royal Statistical Society. Series B (Methodological) 38 (1), 45–53
work page 1976
-
[5]
Demets, D. L. and K. K. G. Lan (1994). Interim analysis: The alpha spending function approach. Statistics in Medicine 13 (13–14), 1341–1352
work page 1994
-
[6]
Dispenzieri, A., J. A. Katzmann, R. A. Kyle, D. R. Larson, T. M. Therneau, C. L. Colby, R. J. Clark, G. P. Mead, S. Kumar, L. J. Melton, et al. (2012). Use of nonclonal serum immunoglobulin free light chains to predict overall survival in the general population. In Mayo Clinic Proceedings, Volume 87, pp. 517–523. Elsevier. 27
work page 2012
-
[7]
Ferguson, T. (1996). A Course in Large Sample Theory . New York: Routledge
work page 1996
-
[8]
Fleming, T. R. and D. P. Harrington (1991). Counting Processes and Survival Analysis . John Wiley & Sons
work page 1991
Show all 31 references
-
[9]
Liu, and C
Gao, P., L. Liu, and C. Mehta (2013). Exact inference for adaptive group sequential designs. Statistics in medicine 32 (23), 3991–4005
2013
-
[10]
Graf, A. C., G. Gutjahr, and W. Brannath (2016). Precision of maximum likelihood estimation in adaptive designs. Statistics in Medicine 35 (6), 922–941. sim.6761
2016
-
[11]
Ivanova, A. and N. Flournoy (2001).A birth and death urn for ternary outcomes: stochastic processes applied to urn models . Chapman and Hall/CRC, Boca Raton
2001
-
[12]
Ivanova, A., W. F. Rosenberger, S. D. Durham, and N. Flournoy (2000). A birth and death urn for randomized clinical trials: Asymptotic methods. Sankhy, Series B 62 (1), 104–118
2000
-
[13]
Jennison, C. and B. Turnbull (1999). Group Sequential Methods with Applications to Clin- ical Trials. Chapman & Hall/CRC Interdisciplinary Statistics. CRC Press
1999
-
[14]
Koopmeiners, J. S., Z. Feng, and M. S. Pepe (2012). Conditional estimation after a two- stage diagnostic biomarker study that allows early termination for futility. Statistics in Medicine 31 (5), 420–435
2012
-
[15]
Lane, A. and N. Flournoy (2012). Two-stage adaptive optimal design with fixed first-stage sample size. Journal of Probability and Statistics 2012 , n/a–n/a
2012
-
[16]
Li, G., W. J. Shih, T. Xie, and J. Lu (2002). A sample size adjustment procedure for clinical trials based on conditional power. Biostatistics 3 (2), 277–287. 28
2002
-
[17]
Liu, A. and W. Hall (1999). Unbiased estimation following a group sequential test. Biometrika 86 (1), 71–78
1999
-
[18]
Liu, A., W. Hall, K. F. Yu, and C. Wu (2006). Estimation following a group sequential test for distributions in the one-parameter exponential family. Statistica Sinica 16 (1), 165–181
2006
-
[19]
Marschner, I. C. and I. M. Schou (2018). Underestimation of treatment effects in sequen- tially monitored clinical trials that did not stop early for benefit. Statistical Methods in Medical Research 0 (0), 0962280218795320
2018
-
[20]
Martens, M. J. and B. R. Logan (2018). A group sequential test for treatment effect based on the fine–gray model. Biometrics 74 (3), 1006–1013
2018
-
[21]
May, C. and N. Flournoy (2009). Asymptotics in response-adaptive designs generated by a two-color, randomly reinforced urn. The Annals of Statistics 37 (2), 1058–1078
2009
-
[22]
Mehta, C. R., P. Bauer, M. Posch, and W. Brannath (2007). Repeated confidence intervals for adaptive group sequential trials. Statistics in Medicine 26 (30), 5422–5433
2007
-
[23]
Molenberghs, A
Milanzi, E., G. Molenberghs, A. Alonso, M. G. Kenward, A. A. Tsiatis, M. Davidian, and G. Verbeke (2015). Estimation after a group sequential trial. Statistics in bio- sciences 7 (2), 187–205
2015
-
[24]
Pampallona, S., A. A. Tsiatis, and K. Kim (2001). Interim monitoring of group sequen- tial trials using spending functions for the type i and type ii error probabilities. Drug Information Journal 35 (4), 1113–1121. 29
2001
-
[25]
Philippou, A. N., G. Roussas, et al. (1973). Asymptotic distribution of the likelihood func- tion in the independent not identically distributed case. The Annals of Statistics 1 (3), 454–471
1973
-
[26]
Proschan, M. A., K. K. G. Lan, and J. T. Wittes (2006). Statistical Monitoring of Clinical Trials: A Unified Approach . Springer
2006
-
[27]
Schou, I. M. and I. C. Marschner (2013). Meta-analysis of clinical trials with early stopping: an investigation of potential bias. Statistics in medicine 32 (28), 4859–4874
2013
-
[28]
Gosho, and A
Shimura, M., M. Gosho, and A. Hirakawa (2017). Comparison of conditional bias- adjusted estimators for interim analysis in clinical trials with survival data. Statistics in medicine 36 (13), 2067–2080
2017
-
[29]
Tarima, S. and N. Flournoy (2019). Asymptotic properties of maximum likelihood estima- tors with sample size recalculation. Statistical Papers 60 , 23–44
2019
-
[30]
Flournoy, and E
Wang, H., N. Flournoy, and E. Kpamegan (2014). A new bounded log-linear regression model. Metrika 77 (5), 695–720
2014
-
[31]
Whitehead, J. (1986). On the bias of maximum likelihood estimation following a sequential test. Biometrika 73 , 573–581. 30
1986
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.