{"id":"024c371f-8591-4986-b7be-e18e2310e912","arxiv_id":"1908.01411","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"In group sequential designs, estimators are mixtures of truncated distributions and common information-fraction calculations are wrong; a new design based on ordered effect sizes avoids them.","lead":"Clinical trials that stop early based on interim results produce estimates and confidence intervals with distorted, non-normal distributions, and this paper derives the exact form of that distortion. The authors also propose a new way to design these trials that avoids the usual 'information fraction' calculations.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 4's local-asymptotic formula (19) omits the noncentrality h in the stopping boundary and integrand, so the claimed asymptotic mixture is not the actual local limit; the reader's Theorem 1 objection, by contrast, does not land because the extra factor is constant in t.","rationale":"The reader's weakest assumption is that equations (14)-(15) invalidate Theorem 1 because Pr_theta(D=k) depends on theta. This objection does not land: in the conditional likelihood ratio, the factor Pr_theta0(D=k)/Pr_theta(D=k) is constant in the observed statistic t once the event D=k is fixed, so multiplying a monotone likelihood ratio by a positive constant preserves monotonicity. Karlin-Rubin therefore still applies; the conditional family need not be the original exponential family. The paper's literal statement LRc(t)=LR(t) is false but harmless. The genuine load-bearing problem is in Section 4: the local asymptotic distribution drops the noncentrality h from the stopping boundary and integrand. This is a concrete mathematical error in the paper's second major quantitative claim, although the qualitative 'mixtures of truncated normals' finding remains true after shifting the truncation limits. The new GSD in Section 3 and the information identity in Section 2 do not depend on the asymptotic formula, so the manuscript is fixable by correcting (19), Section 4.2, and the K-stage formula (22), and by specifying the joint conditional distribution of the xi_j given D=k. The reader's conditional verdict is therefore still appropriate, but for a different reason than the one stated in the reader's weakest_assumption.","tokens_in":14815,"tokens_out":19589,"duration_ms":202633,"concrete_test":"Re-derive equation (19) for the two-stage normal case with n1=n2=100, c1=2.18, h=1 (theta=0.1). Simulate 10^6 trials: draw Z1,Z2 iid N(1,1), stop at stage 1 if Z1>2.18, otherwise continue, and compute V(D)=I(D=1)(Z1-1)+I(D=2)(Z1+Z2-2)/sqrt(2). Compare the empirical CDF of V(D) with the paper's expression, which uses p1=alpha/2, upper limit c1=2.18 and integrand Phi(sqrt(2)v-y)phi(y)/Phi(2.18), and with the corrected expression using p1=1-Phi(1.18), upper limit 1.18, and integrand Phi(sqrt(2)v-(y-1))phi(y-1)/Phi(1.18). If the empirical CDF matches the corrected expression but not the paper's expression, the missing h-shift is confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.1's central asymptotic result is not correct as written. Under theta = h/sqrt(n1), the stage-1 statistic satisfies Z1 ~ N(h,1), so the event D=2 is {Z1 <= c1} = {xi1 <= c1 - h}, and the conditional density of Z1 given D=2 is phi(y-h)/Phi(c1-h). The paper instead writes the integral in (19) with phi(y|D=2) and, in the Pocock example, the integral from -infty to c1 of Phi(sqrt(2)v - y) phi(y)/Pr(y<=c1) dy, which are the h=0 expressions. With n1=n2, V(2) = (xi1+xi2)/sqrt(2), so the correct integrand is Phi(sqrt(2)v - (y-h)) and the upper limit is c1-h; p1 = 1 - Phi(c1-h), not the null value alpha/2. Equation (19) and Section 4.2 therefore do not give the local limiting distribution. The K-stage formula (22) has the same defect: conditioning on D=k imposes shifted truncation constraints on the xi_j, but no such constraints appear. The qualitative conclusion that the limit is a non-normal mixture of truncated normals survives the correction, but the quantitative content of Section 4 is not established as written.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript studies maximum likelihood estimation and information measures in group sequential designs (GSDs) whose stopping rules depend on the parameter of interest. It argues that early stopping changes the support of the data distribution, so MLEs have distributions that are mixtures of truncated distributions; that unconditioned MLEs carry more Fisher information than MLEs conditioned on the stopping stage; that the usual 'information fraction' based on fixed-sample information is not generally correct; and that a design can instead be specified directly by ordered alternative hypotheses and desired stage-wise power. Section 3 proposes such a design and claims, via a monotone likelihood ratio argument, that the resulting tests are most powerful for a fixed alpha-spending function. Section 4 gives local asymptotic distributional formulas for the standardized MLE under two-stage and K-stage GSDs, claiming non-normal mixtures of truncated normals. A Cox regression example illustrates the design. The paper's central qualitative claims are plausible, but the local asymptotic formulas in Section 4 omit the noncentrality parameter and the corresponding truncation shifts, so the quantitative content of that section is not established as written.","tokens_in":15071,"tokens_out":7110,"duration_ms":73596,"significance":"If corrected, the paper would make a useful contribution. Its finite-sample representation of GSD MLEs as mixtures of truncated distributions is clearly illustrated by simulation, and the information decomposition in Eq. (10) — that conditioning on stopping loses exactly the Fisher information in the stopping rule — is clean and potentially valuable for practitioners. The proposed design driven by ordered alternatives is a practical alternative to information-fraction-based designs, and the Cox example demonstrates its applicability. The paper also correctly emphasizes that the canonical multivariate normality assumption fails after early stopping. The main reservation is the local asymptotic theory: as printed, the formulas in Section 4 are the null-limits (h=0) and do not describe the claimed local alternative limits. Because those formulas underlie the paper's principal asymptotic claims, the manuscript needs substantive revision before the results can be relied upon.","major_comments":[{"comment":"Equation (19) omits the local noncentrality parameter. Under theta = h/sqrt(n1), the stage-1 statistic satisfies Z1 ~ N(h,1), so the event D=2 is {Z1 <= c1} = {xi1 <= c1 - h} and the conditional density of Z1 given D=2 is phi(xi1)/Phi(c1-h) with xi1 = Z1 - h, not phi(y)/Pr(y<=c1). The limiting stage-1 stopping probability is p1 = 1 - Phi(c1 - h), not the null value alpha/2. The integral therefore needs a shifted argument and a shifted upper truncation limit; as printed, Eq. (19) gives the h=0 limit. The quantitative local asymptotic claim of Section 4 is not established.","section":"Section 4.1, Eq. (19)"},{"comment":"The K-stage formula (22) has the same defect as Eq. (19). The event r(D)=r(k) corresponds to D=k, which imposes joint constraints on the stage-wise statistics, e.g. xi1 <= c1 - h, xi2 <= c2 - h, and so on, and the stage-wise statistics have non-zero means under local alternatives. None of these shifts or truncation constraints appear in (22), so the displayed expression is not the limiting distribution under theta = h/sqrt(n1). The qualitative claim that the limit is a non-normal mixture of truncated normals may survive a correction, but the formulas as given are incorrect.","section":"Section 4.3, Eq. (22)"},{"comment":"The inequality I(k) <= Ifix(k) at theta = hattheta does not follow from the sign of the second derivative of log fX(k). Comparing (11) with (9) involves different normalizing constants and truncated supports; the integrand's sign alone cannot establish the inequality. In the normal example of Section 2.3, Iobs(k) = n(k) is constant, so I(k) = Ifix(k) rather than a strict inequality. The dashed and dotted curves in Figure 4 and the surrounding text need to be reconciled with this observation.","section":"Section 2.2, Eq. (11)"}],"minor_comments":[{"comment":"The statement that 'For every D=k, LRc(t)=LR(t)' is not correct: from the displayed formulas, LRc(t) equals LR(t) multiplied by Prtheta0(D=k)/Prtheta(D=k). Because this factor is constant in t, the monotone likelihood ratio conclusion still holds, but the displayed equality should be corrected and the proof rephrased accordingly.","section":"Section 3.3, Eqs. (14)-(15)"},{"comment":"The notation p1 = lim_{n1->infinity} Prtheta(D=1) is ambiguous because theta changes with n1 in the local alternative setting; the limiting probability should be written explicitly as a function of h, e.g. 1 - Phi(c1 - h) in the example.","section":"Section 4.1"},{"comment":"There is a typo: 'the probability to reject the converges to 1' should read 'the probability to reject the null hypothesis converges to 1'.","section":"Section 1"},{"comment":"The caption's phrases 'thick right is for stopping at stage 1 and a thick left for stage 2' are unclear; please indicate which curve corresponds to I(1), I(2), Ic(1), and Ic(2) explicitly.","section":"Figure 4 caption"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read through Tarima and Flournoy. The core message is solid: early stopping breaks the joint normality of sequential test statistics, the MLE is a mixture of truncated distributions, and information fractions computed from fixed-sample formulas are not the right quantities. The identity in equation (10) — that the information lost by conditioning on stopping equals the Fisher information in the stopping rule — is clean and worth remembering. The ordered-alternative design in Section 3 is a genuinely practical idea, and the simulation work in Section 2.3 is a nice illustration of the finite-sample effects.\n\nNow the soft spots, in order of severity.\n\nThe big one is Section 4. Equation (19) and the Pocock example in 4.2 are not the actual local limits. Under theta = h/sqrt(n1), the stage-1 statistic has mean h, so the stopping boundary is c1 - h and the conditional density has the shifted normal form phi(y-h)/Phi(c1-h). The paper writes the h = 0 version. The same issue runs through the K-stage formula (22): the truncation constraints are not shifted. The qualitative message — non-normal mixtures of truncated normals — survives, but the quantitative content of that section is not established.\n\nThe reader's objection to Theorem 1 does not land. It is true that LRc(t) is not exactly equal to LR(t); the extra factor Pr_theta0(D=k)/Pr_theta(D=k) is not 1. But that factor is constant in the sufficient statistic t, so the likelihood ratio remains monotone in t. The Karlin-Rubin argument goes through. The proof should state the constant rather than claiming equality, but the theorem itself is fine.\n\nTwo smaller issues. The claim I(k) <= Ifix(k) at theta = hat_theta does not follow from the sign of the integrand; conditioning on a subset can increase expected information. And the design section would benefit from a reproducible solver — there's no code or general algorithm, and the survival data example doesn't specify how the cohort is ordered into stages.\n\nWho gets value: methodologists working on group sequential estimation and design. The identity (10) and the ordered-alternative design are the parts I'd take away. Section 4 needs a corrected derivation before I'd rely on its formulas.\n\nMy vote: send to peer review, but the referee should require rewriting Section 4 with the local shifts and either fixing or rephrasing the LRc equality in Theorem 1.","headline":"Useful message about non-normal MLE mixtures and information loss, but the asymptotic section has a fixable shift error; the UMP claim survives its flawed proof.","tokens_in":15633,"tokens_out":7413,"would_cite":true,"duration_ms":72414,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62L05","62L10","62F12"],"pacs":[],"model":"deepseek-v4-flash","headline":"Interim stopping changes maximum likelihood estimates into mixtures of truncated distributions, and conditioning on the stopping stage loses exactly the Fisher information carried by the stopping rule.","keywords":["group sequential designs","adaptive designs","maximum likelihood estimation","asymptotic distribution theory","interim analyses","local alternative hypotheses","information loss","truncated distributions"],"falsifier":"Simulate the two-stage Pocock design with $X_i\\sim N(h/\\sqrt{n_1},1)$, $n_1=n_2=100$, and $c_1=2.18$ for several $h$, and compare the empirical CDF of $V(D)$ with equation (19); if the mixture-of-truncated-normals limit is not within Monte Carlo error, the paper's asymptotic claim fails. For the most-powerful claim, evaluate the conditional likelihood ratio $LR^c(t)$ at a fixed $k$: any non-monotonicity in $t$ disproves the theorem.","tokens_in":14525,"feed_emoji":"📊","tokens_out":14458,"duration_ms":134640,"temperature":0.7,"pith_summary":"Group sequential designs let a trial stop early, and this paper argues that the stopping rule is part of the statistical model: it cuts the sample space down, so maximum likelihood estimators no longer have their usual distributions. The paper shows that the MLE is a mixture of truncated distributions, one component for each stage at which the trial can stop, and that under local alternatives the asymptotic distribution is a non-normal mixture of truncated normals. It also shows that conditioning on the observed stopping stage discards exactly the Fisher information contained in the stopping rule, so standard information fractions computed from fixed-sample information overstate what an interim look provides. As a remedy, it proposes group sequential designs in which stage-specific sample sizes and critical values are chosen to guarantee a preset power against a sequence of ordered alternatives, avoiding the need to compute adapted information fractions. The practical point is that trials that can stop early need estimators, confidence statements, and design calculations reflecting the adaptation rule actually used.","feed_headline":"Early stopping turns trial estimates into truncated mixtures","feed_subtitle":"Early stopping makes trial estimates non-normal mixtures; conditioning on stopping loses information.","key_machinery":"The central object is the sub-density form of the likelihood under a group sequential design, $f_{X_{(D)}}(x_{(D)}|\\theta)=\\prod_{k=1}^K \\left[f_{X_{(k)}}(x_{(k)}|\\theta)\\right]^{I(D=k)}$, in which each stopping event $D=k$ contributes a density supported only on the region of the sample space that the stopping rule leaves reachable. Conditioning on stopping divides that sub-density by $\\Pr_\\theta(D=k)$, which inserts the term $\\partial^2\\log\\Pr_\\theta(D=k)/\\partial\\theta^2$ into the observed information and is the source of the information-loss identity. The proposed design machinery is a recursive system that fixes stage-specific type I error and powers each stage against a different ordered alternative. The asymptotic machinery is the limiting random variable $r(D)=\\sum_{k=1}^K I(D=k)\\lim n_{(k)}/n_k$, which determines how stagewise normal components combine into the mixture of truncated normals.","core_discovery":"On its own terms, the paper establishes that when the decision to stop at stage $D$ depends on the parameter of interest, the maximum likelihood estimator is $\\hat\\theta(D)=\\sum_{k=1}^K I(D=k)\\hat\\theta_{(k)}$, a mixture whose components are distributions truncated to the region of support left reachable by the stopping rule. Under local alternatives $\\theta=h/\\sqrt{n_1}$, the standardized estimator $V(D)$ converges to a mixture of truncated normal distributions: equation (19) for two stages and equation (22) for $K$ stages, with the mixing weights driven by the limiting ratios $r_{(k)}=\\lim n_{(k)}/n_k$; under fixed alternatives the test statistic degenerates to a point mass. The information result is the identity $I-I^c=E_D[-\\partial^2\\log\\Pr_\\theta(D)/\\partial\\theta^2]$, the Fisher information about $\\theta$ in the stopping rule itself. The design contribution is a group sequential scheme in which a sequence of ordered alternative hypotheses $\\theta_1>\\theta_2>\\cdots>\\theta_K>0$, each powered at $1-\\beta$, determines stage-specific sample sizes and critical values, so adapted information fractions never have to be computed.","pith_inferences":["A simulation check the paper does not run is whether equation (19) still tracks the empirical distribution of $V(D)$ when stagewise estimators come from binary or time-to-event endpoints; the regularity setup suggests it should, and this would extend the result beyond normal and survival-model examples.","The identity for $I-I^c$ supplies a practical diagnostic: estimate the stopping probabilities under a working model and compare conditional with unconditional information; designs whose stopping probabilities depend steeply on the parameter will show the largest penalties.","Treating unreached patient data as structural zeros rather than missing data is a convention the paper chooses explicitly, so likelihood-based model selection and multiple-imputation approaches would systematically disagree with its information calculations; that disagreement is a conceptual choice, not a numerical error.","The ordered-alternative design can be generalized to unequal stagewise error spending, trading early-stage power for smaller expected sample size while keeping the no-information-fraction feature."],"forward_implications":["Trialists should not quote the usual fixed-sample information fraction $n_{(k)}/n_{(K)}$ when planning interim looks; the adapted fraction $I_{(k)}/I_{(K)}$ is the relevant quantity and is generally smaller.","Post-hoc analyses that condition on the observed stopping stage pay a measurable precision penalty equal to the Fisher information about the treatment effect in the stopping rule, and near a stopping boundary the conditional MLE can be badly biased.","Under local alternatives, normal approximations for interim test statistics give way to mixtures of truncated normals, so repeated confidence intervals built on normality can be miscalibrated.","A design powered separately for each ordered alternative gives each interim look a clinically meaningful power guarantee, and its operating characteristics can be read from stagewise rejection probabilities and expected sample size without computing information fractions."],"supporting_citations":[{"why":"Supplies the recursive sub-density formula used throughout to compute stagewise stopping distributions under non-normality.","marker":"Armitage et al. (1969)"},{"why":"Defines the joint canonical normality assumption and the standard GSD framework that the paper argues fails under early stopping.","marker":"Jennison and Turnbull (1999)"},{"why":"Provides the local-asymptotic mixture result for sample-size recalculation that the two-stage and K-stage asymptotic formulas extend.","marker":"Tarima and Flournoy (2019)"},{"why":"Introduced the conditional maximum likelihood estimator whose bias, variance, and information are analyzed here.","marker":"Koopmeiners et al. (2012)"},{"why":"Gives earlier conditional and unconditional information measures in GSDs; the paper rederives them and finds the same information loss.","marker":"Marschner and Schou (2018)"},{"why":"Introduced alpha- and beta-spending functions with a single alternative; the new design extends this to ordered alternatives.","marker":"Pampallona et al. (2001)"},{"why":"Established the parameter-dependent change in support for one-parameter exponential families that motivates the truncation and mixture analysis.","marker":"Liu and Hall (1999)"},{"why":"Recognized truncation in the joint distribution of stage-specific test statistics, a direct antecedent of the mixture-of-truncated-distributions result.","marker":"Schou and Marschner (2013)"}],"fun_headline_variants":["Stopping rules twist trial estimates into mixtures","Interim stops make estimates non-normal mixtures","Conditioning on stopping loses Fisher information","New design skips information-fraction math","Group sequential estimates become truncated mixtures"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that each stage's test is most powerful assumes that conditioning on the trial having reached that stage does not change the relative likelihood of different parameter values; because the probability of stopping at each stage depends on the parameter itself, that assumption can fail, and with it the optimality claim.","fun_headline_variants_meta":{"raw":{"variants":["Stopping rules twist trial estimates into mixtures","Interim stops make estimates non-normal mixtures","Conditioning on stopping loses Fisher information","New design skips information-fraction math","Group sequential estimates become truncated mixtures"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000625,"raw_usage":{"total_tokens":2896,"prompt_tokens":953,"completion_tokens":1943,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":569,"completion_tokens_details":{"reasoning_tokens":1880}},"tokens_in":569,"tokens_out":1943,"duration_ms":16809,"temperature":1.0,"reasoning_tokens":1880,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:14:22.322998+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the two-stage Pocock design with $X_i\\sim N(h/\\sqrt{n_1},1)$, $n_1=n_2=100$, and $c_1=2.18$ for several $h$, and compare the empirical CDF of $V(D)$ with equation (19); if the mixture-of-truncated-normals limit is not within Monte Carlo error, the paper's asymptotic claim fails. For the most-powerful claim, evaluate the conditional likelihood ratio $LR^c(t)$ at a fixed $k$: any non-monotonicity in $t$ disproves the theorem.","supporting_citations":[{"cited_title":"McPherson, and B","cited_arxiv_id":null,"evidence_quote":"Supplies the recursive sub-density formula used throughout to compute stagewise stopping distributions under non-normality."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the joint canonical normality assumption and the standard GSD framework that the paper argues fails under early stopping."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the local-asymptotic mixture result for sample-size recalculation that the two-stage and K-stage asymptotic formulas extend."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduced the conditional maximum likelihood estimator whose bias, variance, and information are analyzed here."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives earlier conditional and unconditional information measures in GSDs; the paper rederives them and finds the same information loss."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduced alpha- and beta-spending functions with a single alternative; the new design extends this to ordered alternatives."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Established the parameter-dependent change in support for one-parameter exponential families that motivates the truncation and mixture analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Recognized truncation in the joint distribution of stage-specific test statistics, a direct antecedent of the mixture-of-truncated-distributions result."}],"review_version":1}