REVIEW 4 major objections 9 minor 29 references
Mitigating the Winner's Curse While Controlling Multiplicity: e-Process Methods for Anytime-Valid Inference in Dose-Ranging Trials
T0 review · 4 major / 9 minor · reviewed 2026-07-09 · glm-5.2
Pith's one-line read Subtract a selection charge, get anytime-valid dose trials
desk verdict Anytime-valid global GO test for dose-ranging trials via a selection-premium charge; the core Type I argument is self-contained, but the companion-paper dependency for the bias interpretation needs independent verification. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The selection premium function sp_L(u), defined as the expected gain from re-selecting the leader after one more update minus the current maximum. The cumulative score G_t(delta), the accumulated charge A_t(delta), and the residual R_t = G_t - A_t. The e-process E_t(delta) built via half-normal mixture over lambda of exp(lambda R_t - lambda^2 V_t / 2). The GO rule: delta_hat_max >= delta_0 + A_t/n_t + q_alpha(V_t; rho)/n_t.
What would settle it
Construct a dose-outcome distribution within the composite null where the accumulated selection premium does not dominate the conditional drift of the cumulative score G_t, causing the residual R_t to have positive drift and the e-process to lose its supermartingale property, violating the Type I error bound in finite samples.
Extended reading notes
Core claim
The central mechanism is the bridge from the selection-premium identity to an e-process. The identity states that the expected value of the running maximum of centered dose-control scores equals the sum of predictable one-step selection premiums. This means the winner's curse is not an unmanageable bias but a quantifiable, state-dependent drift that can be subtracted in real time. Once subtracted, the residual satisfies the conditional drift and sub-Gaussian conditions needed to exponentiate into a nonnegative supermartingale. Mixing over the tuning parameter with a half-normal density yields a closed-form e-process, and inverting the candidate-margin tests produces an anytime lower bound on
Load-bearing premise
The entire construction rests on the selection-premium identity from a companion paper, which says the expected running maximum of the centered dose-control scores equals the sum of expected one-step selection premiums. If this identity or its stopped-time analogue fails under the paper's assumptions, the drift bound on the residual breaks, and the e-process is no longer a valid supermartingale.
Editorial extensions
If this is right
- Trial sponsors can monitor accumulating dose-ranging data continuously and stop at any data-dependent time without inflating the false-GO probability, eliminating the need to pre-specify a fixed interim-analysis schedule.
- The selection-charge ledger provides trial teams with an interpretable, real-time accounting of how much of the observed best-dose effect is attributable to selection optimism versus genuine signal, improving transparency in go/no-go decisions.
- The anytime lower confidence sequence for the best dose effect can be reported alongside the decision rule, giving a continuously updated, valid lower bound on program-level efficacy throughout the trial.
- The method's separation of bias correction and multiplicity control into two additive terms on the effect scale allows practitioners to see the cost of each distortion independently, which could inform design choices about number of doses and monitoring frequency.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper develops an anytime-valid global testing procedure for Phase II dose-ranging trials that simultaneously addresses the winner's curse (selection optimism from choosing the empirical best dose) and Type I error inflation from multiplicity across doses and interim looks. The core construction subtracts a predictable 'selection charge'—derived from a selection-premium identity for the running maximum of dose-control scores—from the cumulative score, yielding a residual with nonpositive conditional drift under the composite null. A mixture-exponential e-process is then built on this residual, providing finite-sample Type I control via Ville's inequality. The authors give plug-in implementations for Gaussian and binary endpoints, prove finite-sample Type I control, derive anytime lower confidence sequences for the best dose effect, and illustrate the method through a worked example and simulations comparing against Dunnett-style procedures and an armwise union e-process.
Significance. The paper addresses a practically important problem in dose-ranging trials: combining winner's-curse correction with anytime-valid multiplicity control. The decision rule has an appealing operational form (raw best effect minus selection charge minus monitoring margin), and the binary implementation yields a closed-form charge depending only on the number of tied leaders, which is computationally transparent. The simulations are informative, particularly Table 4, which demonstrates the Type I inflation from misusing planned-look thresholds at unplanned looks and from omitting the selection charge. The authors are commendably explicit about limitations (no dose-response trend modeling, balanced blocks only, global rather than named-dose inference, no universal power dominance). The mathematical derivation from the drift bound (Eq. 8) through the e-process construction (Theorem 3.1) is standard but correct, and the sub-Gaussian variance bounds are verified directly via Borell's inequality (Gaussian) and Hoeffding's lemma (binary).
major comments (4)
- The central Type I error claim depends on the drift inequality (Eq. 8): E{ΔG_t(δ) | F_{t-1}} ≤ a_t(δ). The proofs of this inequality in Theorem 4.3 (Gaussian, Eq. 23) and Theorem 5.4 (binary, Eq. 33) are presented as self-contained, using coordinatewise monotonicity of the maximum under H_0 (d_k = δ_k − δ ≤ 0) and Jensen's inequality for the covariance bound. I have verified these arguments and they appear sound. The reader's concern about dependence on the companion paper's selection-premium identity (Eq. 5) for the central claim is, on examination, not load-bearing: Eq. 5 is invoked only in Propositions 4.2, 5.3, and 5.6 for the bias-interpretation results, not in the proofs of Theorems 4.3 or 5.4. However, the manuscript should state this separation of concerns more explicitly, since the abstract and introduction (§2.1) prominently feature Eq. 5 as the 'mathematical starting point,' a
- §4.1, Assumption 4.1: The Gaussian model requires a prespecified covariance bound (Σ̄_D ⪰ Σ_D, σ̄²_0 ≥ σ²_0) obtained from 'external information or an independent variance-estimation stage.' The paper acknowledges in §9 that re-estimating variance from the same unblinded efficacy stream is not justified. This is a significant practical limitation that restricts applicability. The authors should comment on how commonly such external variance information is available in Phase II dose-ranging settings, and whether a blinded variance adaptation extension is feasible within the current framework.
- Table 4: The proposed selection-premium rule reports a false-GO rate of 1.35% under the all-boundary null (every δ_k = 0.05), which is substantially below the nominal α = 0.05. This conservatism is expected given the nuisance-robust premium (maximizing over θ ∈ Θ(δ)), but the degree of conservatism is notable. The armwise union e-process achieves 0.80%, which is also conservative but less so. Can the authors quantify how much of the conservatism in the proposed method comes from (a) the selection charge being an upper bound, (b) the half-normal mixture boundary, and (c) the Bonferroni-type structure? This would help practitioners understand the efficiency tradeoff.
- Table 5: The armwise union e-process outperforms the proposed method in GO-by-500 probability for the moderate (60.3% vs. 57.1%) and strong (97.0% vs. 94.8%) alternatives, and is only slightly lower in the near-threshold case (10.0% vs. 11.5%). The authors acknowledge this in the text ('the proposed method is not offered as a uniformly most powerful procedure'), but the distinct advantage of the selection-premium method—its interpretable ledger form and lower confidence sequence for δ*—is not quantified against the armwise union. Can the authors provide a scenario or metric where the selection-premium method has a concrete operational advantage beyond interpretability?
minor comments (9)
- §2.1, Eq. (4): The notation sp_L(u) uses L to denote the law of a 'fresh centered null update,' but the subscript L is not consistently defined across subsequent uses (e.g., sp_G in §4.1, sp^B_r in §5.1). A brief note clarifying the relationship between the generic sp_L and the endpoint-specific versions would help readers.
- §3.3, Eq. (16): The targeted choice ρ = c*_α V* is motivated, but the specific value ρ = 13.41 for the binary simulations is stated without derivation. The text says 'for α = 0.05 the minimizer in (16) is approximately c*_{0.05} = 0.1494,' but the arithmetic connecting 0.1494, V* ≈ 89.7, and ρ = 13.41 should be shown explicitly (13.41 / 89.7 ≈ 0.1495, which checks out, but readers should not have to verify this).
- Figure 1, Panel C: The caption states 'The red curve is the exact expected fixed-look winner optimism E(max_k X̄_{t,k} − μ*).' It would help to clarify how this 'exact' expectation was computed—analytically or by simulation—and whether the green curve's 3,500 simulated paths are the same paths used for the red curve.
- §7, Table 1: The 'Anytime lower floor' column shows negative values (−12.6%, −5.1%, −3.2%) at early looks. A brief note explaining that negative lower floors are expected early (before sufficient evidence accumulates) would prevent confusion for practitioners accustomed to confidence bounds that are always within the parameter space.
- §5.1, Eq. (28): The binary premium sp^B_r(θ) = 1 − (1−θ)^r − θ is derived for the case of r tied leaders. The text should clarify whether 'tied leaders' means tied on cumulative response counts (which is how rt is defined in Eq. 26), and whether ties at zero cumulative counts (early in the trial) are handled the same way.
- References: The companion paper [de la Pena et al., 2026] is cited as arXiv:2602.19481. Given that results from this paper are used for the bias-interpretation propositions (4.2, 5.3, 5.6), its publication status should be noted if available, and the overlap in authorship between the two papers should be transparently disclosed.
- §8.2: The all-boundary null uses p_k = 0.35 for all doses, giving δ_k = 0.05 = δ_0. This is the boundary of H_0(δ_0). It would be informative to also report Type I error under a strict null (e.g., all δ_k = 0 or all δ_k = 0.03) to show how conservatism varies with the distance from the boundary.
- Typo: §1.1, the table defining symbols lists 'εt,k centered noise, X t,k − E(X t,k)' with a stray space after X.
- Typo: §4.1, Eq. (19), the formula sp_G(u; Σ̄_D) = ωφ(d/ω) − d{1 − Φ(d/ω)} should specify that ω² = Var(ε^G_1 − ε^G_2), which is stated in the preceding text but not in the equation itself.
Simulated Author's Rebuttal
We thank the referee for a careful and constructive report. The referee verified the central drift inequality and the e-process construction, and raised four substantive points: (1) clarifying that the selection-premium identity (Eq. 5) is used only for bias interpretation, not for the Type I error proofs; (2) commenting on the practical limitation of requiring prespecified variance bounds in the Gaussian case; (3) requesting a decomposition of the conservatism seen in Table 4; and (4) requesting a concrete operational advantage of the selection-premium method beyond interpretability. We address each below. The manuscript will be revised to incorporate these points.
read point-by-point responses
-
Referee: The manuscript should state more explicitly that Eq. 5 is invoked only in Propositions 4.2, 5.3, and 5.6 for bias-interpretation results, not in the proofs of Theorems 4.3 or 5.4, despite the abstract and introduction featuring Eq. 5 as the 'mathematical starting point.'
Authors: We agree with the referee's observation and will make the separation of concerns explicit in the revised manuscript. Specifically, we will add a sentence at the end of Section 2.1 noting that the selection-premium identity (Eq. 5) is used for the bias-interpretation results (Propositions 4.2, 5.3, and 5.6) but is not invoked in the proofs of the drift inequalities (Theorems 4.3 and 5.4), which rely on coordinatewise monotonicity of the maximum and conditional Jensen's inequality. We will also add a clarifying remark in the introduction to the effect that Eq. 5 provides the conceptual motivation and the bias-correction interpretation, while the Type I error guarantee rests on the drift bound (Eq. 8) and the sub-Gaussian condition (Eq. 9) alone. revision: yes
-
Referee: The Gaussian model requires a prespecified covariance bound from external information or an independent variance-estimation stage. The authors should comment on how commonly such external variance information is available in Phase II dose-ranging settings and whether a blinded variance adaptation extension is feasible.
Authors: This is a genuine practical limitation, and we agree it deserves more discussion. In the revised manuscript we will add commentary to Section 4.1 and Section 9 addressing both parts of the question. Regarding availability: in Phase II dose-ranging trials, external variance information is often available from Phase I data, prior trials of the same compound in related indications, or published benchmarks for the endpoint (e.g., change-from-baseline in blood pressure, PANSS scores). When such information is unavailable, an independent pilot or blinded interim variance estimation is common practice. Regarding feasibility of blinded adaptation: the current proof requires the variance bound to be deterministic and prespecified because the e-process supermartingale property depends on the sub-Gaussian proxy being predictable. A blinded variance update (using only the pooled treatment-control data without unblinding dose assignments) would preserve predictability of the bound and is feasible in principle, since the pooled variance estimate does not reveal dose-control contrasts. We will note this as a concrete extension and sketch why the proof structure accommodates it, while flagging that unblinded variance re-estimation from the efficacy stream remains outside the current guarantee. revision: partial
-
Referee: Table 4 shows the proposed method has a false-GO rate of 1.35% under the all-boundary null, well below the nominal 0.05. Can the authors quantify how much of the conservatism comes from (a) the selection charge being an upper bound, (b) the half-normal mixture boundary, and (c) the Bonferroni-type structure?
Authors: We can provide a partial decomposition. The three sources of conservatism are partially separable through additional simulation runs that we will add to the revised manuscript. (a) The selection charge being an upper bound: the 'no selection charge' row in Table 4 (false-GO rate 7.26%) is not a clean decomposition because removing the charge also removes the drift correction, so the 7.26% figure reflects Type I inflation rather than pure conservatism. However, we can quantify the charge's contribution to conservatism by comparing the false-GO rate with the exact-equal-margin charge (using the true boundary rate theta rather than the supremum over Theta(delta)) versus the robust charge. (b) The half-normal mixture boundary: we can isolate this by computing the false-GO rate of a single-dose (K=1) e-process with the same mixture boundary and Vt = t/2, which removes both the selection charge and the multiplicity structure. (c) The Bonferroni-type structure: the armwise union e-process (0.80%) provides a comparison point, as it uses Bonferroni allocation but no selection charge. We will add a small supplementary table with these decomposition runs. We note that the three sources interact (e.g., the selection charge and the mixture boundary both depend on the same Vt), so the decomposition is additive only approximately. revision: partial
-
Referee: Table 5 shows the armwise union e-process outperforms the proposed method in GO-by-500 probability for moderate and strong alternatives. Can the authors provide a scenario or metric where the selection-premium method has a concrete operational advantage beyond interpretability?
Authors: We appreciate this pointed question and will address it in two ways in the revision. First, the anytime lower confidence sequence for delta* (Theorem 6.1) is a concrete inferential output that the armwise union e-process does not provide: the armwise procedure controls each named-dose null but does not yield a lower confidence bound for the best true effect delta*. We will make this contrast explicit by noting that the selection-premium method's lower floor can be reported at any look as 'the best true dose effect is at least X% with anytime validity,' which is a distinct operational deliverable. Second, we will add a simulation scenario with a larger number of doses (e.g., K=10 or K=12) where the Bonferroni penalty in the armwise union becomes more severe while the selection-premium charge depends on the number of tied leaders (which typically decreases as K grows, since ties become less persistent). In such settings, we expect the selection-premium method to become more competitive or superior in GO probability. We are honest that we have not yet run this simulation and cannot guarantee the outcome; if the advantage does not materialize, we will report that and frame the lower confidence sequence as the primary distinguishing output. revision: partial
Circularity Check
No significant circularity: the central Type I control claim is proven self-containedly; the self-cited selection premium identity is used only for interpretive bias-correction results.
full rationale
The paper's central claim—finite-sample Type I error control via the e-process in Theorem 3.1—depends on two conditions: the drift bound (Eq. 8) and the sub-Gaussian condition (Eq. 9). Both are proven directly and self-containedly in Theorem 4.3 (Gaussian, via coordinatewise monotonicity of the maximum under H_0 and Jensen's inequality for the covariance bound) and Theorem 5.4 (binary, via p_k ≤ p_0 + δ and Hoeffding's lemma). Neither proof invokes the selection premium identity (Eq. 5) from the companion paper [de la Pena et al., 2026]. The e-process construction itself (Theorem 3.1) is standard: given nonpositive drift b_t ≥ 0 and sub-Gaussian increments, the exponential supermartingale L_t(λ;δ) = exp{λR_t - λ²V_t/2} is a supermartingale, and Ville's inequality yields Type I control. The self-citation to [de la Pena et al., 2026] serves three roles: (1) defining the selection premium function sp_L (Eq. 4), which is a straightforward mathematical definition verifiable independently; (2) stating the identity E[M_t] = Σ E[sp_L(S_{s-1})] (Eq. 5), which is used only in Propositions 4.2, 5.3, and 5.6 for the interpretive claim that the accumulated charge equals exact winner's-curse bias at the equal-margin reference; and (3) citing envelope and decay properties of sp_L, which support the heterogeneous-effect interpretation but not the validity of the test. If Eq. 5 were false, the charge would lose its bias-correction interpretation, but the e-process would remain valid because the drift inequality (Eq. 8) is established by direct calculation. The self-citation is therefore not load-bearing for the central result. The only minor concern is that the selection premium function's definition originates from the companion paper, but since it is a parameter-free mathematical definition (an expectation of a maximum minus the current maximum), it does not create circularity. Score 2 reflects the presence of a self-citation that is not load-bearing for the central claim.
Assumptions & free parameters
free parameters (2)
- rho (half-normal tuning constant) =
13.41
- V* (target variance budget) =
~89.7
assumptions (3)
- domain assumption Selection premium identity: E[M_t] = sum of E{sp_L(S_{s-1})}
- standard math Sub-Gaussian conditional increments
- domain assumption Prespecified covariance upper bound
Cite this review
Pith. "Pith review of Mitigating the Winner's Curse While Controlling Multiplicity: e-Process Methods for Anytime-Valid Inference in Dose-Ranging Trials." pith.science (2026). https://pith.science/paper/NSKQ2ZCD
@misc{pith2026260707699,
author = {Pith},
title = {Pith review of: Mitigating the Winner's Curse While Controlling Multiplicity: e-Process Methods for Anytime-Valid Inference in Dose-Ranging Trials},
year = {2026},
howpublished = {\url{https://pith.science/paper/NSKQ2ZCD}},
note = {Machine review of arXiv:2607.07699}
}
read the original abstract
Phase II dose-ranging trials often report the largest observed dose-control effect while inspecting accumulating data repeatedly. This creates two coupled distortions: selection optimism from choosing the empirical winner, known as the winner's curse, and Type I error inflation from multiplicity across doses and interim looks. We develop an anytime-valid procedure for testing whether the best true dose effect exceeds a clinically meaningful margin. The mathematical starting point is a recent selection-premium identity for the running maximum: for dose-control scores, the expected gain from re-selecting the current leader becomes a predictable selection charge. Subtracting this charge gives a residual with nonpositive drift under the composite null; applying a one-sided mixture-exponential construction then yields an e-process and hence an anytime-valid global test. The resulting rule has a transparent ledger form: raw best effect minus selection charge minus monitoring margin, and a ``GO'' decision is made only when the remaining evidence still exceeds the clinical margin. We give plug-in implementations for Gaussian and binary outcomes, prove finite-sample Type I control and anytime lower confidence bounds for the best dose effect, and illustrate the method through a worked example and simulations.
Figures
Reference graph
Works this paper leans on
-
[1]
arXiv preprint arXiv:2602.19481 , year=
A Selection Premium Decomposition for the Expected Maximum of Random Walks , author=. arXiv preprint arXiv:2602.19481 , year=
-
[2]
Journal of the American Statistical Association , volume=
A multiple comparison procedure for comparing several treatments with a control , author=. Journal of the American Statistical Association , volume=. 1955 , publisher=
work page 1955
-
[3]
Statistics in medicine , volume=
Sequential designs for phase III clinical trials incorporating treatment selection , author=. Statistics in medicine , volume=. 2003 , publisher=
work page 2003
-
[4]
A generalized Dunnett test for multi-arm multi-stage clinical studies with treatment selection , author=. Biometrika , volume=. 2012 , publisher=
work page 2012
-
[5]
Statistical methods in medical research , volume=
A multi-stage drop-the-losers design for multi-arm clinical trials , author=. Statistical methods in medical research , volume=. 2017 , publisher=
work page 2017
-
[6]
Evaluation of experiments with adaptive interim analyses , author=. Biometrics , pages=. 1994 , publisher=
work page 1994
-
[7]
Statistics in medicine , volume=
Adaptive Dunnett tests for treatment selection , author=. Statistics in medicine , volume=. 2008 , publisher=
work page 2008
-
[8]
Biometrical Journal: Journal of Mathematical Methods in Biosciences , volume=
A comparison of methods for adaptive treatment selection , author=. Biometrical Journal: Journal of Mathematical Methods in Biosciences , volume=. 2008 , publisher=
work page 2008
Show all 29 references
-
[9]
Statistical Methods in Medical Research , volume=
Comparing the MAMS framework with the combination method in multi-arm adaptive trials with binary outcomes , author=. Statistical Methods in Medical Research , volume=. 2019 , publisher=
2019
-
[10]
Biometrics , volume=
Combining multiple comparisons and modeling techniques in dose-response studies , author=. Biometrics , volume=. 2005 , publisher=
2005
-
[11]
The Quarterly Journal of Economics , volume=
Inference on winners , author=. The Quarterly Journal of Economics , volume=. 2024 , publisher=
2024
-
[12]
The Annals of Statistics , volume=
Time-uniform, nonparametric, nonasymptotic confidence sequences , author=. The Annals of Statistics , volume=. 2021 , publisher=
2021
-
[13]
Statistical Science , volume=
Game-theoretic statistics and safe anytime-valid inference , author=. Statistical Science , volume=. 2023 , publisher=
2023
-
[14]
arXiv preprint arXiv:2603.17925 , year=
Multi-Armed Sequential Hypothesis Testing by Betting , author=. arXiv preprint arXiv:2603.17925 , year=
-
[15]
Time-uniform Chernoff bounds via nonnegative supermartingales , author=
-
[16]
2024 , publisher=
Clinical trials: a methodologic perspective , author=. 2024 , publisher=
2024
-
[17]
1994 , howpublished =
1994
-
[18]
Adaptive Designs for Clinical Trials of Drugs and Biologics: Guidance for Industry , year =
-
[19]
Multiple Endpoints in Clinical Trials: Guidance for Industry , year =
-
[20]
Optimizing the Dosage of Human Prescription Drugs and Biological Products for the Treatment of Oncologic Diseases: Guidance for Industry , year =
-
[21]
2015 , howpublished =
Qualification of the. 2015 , howpublished =
2015
-
[22]
1999 , publisher=
Group sequential methods with applications to clinical trials , author=. 1999 , publisher=
1999
-
[23]
2010 , publisher=
Bayesian adaptive methods for clinical trials , author=. 2010 , publisher=
2010
-
[24]
Statistics in medicine , volume=
Optimal design of multi-arm multi-stage trials , author=. Statistics in medicine , volume=. 2012 , publisher=
2012
-
[25]
Scandinavian journal of statistics , pages=
A simple sequentially rejective multiple test procedure , author=. Scandinavian journal of statistics , pages=. 1979 , publisher=
1979
-
[26]
Statistics in medicine , volume=
Model-based dose finding under model uncertainty using general parametric models , author=. Statistics in medicine , volume=. 2014 , publisher=
2014
-
[27]
2014 , howpublished =
Qualification Opinion of. 2014 , howpublished =
2014
-
[28]
Journal of Statistical Software , volume=
The R package MAMS for designing multi-arm multi-stage clinical trials , author=. Journal of Statistical Software , volume=
-
[29]
Summer school on machine learning , pages=
Concentration inequalities , author=. Summer school on machine learning , pages=. 2003 , publisher=
2003
Reviewed July 9, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.