{"id":"bf2c833d-5bf0-4c54-bd35-0c2a12b07581","arxiv_id":"2607.07699","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":2,"one_line_summary":"The paper derives an anytime-valid global test for dose-ranging trials that subtracts a predictable 'selection charge' from the running maximum dose effect to correct winner's curse bias while controlling Type I error.","lead":"This paper develops a statistical method for dose-ranging clinical trials that corrects for the 'winner's curse' (overestimating the best dose's effect) while allowing flexible, repeated data monitoring without inflating false-positive rates. A smart generalist might read it to understand how modern 'anytime-valid' statistics can make drug-trial decisions more reliable and transparent.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"Reader's concern about Eq. 5 is not load-bearing for the central Type I claim; the drift bound (Eq. 8) is proven directly in Theorems 4.3 and 5.4 without invoking the companion paper's identity.","rationale":"The reader's verdict is CONDITIONAL based on the concern that the selection premium identity (Eq. 5) from the companion paper is unverified and load-bearing. My careful reading shows this concern does not land on the central claim. The Type I error control (Theorem 3.1, Eq. 10–12) rests on the drift bound (Eq. 8) and sub-Gaussian condition (Eq. 9), both of which are proven directly in Theorems 4.3 and 5.4 without invoking Eq. 5. The companion paper's identity is used only for the bias-interpretation propositions (4.2, 5.3, 5.6), which are secondary results about the winner's-curse correction interpretation. Even if the companion paper contained an error, the e-process would remain valid because the drift inequality is established independently. The mathematical argument for the central claim is sound: the monotonicity argument (d_k ≤ 0 under H_0 implies the heterogeneous drift is bounded by the equal-margin premium) is elementary and correct; the sub-Gaussian verification via Borell/Hoeffding is standard; the mixture-exponential construction and Ville's inequality are well-established. The paper does have acknowledged practical limitations (conservatism: 1.35% vs 5% Type I in Table 4; lower power than the armwise union e-process in some scenarios per Table 5), but these are honestly reported and do not constitute correctness risks. The lack of code release and the unavailable companion paper are legitimate transparency concerns, but they do not affect the validity of the central mathematical argument as presented. I would keep the verdict at CONDITIONAL but for a different reason than the reader states: the condition should be code release for reproducibility of simulations, not verification of Eq. 5. The correctness risk for the central claim is low, not unknown.","tokens_in":17081,"tokens_out":8179,"duration_ms":655878,"concrete_test":"Verify the drift bound (Eq. 8) directly by simulation under the all-boundary null (p_k = 0.35, p_0 = 0.30, δ_0 = 0.05, K = 6 binary case). At each update t across 100,000 simulated paths, compute the empirical conditional drift E[ΔG_t(δ_0) | F_{t-1}] (by conditioning on the pre-update state) and compare it to the selection charge sp^B_{r_t}(δ_0). If the empirical drift exceeds the charge for any state configuration, the drift bound fails and the e-process validity is in question. This test is self-contained and does not require the companion paper.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader identifies the selection premium identity (Eq. 5) from the companion paper [de la Pena et al., 2026] as the weakest assumption underlying the central claim. However, Eq. 5 is an unconditional identity about E[M_t] = Σ E[sp_L(S_{s-1})], used in Propositions 4.2 and 5.3 to interpret the accumulated charge as an exact winner's-curse bias correction at the equal-margin reference. The central claim — finite-sample Type I control via the e-process in Theorem 3.1 — depends on the conditional drift bound (Eq. 8): E{ΔG_t(δ) | F_{t-1}} ≤ a_t(δ). This bound is proven directly and self-containedly in Theorem 4.3 (Gaussian) and Theorem 5.4 (binary), using coordinatewise monotonicity of the maximum (d_k = δ_k − δ ≤ 0 under H_0) and Jensen's inequality for the covariance bound. No step in these proofs invokes Eq. 5 or its stopped-time analogue. The sub-Gaussian condition (Eq. 9) is also verified directly via Borell's inequality (Gaussian) and Hoeffding's lemma (binary). The e-process construction (Theorem 3.1) is then standard: b_t = a_t − E{ΔG_t|F_{t-1}} ≥ 0 implies e^{−λb_t} ≤ 1 for λ ≥ 0, yielding the supermartingale property, and Ville's inequality gives Type I control. The companion paper's identity is needed only for the secondary bias-interpretation results (Propositions 4.2, 5.3, 5.6), not for the validity of the test. If Eq. 5 were wrong, the charge would lose its interpretation as an exact bias correction, but the e-process would remain valid because the drift inequality still holds. The reader's concern thus targets a secondary claim, not the load-bearing one.","agreement_with_reader":"disagree"},"referee_report":{"model":"glm-5.2","summary":"This paper develops an anytime-valid global testing procedure for Phase II dose-ranging trials that simultaneously addresses the winner's curse (selection optimism from choosing the empirical best dose) and Type I error inflation from multiplicity across doses and interim looks. The core construction subtracts a predictable 'selection charge'—derived from a selection-premium identity for the running maximum of dose-control scores—from the cumulative score, yielding a residual with nonpositive conditional drift under the composite null. A mixture-exponential e-process is then built on this residual, providing finite-sample Type I control via Ville's inequality. The authors give plug-in implementations for Gaussian and binary endpoints, prove finite-sample Type I control, derive anytime lower confidence sequences for the best dose effect, and illustrate the method through a worked example and simulations comparing against Dunnett-style procedures and an armwise union e-process.","tokens_in":17514,"tokens_out":2169,"duration_ms":411746,"significance":"The paper addresses a practically important problem in dose-ranging trials: combining winner's-curse correction with anytime-valid multiplicity control. The decision rule has an appealing operational form (raw best effect minus selection charge minus monitoring margin), and the binary implementation yields a closed-form charge depending only on the number of tied leaders, which is computationally transparent. The simulations are informative, particularly Table 4, which demonstrates the Type I inflation from misusing planned-look thresholds at unplanned looks and from omitting the selection charge. The authors are commendably explicit about limitations (no dose-response trend modeling, balanced blocks only, global rather than named-dose inference, no universal power dominance). The mathematical derivation from the drift bound (Eq. 8) through the e-process construction (Theorem 3.1) is standard but correct, and the sub-Gaussian variance bounds are verified directly via Borell's inequality (Gaussian) and Hoeffding's lemma (binary).","major_comments":[{"comment":"The central Type I error claim depends on the drift inequality (Eq. 8): E{ΔG_t(δ) | F_{t-1}} ≤ a_t(δ). The proofs of this inequality in Theorem 4.3 (Gaussian, Eq. 23) and Theorem 5.4 (binary, Eq. 33) are presented as self-contained, using coordinatewise monotonicity of the maximum under H_0 (d_k = δ_k − δ ≤ 0) and Jensen's inequality for the covariance bound. I have verified these arguments and they appear sound. The reader's concern about dependence on the companion paper's selection-premium identity (Eq. 5) for the central claim is, on examination, not load-bearing: Eq. 5 is invoked only in Propositions 4.2, 5.3, and 5.6 for the bias-interpretation results, not in the proofs of Theorems 4.3 or 5.4. However, the manuscript should state this separation of concerns more explicitly, since the abstract and introduction (§2.1) prominently feature Eq. 5 as the 'mathematical starting point,' a","section":null},{"comment":"§4.1, Assumption 4.1: The Gaussian model requires a prespecified covariance bound (Σ̄_D ⪰ Σ_D, σ̄²_0 ≥ σ²_0) obtained from 'external information or an independent variance-estimation stage.' The paper acknowledges in §9 that re-estimating variance from the same unblinded efficacy stream is not justified. This is a significant practical limitation that restricts applicability. The authors should comment on how commonly such external variance information is available in Phase II dose-ranging settings, and whether a blinded variance adaptation extension is feasible within the current framework.","section":null},{"comment":"Table 4: The proposed selection-premium rule reports a false-GO rate of 1.35% under the all-boundary null (every δ_k = 0.05), which is substantially below the nominal α = 0.05. This conservatism is expected given the nuisance-robust premium (maximizing over θ ∈ Θ(δ)), but the degree of conservatism is notable. The armwise union e-process achieves 0.80%, which is also conservative but less so. Can the authors quantify how much of the conservatism in the proposed method comes from (a) the selection charge being an upper bound, (b) the half-normal mixture boundary, and (c) the Bonferroni-type structure? This would help practitioners understand the efficiency tradeoff.","section":null},{"comment":"Table 5: The armwise union e-process outperforms the proposed method in GO-by-500 probability for the moderate (60.3% vs. 57.1%) and strong (97.0% vs. 94.8%) alternatives, and is only slightly lower in the near-threshold case (10.0% vs. 11.5%). The authors acknowledge this in the text ('the proposed method is not offered as a uniformly most powerful procedure'), but the distinct advantage of the selection-premium method—its interpretable ledger form and lower confidence sequence for δ*—is not quantified against the armwise union. Can the authors provide a scenario or metric where the selection-premium method has a concrete operational advantage beyond interpretability?","section":null}],"minor_comments":[{"comment":"§2.1, Eq. (4): The notation sp_L(u) uses L to denote the law of a 'fresh centered null update,' but the subscript L is not consistently defined across subsequent uses (e.g., sp_G in §4.1, sp^B_r in §5.1). A brief note clarifying the relationship between the generic sp_L and the endpoint-specific versions would help readers.","section":null},{"comment":"§3.3, Eq. (16): The targeted choice ρ = c*_α V* is motivated, but the specific value ρ = 13.41 for the binary simulations is stated without derivation. The text says 'for α = 0.05 the minimizer in (16) is approximately c*_{0.05} = 0.1494,' but the arithmetic connecting 0.1494, V* ≈ 89.7, and ρ = 13.41 should be shown explicitly (13.41 / 89.7 ≈ 0.1495, which checks out, but readers should not have to verify this).","section":null},{"comment":"Figure 1, Panel C: The caption states 'The red curve is the exact expected fixed-look winner optimism E(max_k X̄_{t,k} − μ*).' It would help to clarify how this 'exact' expectation was computed—analytically or by simulation—and whether the green curve's 3,500 simulated paths are the same paths used for the red curve.","section":null},{"comment":"§7, Table 1: The 'Anytime lower floor' column shows negative values (−12.6%, −5.1%, −3.2%) at early looks. A brief note explaining that negative lower floors are expected early (before sufficient evidence accumulates) would prevent confusion for practitioners accustomed to confidence bounds that are always within the parameter space.","section":null},{"comment":"§5.1, Eq. (28): The binary premium sp^B_r(θ) = 1 − (1−θ)^r − θ is derived for the case of r tied leaders. The text should clarify whether 'tied leaders' means tied on cumulative response counts (which is how rt is defined in Eq. 26), and whether ties at zero cumulative counts (early in the trial) are handled the same way.","section":null},{"comment":"References: The companion paper [de la Pena et al., 2026] is cited as arXiv:2602.19481. Given that results from this paper are used for the bias-interpretation propositions (4.2, 5.3, 5.6), its publication status should be noted if available, and the overlap in authorship between the two papers should be transparently disclosed.","section":null},{"comment":"§8.2: The all-boundary null uses p_k = 0.35 for all doses, giving δ_k = 0.05 = δ_0. This is the boundary of H_0(δ_0). It would be informative to also report Type I error under a strict null (e.g., all δ_k = 0 or all δ_k = 0.03) to show how conservatism varies with the distance from the boundary.","section":null},{"comment":"Typo: §1.1, the table defining symbols lists 'εt,k centered noise, X t,k − E(X t,k)' with a stray space after X.","section":null},{"comment":"Typo: §4.1, Eq. (19), the formula sp_G(u; Σ̄_D) = ωφ(d/ω) − d{1 − Φ(d/ω)} should specify that ω² = Var(ε^G_1 − ε^G_2), which is stated in the preceding text but not in the equation itself.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The reader's report raised a concern about circularity stemming from the companion paper [de la Pena et al., 2026], specifically that the selection-premium identity (Eq. 5) is unproven in the current text and the companion paper shares overlapping authors. On examination, this concern does not land as a load-bearing issue for the central Type I claim: the drift bounds in Theorems 4.3 and 5.4 are proven directly without invoking Eq. 5, and the e-process construction via Ville's inequality is standard. The companion paper's identity is used only for the secondary bias-interpretation results. That said, the overlap in authorship between this paper and the cited companion should be transparently disclosed, and the manuscript should clarify which results depend on the companion paper and which are self-contained. The paper is a solid methodological contribution; the issues raised are presentational and clarificatory rather than structural."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful and constructive report. The referee verified the central drift inequality and the e-process construction, and raised four substantive points: (1) clarifying that the selection-premium identity (Eq. 5) is used only for bias interpretation, not for the Type I error proofs; (2) commenting on the practical limitation of requiring prespecified variance bounds in the Gaussian case; (3) requesting a decomposition of the conservatism seen in Table 4; and (4) requesting a concrete operational advantage of the selection-premium method beyond interpretability. We address each below. The manuscript will be revised to incorporate these points.","responses":[{"response":"We agree with the referee's observation and will make the separation of concerns explicit in the revised manuscript. Specifically, we will add a sentence at the end of Section 2.1 noting that the selection-premium identity (Eq. 5) is used for the bias-interpretation results (Propositions 4.2, 5.3, and 5.6) but is not invoked in the proofs of the drift inequalities (Theorems 4.3 and 5.4), which rely on coordinatewise monotonicity of the maximum and conditional Jensen's inequality. We will also add a clarifying remark in the introduction to the effect that Eq. 5 provides the conceptual motivation and the bias-correction interpretation, while the Type I error guarantee rests on the drift bound (Eq. 8) and the sub-Gaussian condition (Eq. 9) alone.","revision_made":"yes","referee_comment":"The manuscript should state more explicitly that Eq. 5 is invoked only in Propositions 4.2, 5.3, and 5.6 for bias-interpretation results, not in the proofs of Theorems 4.3 or 5.4, despite the abstract and introduction featuring Eq. 5 as the 'mathematical starting point.'"},{"response":"This is a genuine practical limitation, and we agree it deserves more discussion. In the revised manuscript we will add commentary to Section 4.1 and Section 9 addressing both parts of the question. Regarding availability: in Phase II dose-ranging trials, external variance information is often available from Phase I data, prior trials of the same compound in related indications, or published benchmarks for the endpoint (e.g., change-from-baseline in blood pressure, PANSS scores). When such information is unavailable, an independent pilot or blinded interim variance estimation is common practice. Regarding feasibility of blinded adaptation: the current proof requires the variance bound to be deterministic and prespecified because the e-process supermartingale property depends on the sub-Gaussian proxy being predictable. A blinded variance update (using only the pooled treatment-control data without unblinding dose assignments) would preserve predictability of the bound and is feasible in principle, since the pooled variance estimate does not reveal dose-control contrasts. We will note this as a concrete extension and sketch why the proof structure accommodates it, while flagging that unblinded variance re-estimation from the efficacy stream remains outside the current guarantee.","revision_made":"partial","referee_comment":"The Gaussian model requires a prespecified covariance bound from external information or an independent variance-estimation stage. The authors should comment on how commonly such external variance information is available in Phase II dose-ranging settings and whether a blinded variance adaptation extension is feasible."},{"response":"We can provide a partial decomposition. The three sources of conservatism are partially separable through additional simulation runs that we will add to the revised manuscript. (a) The selection charge being an upper bound: the 'no selection charge' row in Table 4 (false-GO rate 7.26%) is not a clean decomposition because removing the charge also removes the drift correction, so the 7.26% figure reflects Type I inflation rather than pure conservatism. However, we can quantify the charge's contribution to conservatism by comparing the false-GO rate with the exact-equal-margin charge (using the true boundary rate theta rather than the supremum over Theta(delta)) versus the robust charge. (b) The half-normal mixture boundary: we can isolate this by computing the false-GO rate of a single-dose (K=1) e-process with the same mixture boundary and Vt = t/2, which removes both the selection charge and the multiplicity structure. (c) The Bonferroni-type structure: the armwise union e-process (0.80%) provides a comparison point, as it uses Bonferroni allocation but no selection charge. We will add a small supplementary table with these decomposition runs. We note that the three sources interact (e.g., the selection charge and the mixture boundary both depend on the same Vt), so the decomposition is additive only approximately.","revision_made":"partial","referee_comment":"Table 4 shows the proposed method has a false-GO rate of 1.35% under the all-boundary null, well below the nominal 0.05. Can the authors quantify how much of the conservatism comes from (a) the selection charge being an upper bound, (b) the half-normal mixture boundary, and (c) the Bonferroni-type structure?"},{"response":"We appreciate this pointed question and will address it in two ways in the revision. First, the anytime lower confidence sequence for delta* (Theorem 6.1) is a concrete inferential output that the armwise union e-process does not provide: the armwise procedure controls each named-dose null but does not yield a lower confidence bound for the best true effect delta*. We will make this contrast explicit by noting that the selection-premium method's lower floor can be reported at any look as 'the best true dose effect is at least X% with anytime validity,' which is a distinct operational deliverable. Second, we will add a simulation scenario with a larger number of doses (e.g., K=10 or K=12) where the Bonferroni penalty in the armwise union becomes more severe while the selection-premium charge depends on the number of tied leaders (which typically decreases as K grows, since ties become less persistent). In such settings, we expect the selection-premium method to become more competitive or superior in GO probability. We are honest that we have not yet run this simulation and cannot guarantee the outcome; if the advantage does not materialize, we will report that and frame the lower confidence sequence as the primary distinguishing output.","revision_made":"partial","referee_comment":"Table 5 shows the armwise union e-process outperforms the proposed method in GO-by-500 probability for moderate and strong alternatives. Can the authors provide a scenario or metric where the selection-premium method has a concrete operational advantage beyond interpretability?"}],"tokens_in":17464,"tokens_out":1406,"duration_ms":185641,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"The main thing to know: this paper builds an anytime-valid global GO test for Phase II dose-ranging trials by subtracting a predictable 'selection charge' from the running maximum dose-control score, then applying a standard mixture-exponential e-process. The decision rule is transparent — raw best effect minus selection charge minus monitoring margin, and you declare GO only if what remains clears the clinical margin. That is a clean and useful contribution to a genuine problem.","headline":"Anytime-valid global GO test for dose-ranging trials via a selection-premium charge; the core Type I argument is self-contained, but the companion-paper dependency for the bias interpretation needs independent verification.","tokens_in":18202,"tokens_out":168,"would_cite":true,"duration_ms":118999,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Subtract a selection charge, get anytime-valid dose trials","keywords":[],"falsifier":"Construct a dose-outcome distribution within the composite null where the accumulated selection premium does not dominate the conditional drift of the cumulative score G_t, causing the residual R_t to have positive drift and the e-process to lose its supermartingale property, violating the Type I error bound in finite samples.","tokens_in":17416,"feed_emoji":"📊","tokens_out":906,"duration_ms":200411,"temperature":0.7,"pith_summary":"Phase II dose-ranging trials that pick the best-looking dose and peek at data repeatedly face two distortions: the winner's curse (the selected dose looks better than it is) and inflated false-positive rates from multiple comparisons across doses and looks. This paper proposes a single corrective device for both. The key object is the selection premium: the expected one-step gain from being allowed to re-select the current leader after the next data block arrives. Under a composite null where no dose truly clears a clinical margin, this premium is a predictable, nonnegative quantity that exactly equals the upward drift of the running maximum dose-control score. Subtracting the accumulated premium from the raw maximum yields a residual with nonpositive drift. A one-sided mixture-exponential construction then converts this residual into an e-process, a nonnegative supermartingale whose crossing probability is bounded by Ville's inequality. The resulting decision rule is a transparent ledger: declare GO only when the raw best observed effect exceeds the clinical margin after paying both a per-look selection charge and a monitoring margin for repeated inspection. The paper provides plug-in implementations for Gaussian and binary endpoints, proves finite-sample Type I error control at arbitrary data-dependent looks, and derives anytime lower confidence sequences for the best true dose effect.","feed_headline":"Subtract a selection charge, get anytime-valid dose trials","feed_subtitle":"A predictable 'selection premium' turns the winner's curse into a drift charge, yielding an e-process that controls false-GO probability at ","key_machinery":"The selection premium function sp_L(u), defined as the expected gain from re-selecting the leader after one more update minus the current maximum. The cumulative score G_t(delta), the accumulated charge A_t(delta), and the residual R_t = G_t - A_t. The e-process E_t(delta) built via half-normal mixture over lambda of exp(lambda R_t - lambda^2 V_t / 2). The GO rule: delta_hat_max >= delta_0 + A_t/n_t + q_alpha(V_t; rho)/n_t.","core_discovery":"The central mechanism is the bridge from the selection-premium identity to an e-process. The identity states that the expected value of the running maximum of centered dose-control scores equals the sum of predictable one-step selection premiums. This means the winner's curse is not an unmanageable bias but a quantifiable, state-dependent drift that can be subtracted in real time. Once subtracted, the residual satisfies the conditional drift and sub-Gaussian conditions needed to exponentiate into a nonnegative supermartingale. Mixing over the tuning parameter with a half-normal density yields a closed-form e-process, and inverting the candidate-margin tests produces an anytime lower bound on","pith_inferences":[],"forward_implications":["Trial sponsors can monitor accumulating dose-ranging data continuously and stop at any data-dependent time without inflating the false-GO probability, eliminating the need to pre-specify a fixed interim-analysis schedule.","The selection-charge ledger provides trial teams with an interpretable, real-time accounting of how much of the observed best-dose effect is attributable to selection optimism versus genuine signal, improving transparency in go/no-go decisions.","The anytime lower confidence sequence for the best dose effect can be reported alongside the decision rule, giving a continuously updated, valid lower bound on program-level efficacy throughout the trial.","The method's separation of bias correction and multiplicity control into two additive terms on the effect scale allows practitioners to see the cost of each distortion independently, which could inform design choices about number of doses and monitoring frequency."],"fun_headline_variants":["Anytime-valid dose trials by subtracting a selection charge","Quantify the winner's curse for anytime-valid dose testing","A ledger approach to the winner's curse in dose-ranging trials","e-Process controls multiplicity in dose trials via a selection charge","Track a selection charge for anytime-valid dose-ranging trials"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The entire construction rests on the selection-premium identity from a companion paper, which says the expected running maximum of the centered dose-control scores equals the sum of expected one-step selection premiums. If this identity or its stopped-time analogue fails under the paper's assumptions, the drift bound on the residual breaks, and the e-process is no longer a valid supermartingale.","fun_headline_variants_meta":{"raw":{"variants":["Anytime-valid dose trials by subtracting a selection charge","Quantify the winner's curse for anytime-valid dose testing","A ledger approach to the winner's curse in dose-ranging trials","e-Process controls multiplicity in dose trials via a selection charge","Track a selection charge for anytime-valid dose-ranging trials","Make the winner's curse a predictable charge for dose trials"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":1284,"prompt_tokens":536,"completion_tokens":748,"prompt_tokens_details":null},"tokens_in":536,"tokens_out":748,"duration_ms":25311,"temperature":1.0,"reasoning_tokens":720,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-09T01:48:25.953113+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"Construct a dose-outcome distribution within the composite null where the accumulated selection premium does not dominate the conditional drift of the cumulative score G_t, causing the residual R_t to have positive drift and the e-process to lose its supermartingale property, violating the Type I error bound in finite samples.","supporting_citations":[],"review_version":1}