REVIEW 2 major objections 4 minor 11 references
Likelihood-Ratio E-Value Monitoring for Benchmark-Based Decisions in Early-Phase Oncology Trials
T0 review · 2 major / 4 minor · reviewed 2026-07-11 · grok-4.5
Pith's one-line read A likelihood-ratio e-process rule for phase II oncology gives count boundaries without Bayesian priors and cuts expected sample size while matching type I error.
desk verdict Solid, practical packaging of LR e-processes into BOP2-style count tables; the EN win for CEVAM-T is real under the stated calibration but not a pure out-of-sample test of the evidence scale. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The efficacy and reverse-futility likelihood-ratio e-processes ELR_n(p1,p0) and FLR_n(p0,p1), which are nonnegative supermartingales under their respective composite nulls and convert, by monotonicity, into explicit integer count boundaries reff(n;α) and rfut(n;β).
What would settle it
Re-run the same Monte Carlo comparison on a materially different look schedule or benchmark pair (for example denser early looks or p0=0.05, p1=0.20); if CEVAM-T no longer controls type I error near the nominal level or loses its expected-sample-size advantage while keeping similar power, the efficiency claim fails for that setting.
Extended reading notes
Core claim
For binary efficacy, the likelihood ratio of a design alternative p1 against the unacceptable rate p0 is an e-process under the composite null p ≤ p0; its reverse is an e-process under p ≥ p1. Both are monotone in the cumulative response count, so they translate into integer efficacy and futility boundaries that can be pretabulated. Planned-look calibration of the two evidence thresholds yields a design (CEVAM-T) that, on the standard binary benchmarks, matches type I error and rejection probability while achieving the smallest expected sample size of the designs compared.
Load-bearing premise
That a design-stage grid search over the two evidence thresholds, calibrated on one binary look schedule and one (p0,p1) pair, produces a fair efficiency comparison rather than an overfit to that particular configuration.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes CEVAM, a likelihood-ratio e-value framework for protocol-ready interim monitoring of binary (and, in the supplement, complex) endpoints against prespecified clinical benchmarks in early-phase oncology trials. For binary efficacy, it constructs an efficacy LR e-process targeted at design alternative p1 versus the least-favorable null boundary p0, and a reverse LR futility process, both of which are nonnegative supermartingales under the respective composite nulls (Theorem 2.1). Monotonicity yields pretabulated integer response-count boundaries (Proposition 2.2). The authors distinguish a fixed-threshold anytime-valid version from planned-look calibrated versions (CEVAM-NT and CEVAM-T) that grid-search evidence thresholds to control planned-look type I error. In binary simulations matching published BOP2-FE settings (N=40, looks 10–40, p0=0.20, p1=0.40, α*=0.10), CEVAM-T controlled nominal type I error and achieved the smallest expected sample size across true response rates while retaining rejection probabilities close to BOP2-FE (Tables 1–2). A retrospective re-analysis of the TREND trial pCR result is used as illustration.
Significance. If the results hold, CEVAM supplies a practically useful, analysis-prior-free alternative to BOP2/BOP2-FE-style posterior-cutoff designs while preserving the protocol-ready count-boundary format that makes those designs attractive. The fixed-threshold construction rests on standard e-process/supermartingale theory and is independently valuable for optional or irregular monitoring; the pretabulated monotone boundaries and public R code further strengthen implementability. The efficiency claim for CEVAM-T is of direct interest to early-phase oncology design, even though it is obtained under ordinary design-stage calibration rather than a pure out-of-sample test of the evidence scale. The work is a solid methodological contribution at the interface of sequential evidence and phase II oncology practice.
major comments (2)
- Section 2.2 and 3.1 / Table 2: The central efficiency claim for CEVAM-T (smallest EN across all true response rates while retaining similar PRN) is obtained by grid-searching (α_eff, β) on the same look schedule and (p0, p1, N) used for evaluation, with selection by closeness of planned-look type I error to α* then power/EN0. This is ordinary for calibrated phase II designs (and analogous to BOP2/BOP2-FE), but the manuscript should more explicitly state that the ranking is not a pure out-of-sample test of the LR evidence scale and should report at least one sensitivity check under a different look schedule or (p0, p1) pair so that readers can judge robustness of the EN advantage.
- Section 3.2 / Figure 1 and Table 2: CEVAM-T (and CEVAM-NT) allocate a larger fraction of false-positive probability at early looks than the BOP2-FE comparators. While overall type I error is controlled at the final look, the manuscript should discuss whether this early alpha spending is clinically acceptable for oncology phase II (e.g., risk of early false success stopping) and whether the calibration criterion could be modified to constrain early spending if desired.
minor comments (4)
- Section 4 / Table 3: The TREND re-analysis uses n=44 efficacy-evaluable patients rather than the protocol-specified interim of 48; the text already notes this, but a one-sentence caveat that the illustration is not a reconstruction of original trial conduct would further reduce any risk of over-interpretation.
- Proposition 2.2 / Eqs. (8)–(9): The ceiling/floor expressions for reff and rfut are useful; a brief numerical check that they match the Table 1 boundaries for the simulation setting would help readers implement the formulas.
- Throughout: Occasional typesetting artifacts (e.g., “Gr¨ unwald”, “Heff 0”) and the arXiv-style line breaks should be cleaned for journal production.
- References: BOP2-FE is cited as advance online publication; ensure the final citation is updated if a volume/page version is available at revision.
Circularity Check
No significant circularity: e-process validity and pretabulated boundaries are self-contained; calibration is ordinary design-stage OC search, not a prediction forced by construction.
full rationale
The paper's core derivation is standard and non-circular. Theorem 2.1 establishes that the likelihood-ratio process ELR_n(p1,p0) is a nonnegative supermartingale under the composite null p≤p0 (and the reverse process under p≥p1), yielding Ville-type bounds; this rests on classical martingale theory for Bernoulli likelihood ratios, not on any quantity defined from the target operating characteristics. Proposition 2.2 then converts the monotone LR statistics into explicit integer count boundaries reff(n;α) and rfut(n;β) by solving the inequalities for Rn; the formulas are algebraic rearrangements of the LR expressions, not fits to data. The planned-look versions CEVAM-NT and CEVAM-T select α_eff (and β) by grid search so that planned-look type I error is near the nominal α*, then break ties by power and EN0. That is ordinary design-stage operating-characteristic calibration of the same kind used by the BOP2/BOP2-FE comparators; it does not redefine the evidence process or force the reported EN advantage by construction. The simulation (Table 2) and TREND re-analysis simply evaluate the resulting pretabulated rules under the stated benchmarks. There is no self-definitional loop, no fitted input renamed as a prediction of a closely related quantity, no load-bearing self-citation uniqueness theorem, and no renaming of a known empirical pattern as a first-principles result. The reader's calibration-overfit concern is a fairness/generalizability issue for the efficiency comparison, not circularity of the derivation chain. Score 0 is therefore appropriate.
Assumptions & free parameters
free parameters (4)
- α_eff (efficacy evidence-threshold parameter)
- β (futility evidence-threshold parameter)
- design alternative p1 and null p0
- look schedule N and maximum N
assumptions (4)
- standard math Likelihood-ratio process ELR_n(p1,p0) is a nonnegative supermartingale under every p≤p0, so Ville’s inequality yields time-uniform type I error ≤α (Theorem 2.1).
- domain assumption Patient outcomes are i.i.d. Bernoulli(p) with fixed unknown p.
- domain assumption Clinically meaningful benchmarks satisfy 0<p0<p1<1 and are known at design time.
- ad hoc to paper Planned-look type I error control after grid-search calibration is an acceptable substitute for anytime-valid guarantees in conventional finite-look phase II trials.
invented entities (2)
-
CEVAM (calibrated e-value monitoring) framework and its NT/T variants
-
Reverse likelihood-ratio futility e-process F_LR_n(p0,p1)
Cite this review
Pith. "Pith review of Likelihood-Ratio E-Value Monitoring for Benchmark-Based Decisions in Early-Phase Oncology Trials." pith.science (2026). https://pith.science/paper/U4JMPDCC
@misc{pith2026260704571,
author = {Pith},
title = {Pith review of: Likelihood-Ratio E-Value Monitoring for Benchmark-Based Decisions in Early-Phase Oncology Trials},
year = {2026},
howpublished = {\url{https://pith.science/paper/U4JMPDCC}},
note = {Machine review of arXiv:2607.04571}
}
read the original abstract
Early-phase oncology trials often require protocol-ready interim rules for deciding whether an experimental regimen shows sufficient activity relative to prespecified clinical benchmarks. Existing Bayesian optimal phase II designs provide calibrated count boundaries for such decisions, but evidence is quantified through posterior probabilities and sample-size-dependent cutoff functions. We propose calibrated e-value monitoring (CEVAM), a likelihood-ratio evidence framework that defines monitoring evidence directly relative to prespecified clinical benchmarks, without requiring a Bayesian analysis prior or sample-size-dependent posterior-probability cutoff functions. For a binary efficacy endpoint, CEVAM constructs efficacy and reverse-futility e-processes targeted to clinically meaningful benchmark rates and converts them into monotone response-count boundaries. We distinguish a fixed-threshold e-process version, which has an anytime-valid interpretation under optional monitoring, from planned-look calibrated versions designed for conventional finite-look phase II trials. In simulations based on binary benchmark settings used in existing posterior-cutoff phase II designs, the proposed tuned planned-look rule controlled the nominal type I error and achieved the smallest expected sample size across all evaluated settings while retaining similar rejection probabilities. In an application to an actual phase II breast cancer trial, CEVAM classified the reported pathological complete response result as sufficient evidence for early success stopping. Extensions to categorical and multicomponent endpoints are provided in the Supplementary Material. CEVAM offers an analysis-prior-free, protocol-ready likelihood-ratio evidence scale for binary benchmark monitoring in early-phase oncology trials.
Figures
Reference graph
Works this paper leans on
-
[1]
Signal Transduction and Targeted Therapy , volume=
Efficacy, safety and exploratory analysis of neoadjuvant tislelizumab (a PD-1 inhibitor) plus nab-paclitaxel followed by epirubicin/cyclophosphamide for triple-negative breast cancer: a phase 2 TREND trial , author=. Signal Transduction and Targeted Therapy , volume=. 2025 , doi=
2025
-
[2]
Game-theoretic statistics and safe anytime-valid inference , journal =
Ramdas, Aaditya and Gr. Game-theoretic statistics and safe anytime-valid inference , journal =. 2023 , volume =
2023
-
[3]
2019 , isbn =
Shafer, Glenn and Vovk, Vladimir , title =. 2019 , isbn =
2019
-
[4]
Controlled Clinical Trials , year =
Simon, Richard , title =. Controlled Clinical Trials , year =
-
[5]
and Simon, Richard , title =
Thall, Peter F. and Simon, Richard , title =. Biometrics , year =
-
[6]
Statistics in Medicine , year =
Sambucini, Valeria , title =. Statistics in Medicine , year =
-
[7]
Statistics in Medicine , year =
Cai, Chunyan and Liu, Suyu and Yuan, Ying , title =. Statistics in Medicine , year =
-
[8]
Jack and Yuan, Ying , title =
Zhou, Heng and Lee, J. Jack and Yuan, Ying , title =. Statistics in Medicine , year =
Show all 11 references
-
[9]
and Takeda, Kentaro , title =
Xu, Xinling and Hashimoto, Atsuki and Yimer, Belay B. and Takeda, Kentaro , title =. Journal of Biopharmaceutical Statistics , year =. doi:10.1080/10543406.2025.2558142 , note =
2025 doi
-
[10]
The Annals of Statistics , year =
Vovk, Vladimir and Wang, Ruodu , title =. The Annals of Statistics , year =
-
[11]
Safe Testing , journal =
Gr. Safe Testing , journal =. 2024 , volume =
2024
Reviewed July 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.