Pith. sign in

REVIEW 2 major objections 4 minor 11 references

Likelihood-Ratio E-Value Monitoring for Benchmark-Based Decisions in Early-Phase Oncology Trials

T0 review · 2 major / 4 minor · reviewed 2026-07-11 · grok-4.5

Pith's one-line read A likelihood-ratio e-process rule for phase II oncology gives count boundaries without Bayesian priors and cuts expected sample size while matching type I error.

desk verdict Solid, practical packaging of LR e-processes into BOP2-style count tables; the EN win for CEVAM-T is real under the stated calibration but not a pure out-of-sample test of the evidence scale. read the letter →

arxiv 2607.04571 v1 pith:U4JMPDCC submitted 2026-07-06 stat.ME

classification stat.ME MSC 62L0562P10
keywords BayesianphaseIIdesigncomplexendpointsearlystoppinge-processesoncologytrialssequentialmonitoringlikelihood-ratioevidencebenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Early oncology trials need simple, protocol-ready rules that stop accrual early when a regimen is clearly better or worse than a clinical benchmark response rate. Existing Bayesian optimal designs already give such count tables, but they measure evidence with posterior probabilities that depend on an analysis prior and on sample-size-dependent cutoffs. This paper proposes CEVAM: form the ordinary likelihood ratio of a design alternative against the unacceptable rate (and the reverse ratio for futility), treat those ratios as e-processes, and convert them into monotone response-count boundaries that can be written down before the trial starts. A fixed-threshold version stays valid under optional monitoring; planned-look calibrated versions (especially the fully tuned CEVAM-T) are tuned only for conventional finite interim looks. In simulations that reuse the binary settings of the leading Bayesian designs, CEVAM-T controlled the nominal type I error and produced the smallest expected sample size at every true response rate while keeping rejection probabilities close to the Bayesian competitors. The same boundaries reclassified a published breast-cancer pathological-complete-response result as early success. The practical claim is that investigators can keep the familiar boundary-table workflow while replacing prior-dependent posterior cutoffs with an explicit likelihood-ratio evidence scale.

What carries the argument

The efficacy and reverse-futility likelihood-ratio e-processes ELR_n(p1,p0) and FLR_n(p0,p1), which are nonnegative supermartingales under their respective composite nulls and convert, by monotonicity, into explicit integer count boundaries reff(n;α) and rfut(n;β).

What would settle it

Re-run the same Monte Carlo comparison on a materially different look schedule or benchmark pair (for example denser early looks or p0=0.05, p1=0.20); if CEVAM-T no longer controls type I error near the nominal level or loses its expected-sample-size advantage while keeping similar power, the efficiency claim fails for that setting.

Watch

Extended reading notes

Core claim

For binary efficacy, the likelihood ratio of a design alternative p1 against the unacceptable rate p0 is an e-process under the composite null p ≤ p0; its reverse is an e-process under p ≥ p1. Both are monotone in the cumulative response count, so they translate into integer efficacy and futility boundaries that can be pretabulated. Planned-look calibration of the two evidence thresholds yields a design (CEVAM-T) that, on the standard binary benchmarks, matches type I error and rejection probability while achieving the smallest expected sample size of the designs compared.

Load-bearing premise

That a design-stage grid search over the two evidence thresholds, calibrated on one binary look schedule and one (p0,p1) pair, produces a fair efficiency comparison rather than an overfit to that particular configuration.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper proposes CEVAM, a likelihood-ratio e-value framework for protocol-ready interim monitoring of binary (and, in the supplement, complex) endpoints against prespecified clinical benchmarks in early-phase oncology trials. For binary efficacy, it constructs an efficacy LR e-process targeted at design alternative p1 versus the least-favorable null boundary p0, and a reverse LR futility process, both of which are nonnegative supermartingales under the respective composite nulls (Theorem 2.1). Monotonicity yields pretabulated integer response-count boundaries (Proposition 2.2). The authors distinguish a fixed-threshold anytime-valid version from planned-look calibrated versions (CEVAM-NT and CEVAM-T) that grid-search evidence thresholds to control planned-look type I error. In binary simulations matching published BOP2-FE settings (N=40, looks 10–40, p0=0.20, p1=0.40, α*=0.10), CEVAM-T controlled nominal type I error and achieved the smallest expected sample size across true response rates while retaining rejection probabilities close to BOP2-FE (Tables 1–2). A retrospective re-analysis of the TREND trial pCR result is used as illustration.

Significance. If the results hold, CEVAM supplies a practically useful, analysis-prior-free alternative to BOP2/BOP2-FE-style posterior-cutoff designs while preserving the protocol-ready count-boundary format that makes those designs attractive. The fixed-threshold construction rests on standard e-process/supermartingale theory and is independently valuable for optional or irregular monitoring; the pretabulated monotone boundaries and public R code further strengthen implementability. The efficiency claim for CEVAM-T is of direct interest to early-phase oncology design, even though it is obtained under ordinary design-stage calibration rather than a pure out-of-sample test of the evidence scale. The work is a solid methodological contribution at the interface of sequential evidence and phase II oncology practice.

major comments (2)
  1. Section 2.2 and 3.1 / Table 2: The central efficiency claim for CEVAM-T (smallest EN across all true response rates while retaining similar PRN) is obtained by grid-searching (α_eff, β) on the same look schedule and (p0, p1, N) used for evaluation, with selection by closeness of planned-look type I error to α* then power/EN0. This is ordinary for calibrated phase II designs (and analogous to BOP2/BOP2-FE), but the manuscript should more explicitly state that the ranking is not a pure out-of-sample test of the LR evidence scale and should report at least one sensitivity check under a different look schedule or (p0, p1) pair so that readers can judge robustness of the EN advantage.
  2. Section 3.2 / Figure 1 and Table 2: CEVAM-T (and CEVAM-NT) allocate a larger fraction of false-positive probability at early looks than the BOP2-FE comparators. While overall type I error is controlled at the final look, the manuscript should discuss whether this early alpha spending is clinically acceptable for oncology phase II (e.g., risk of early false success stopping) and whether the calibration criterion could be modified to constrain early spending if desired.
minor comments (4)
  1. Section 4 / Table 3: The TREND re-analysis uses n=44 efficacy-evaluable patients rather than the protocol-specified interim of 48; the text already notes this, but a one-sentence caveat that the illustration is not a reconstruction of original trial conduct would further reduce any risk of over-interpretation.
  2. Proposition 2.2 / Eqs. (8)–(9): The ceiling/floor expressions for reff and rfut are useful; a brief numerical check that they match the Table 1 boundaries for the simulation setting would help readers implement the formulas.
  3. Throughout: Occasional typesetting artifacts (e.g., “Gr¨ unwald”, “Heff 0”) and the arXiv-style line breaks should be cleaned for journal production.
  4. References: BOP2-FE is cited as advance online publication; ensure the final citation is updated if a volume/page version is available at revision.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: e-process validity and pretabulated boundaries are self-contained; calibration is ordinary design-stage OC search, not a prediction forced by construction.

full rationale

The paper's core derivation is standard and non-circular. Theorem 2.1 establishes that the likelihood-ratio process ELR_n(p1,p0) is a nonnegative supermartingale under the composite null p≤p0 (and the reverse process under p≥p1), yielding Ville-type bounds; this rests on classical martingale theory for Bernoulli likelihood ratios, not on any quantity defined from the target operating characteristics. Proposition 2.2 then converts the monotone LR statistics into explicit integer count boundaries reff(n;α) and rfut(n;β) by solving the inequalities for Rn; the formulas are algebraic rearrangements of the LR expressions, not fits to data. The planned-look versions CEVAM-NT and CEVAM-T select α_eff (and β) by grid search so that planned-look type I error is near the nominal α*, then break ties by power and EN0. That is ordinary design-stage operating-characteristic calibration of the same kind used by the BOP2/BOP2-FE comparators; it does not redefine the evidence process or force the reported EN advantage by construction. The simulation (Table 2) and TREND re-analysis simply evaluate the resulting pretabulated rules under the stated benchmarks. There is no self-definitional loop, no fitted input renamed as a prediction of a closely related quantity, no load-bearing self-citation uniqueness theorem, and no renaming of a known empirical pattern as a first-principles result. The reader's calibration-overfit concern is a fairness/generalizability issue for the efficiency comparison, not circularity of the derivation chain. Score 0 is therefore appropriate.

Assumptions & free parameters 4 free parameters · 4 assumptions · 2 invented entities

Central claims rest on classical supermartingale/Ville theory, Bernoulli i.i.d. sampling, and the clinical practice of fixing two benchmark rates (p0, p1). Free parameters are the usual design knobs plus the calibration grids; no new physical entities are postulated. The invented constructs are methodological (the CEVAM rules themselves).

free parameters (4)
  • α_eff (efficacy evidence-threshold parameter)
    Selected by grid search over A={0.01,...,1.00} to meet planned-look type I error; not fixed by theory.
  • β (futility evidence-threshold parameter)
    Fixed at 0.10 for CEVAM-NT or grid-searched over B for CEVAM-T; controls aggressiveness of reverse-futility stopping.
  • design alternative p1 and null p0
    Clinically chosen betting points that define the LR processes; operating characteristics depend on these choices.
  • look schedule N and maximum N
    Prespecified monitoring times used both for calibration and evaluation; efficiency claims are schedule-dependent.
assumptions (4)
  • standard math Likelihood-ratio process ELR_n(p1,p0) is a nonnegative supermartingale under every p≤p0, so Ville’s inequality yields time-uniform type I error ≤α (Theorem 2.1).
    Standard e-process / betting-martingale theory applied to Bernoulli LR; proof deferred to Supplementary Material.
  • domain assumption Patient outcomes are i.i.d. Bernoulli(p) with fixed unknown p.
    Stated at the opening of Section 2.1; required for the product LR and martingale property.
  • domain assumption Clinically meaningful benchmarks satisfy 0<p0<p1<1 and are known at design time.
    Required for the LR construction and for the interpretation of efficacy/futility claims.
  • ad hoc to paper Planned-look type I error control after grid-search calibration is an acceptable substitute for anytime-valid guarantees in conventional finite-look phase II trials.
    Explicitly distinguished in Section 2.2; the calibrated versions deliberately drop optional-monitoring validity.
invented entities (2)
  • CEVAM (calibrated e-value monitoring) framework and its NT/T variants
    purpose: Package LR e-processes into protocol-ready count boundaries for oncology benchmark monitoring without analysis priors.
    Methodological construct defined by the paper; independent evidence is the simulation and TREND illustration, not an external physical prediction.
  • Reverse likelihood-ratio futility e-process F_LR_n(p0,p1)
    purpose: Provide a symmetric evidence scale for early futility stopping against the target rate p1.
    Standard reverse LR, specialized here as a monitoring statistic with pretabulated count boundaries.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Likelihood-Ratio E-Value Monitoring for Benchmark-Based Decisions in Early-Phase Oncology Trials." pith.science (2026). https://pith.science/paper/U4JMPDCC

@misc{pith2026260704571,
  author       = {Pith},
  title        = {Pith review of: Likelihood-Ratio E-Value Monitoring for Benchmark-Based Decisions in Early-Phase Oncology Trials},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/U4JMPDCC}},
  note         = {Machine review of arXiv:2607.04571}
}
read the original abstract

Early-phase oncology trials often require protocol-ready interim rules for deciding whether an experimental regimen shows sufficient activity relative to prespecified clinical benchmarks. Existing Bayesian optimal phase II designs provide calibrated count boundaries for such decisions, but evidence is quantified through posterior probabilities and sample-size-dependent cutoff functions. We propose calibrated e-value monitoring (CEVAM), a likelihood-ratio evidence framework that defines monitoring evidence directly relative to prespecified clinical benchmarks, without requiring a Bayesian analysis prior or sample-size-dependent posterior-probability cutoff functions. For a binary efficacy endpoint, CEVAM constructs efficacy and reverse-futility e-processes targeted to clinically meaningful benchmark rates and converts them into monotone response-count boundaries. We distinguish a fixed-threshold e-process version, which has an anytime-valid interpretation under optional monitoring, from planned-look calibrated versions designed for conventional finite-look phase II trials. In simulations based on binary benchmark settings used in existing posterior-cutoff phase II designs, the proposed tuned planned-look rule controlled the nominal type I error and achieved the smallest expected sample size across all evaluated settings while retaining similar rejection probabilities. In an application to an actual phase II breast cancer trial, CEVAM classified the reported pathological complete response result as sufficient evidence for early success stopping. Extensions to categorical and multicomponent endpoints are provided in the Supplementary Material. CEVAM offers an analysis-prior-free, protocol-ready likelihood-ratio evidence scale for binary benchmark monitoring in early-phase oncology trials.

Figures

Figures reproduced from arXiv: 2607.04571 by the authors.

Figure 1
Figure 1. Descriptive cumulative false-positive probability curves under the binary design [PITH_FULL_IMAGE:figures/full_fig_p026_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

11 extracted references

  1. [1]

    Signal Transduction and Targeted Therapy , volume=

    Efficacy, safety and exploratory analysis of neoadjuvant tislelizumab (a PD-1 inhibitor) plus nab-paclitaxel followed by epirubicin/cyclophosphamide for triple-negative breast cancer: a phase 2 TREND trial , author=. Signal Transduction and Targeted Therapy , volume=. 2025 , doi=

  2. [2]

    Game-theoretic statistics and safe anytime-valid inference , journal =

    Ramdas, Aaditya and Gr. Game-theoretic statistics and safe anytime-valid inference , journal =. 2023 , volume =

  3. [3]

    2019 , isbn =

    Shafer, Glenn and Vovk, Vladimir , title =. 2019 , isbn =

  4. [4]

    Controlled Clinical Trials , year =

    Simon, Richard , title =. Controlled Clinical Trials , year =

  5. [5]

    and Simon, Richard , title =

    Thall, Peter F. and Simon, Richard , title =. Biometrics , year =

  6. [6]

    Statistics in Medicine , year =

    Sambucini, Valeria , title =. Statistics in Medicine , year =

  7. [7]

    Statistics in Medicine , year =

    Cai, Chunyan and Liu, Suyu and Yuan, Ying , title =. Statistics in Medicine , year =

  8. [8]

    Jack and Yuan, Ying , title =

    Zhou, Heng and Lee, J. Jack and Yuan, Ying , title =. Statistics in Medicine , year =

Show all 11 references
  1. [9]

    and Takeda, Kentaro , title =

    Xu, Xinling and Hashimoto, Atsuki and Yimer, Belay B. and Takeda, Kentaro , title =. Journal of Biopharmaceutical Statistics , year =. doi:10.1080/10543406.2025.2558142 , note =

  2. [10]

    The Annals of Statistics , year =

    Vovk, Vladimir and Wang, Ruodu , title =. The Annals of Statistics , year =

  3. [11]

    Safe Testing , journal =

    Gr. Safe Testing , journal =. 2024 , volume =

Pith tools

Reviewed July 11, 2026 · model on record in the stance chip above.