REVIEW 4 major objections 6 minor 47 references
The paper claims that any fixed-sample hypothesis test can be made valid under arbitrary stopping by monitoring the probability that it would reject at its planned end, with Type I error controlled at the original α and near-optimal power.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 23:23 UTC pith:7FXXXAKX
load-bearing objection The simple-null version of this paper is a real contribution; the composite and censored-data extensions are not backed by the same martingale argument. the 4 major comments →
Predicting fixed-sample test decisions enables anytime-valid inference
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
At n<N define Q_n = P_0(T_N^(n) ∈ C_{αγ} | F_n), where T_N^(n) is the fixed-sample statistic computed from the observed data plus future values imputed from the null, and C_{αγ} is a slightly tightened rejection region with level αγ. The rule stops and rejects when Q_n ≥ γ. Under the null, (Q_n) is a martingale, so the martingale maximal inequality bounds the probability of ever crossing γ by Q_0/γ = α. Thus any test of the form 'reject when T_N ∈ C_α' is adapted to an anytime-valid test by rejecting at stage n when Q_n ≥ γ. The authors argue this converts fixed-sample tests—parametric or nonparametric—into sequential tests without replacing them, matching fixed-sample power with about a 2%
What carries the argument
The predictive rejection probability Q_n — the null-probability that the completed fixed-sample test would reject at N given the data observed so far — is the central object. It is computed by treating future observations as missing data and imputing them from the null, either analytically or by Monte Carlo. Its martingale property under the null is the engine of error control: the maximal inequality converts a threshold crossing (Q_n ≥ γ) into a bound on the probability of any false rejection. The predictive critical region C_{αγ}, a level-αγ subset of the original region C_α, is the calibration device that makes the final Type I error equal to α rather than αγ.
Load-bearing premise
For the Type I error guarantee to hold, the distribution used to fill in the unobserved future data must be the true conditional distribution of those future data under the null; when the null is composite or nonparametric and the supplement substitutes plug-in estimates (S11.4.3) or imputes a death time for every unobserved patient (S10), that equality is an assumption, not a consequence of the theorem.
What would settle it
Generate data under a true null where the paper's imputation scheme is misspecified — for example, censored survival times with censoring dependent on covariates while imputing every future patient's outcome as a death time — run the sequential rule and count rejections before N in many simulations. If the empirical false-positive rate exceeds α, the claimed anytime-valid guarantee fails, locating the breakdown in the imputation model rather than in the martingale argument. A cleaner check: for any proposed imputation scheme, estimate E(Q_{n+1} | F_n) from simulations under the true null; if i
If this is right
- A clinical trial or A/B test can be monitored continuously, with no prespecified interim-analysis schedule, and still claim the same Type I error as the original fixed-sample design.
- Anytime-valid versions of nonparametric tests such as the two-sample Kolmogorov-Smirnov and log-rank tests follow from the same construction, without needing a likelihood ratio or a model for the alternative.
- The original fixed-sample test is not discarded; its statistic, critical region, and scientific interpretation remain, with power matched by a maximum sample size roughly 2% larger.
- Under the alternative, stopping typically occurs well before the planned end — mean stopping times around 384–388 for a 500-sample normal-mean design, and 7 months instead of 15 in the stroke-trial example — while the realized power can even exceed the fixed-sample power at the same effect size.
- An analogous prediction under the design alternative gives a safe futility rule: stop without rejecting when the updated probability of failing to detect the design effect exceeds a threshold; this cannot inflate Type I error.
Where Pith is reading between the lines
- Going beyond the paper: the construction is a specific instance of a more general recipe — take any decision rule that is a deterministic function of a complete dataset and make it sequential by predicting its final output under a null or reference imputation model. Hypothesis tests are the first application; point estimation, classification, or ranking rules could be converted the same way.
- Going beyond the paper: for composite nulls and nonparametric settings where the null does not fully determine the conditional law of future data, the Type I error guarantee is only as good as the chosen imputation scheme. The supplement itself substitutes plug-in predictive nulls and imputes a death time for every unobserved patient; a sensitivity analysis that varies the imputation scheme within
- Going beyond the paper: the reported power and sample-savings figures come from calibrated normal-theory designs and Monte Carlo runs; small-sample, heavy-tailed, or mis-specified settings are likely to show different trade-offs, so simulation studies over a range of null and alternative distributions would be needed before using the numbers as design guarantees.
- Going beyond the paper: the martingale property suggests a direct route to sequential confidence intervals and confidence sequences by inverting the test at every parameter value; the paper sketches one-sided and two-sided intervals for a normal mean, and this could extend to nonparametric quantities via the same predictive completion.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a 'predictive anytime-valid testing' procedure. For a fixed-sample test of sample size N with rejection region C_α, at each interim n<N it computes Q_n = P_0(T_N^{(n)} ∈ C_{αγ} | F_n), the probability under the null that the completed-sample test statistic will fall in a slightly tightened rejection region, with future observations imputed from a null-conditional distribution. The stopping rule rejects H_0 at the first n with Q_n ≥ γ, and the authors set the tightened level to αγ so that Doob's inequality yields a Type I error bound of α. The paper develops closed forms for the one-sided normal-mean test, gives simulation results on power and sample savings, applies the method to a two-sample log-rank analysis of the International Stroke Trial, and sketches extensions to confidence intervals, composite nulls, nonparametric tests, and futility stopping. The central claim is that any fixed-sample hypothesis test can be transformed into an anytime-valid test with near-optimal power and substantial sample savings under alternatives.
Significance. If the claimed generality were established, this would be a practically significant contribution: it would permit continuous, unplanned monitoring of a classical fixed-sample test while retaining Type I error control and avoiding the power losses associated with e-process methods. The simple-null version is an elegant and correct contribution: Proposition 1 (SM S3) and the Doob/Ville bound (SM S4) are rigorous, the Gaussian example admits an exact closed-form Q_n, and the Monte Carlo scheme for Q_n is straightforward and parallelizable. The confidence-interval extension in SM S7 is also attractive. However, the paper's headline claims—'any fixed-sample hypothesis test', 'broad applicability' including nonparametric and composite null settings—are not supported by the proofs provided. In the composite and nonparametric extensions the imputation distributions are not the true null-conditional laws required by Proposition 1; consequently the claimed Type I error control, especially for the International Stroke Trial analysis, is unproved. The paper should either substantially narrow its claims to fully specified nulls or supply rigorous justifications for the extended settings.
major comments (4)
- [SM S11.4.3, Eq. (3)] The composite-null extension defines a sequential null hypothesis H_0: X_{n+1} ~ f(·|θ̂_n), with plug-in MLEs. This is not the conditional law under the original composite null H_0: θ∈Θ_0. Proposition 1 (SM S3) requires imputation from the true null-conditional law. A martingale under this artificial plug-in predictive null does not imply the Doob inequality bound uniformly over θ∈Θ_0. The statement in S11.4.3 that 'the sequence (Q_n) is a martingale under the null hypothesis' is therefore unsupported, and the claimed Type I error control for composite nulls does not follow without additional argument.
- [SM S10, IST analysis (Figure S11)] The International Stroke Trial analysis uses the two-sample log-rank test and imputes future outcomes by sampling a death time for every patient whose event/censoring time has not yet been observed. Under the null, the conditional law of future data includes censoring; imputing all unobserved patients as deaths is not the predictive distribution defined in SM S2.1. The paper asserts that the IST sequential procedure 'can be rejected safely after 7 months' with Type I error 0.05, but no proof is given that the resulting Q_t process is a martingale or supermartingale under the null. The distribution-free property of the KS statistic invoked in S10 does not establish conditional invariance of the imputed-completion statistic given partial observation.
- [SM S11.4.1] For the composite null H_0: θ≤θ_0, the paper claims that (Q_n) is a supermartingale and justifies it by an inequality asserted to hold 'regardless of which one this is'. This is not generally true for arbitrary composite nulls; it relies on a stochastic-ordering/monotone-likelihood property that happens to hold in the normal one-sided example but is not established for 'any fixed-sample hypothesis test'. Without a general proof or explicitly stated conditions, the abstract's 'any' claim is an overgeneralization.
- [Main text 'Near-optimal power' and Table S1] The claims of 'near-optimal power' and a '~2% sample size increase' to match fixed-sample power are supported only for the one-sided normal-mean test with a particular design alternative. Table S1 reports only that setting; no simulations or analytical results are provided for other tests. The comparison with an e-process method (Figure 4) uses one specific e-process with a standard normal prior, which may not be representative. The power/sample-saving claims should be stated as empirical findings for these examples, not as general guarantees.
minor comments (6)
- [SM S4 and SM S2.4] The notation is confusing: 'αγ' is used both as the product α×γ and as a subscript in C_{αγ}; the relationship α̃ = αγ is introduced twice. Please define once and use consistently.
- [SM S10] The text states the critical value for the log-rank test is '3.84' but later says 'using the critical value of 3.93'. Please reconcile.
- [Figure 5 caption and SM S10] The caption says 'death including censoring times', while the text says 'death and censoring times'. Please use consistent terminology.
- [SM S6] Typo: 'fized' should be 'fixed'.
- [SM S4] Typo: 'gauranteeing' should be 'guaranteeing'.
- [SM S10 and S11] The terms P_0, F_n, and Q_n are used both for the original full-sample objects and for the sequentially conditioned versions; this is a source of confusion in the nonparametric sections. A more explicit subscripting scheme would help.
Circularity Check
Anytime-validity proof is self-contained (martingale + Ville inequality); no prediction reduces to a fit. Composite/nonparametric gaps are correctness risks, not circularity.
full rationale
The paper's central result is Proposition 1 (SM Section S3): under the null, Q_n = P_0(T_N^{(n)} in C | F_n) is a martingale when future data are imputed from the null-conditional law. The proof uses the tower property directly and is self-contained. Section S4 derives P_0(max Q_n >= gamma) <= Q_0/gamma = (alpha*gamma)/gamma = alpha by Ville/Doob, giving Type I error control without fitting any data. The threshold gamma is user-set; alpha*gamma is a design choice, not a fitted parameter. Power and sample-size claims (Section S5, Table S1, Figures 3-4) are standard normal-power calculations and Monte Carlo evaluations of the defined procedure, not retro-fitted predictions. The only self-citations (Fong, Holmes, Walker 2023, in SM Section S1) support a philosophical remark about predictive modeling and are not load-bearing for the theorem. The composite-null and survival extensions (SM Sections S10, S11.4.3) replace the true null conditional law with plug-in or ad hoc imputations (e.g., 'all individuals will be sampled to provide a death time'); this is an assumption-match or proof gap for those extensions rather than a circular reduction, because the martingale claim for the redefined null is still a forward derivation. No fitted parameter is renamed as a prediction; no external uniqueness theorem is invoked; no known result is simply renamed. Hence no significant circularity.
Axiom & Free-Parameter Ledger
free parameters (1)
- γ (rejection threshold) =
0.95 (default)
axioms (4)
- standard math Doob/Ville maximal inequality for nonnegative supermartingales
- domain assumption The imputation model P_0(X_{n+1:N}|X_{1:n}) is the true conditional law of the future observations under the null.
- ad hoc to paper Data-dependent predictive nulls (plug-in MLEs in Eq. (3), or N(θ̄_n, S_n^2) in S11.1) define a valid null for Type I error purposes.
- ad hoc to paper For the two-sample KS/log-rank procedures, imputing future values as uniform on (t,T), or as death times for all unobserved patients, gives the correct predictive distribution.
read the original abstract
Statistical hypothesis tests typically use prespecified sample sizes, yet data often arrive sequentially. Interim analyses invalidate classical error guarantees, while existing sequential methods require rigid testing preschedules or incur substantial losses in statistical power. We introduce a simple procedure that transforms any fixed-sample hypothesis test into an anytime-valid test while ensuring Type-I error control and near-optimal power with substantial sample savings when the null hypothesis is false. At each step, the procedure predicts the probability that a classical test would reject the null hypothesis at its fixed-sample size, treating future observations as missing data under the null hypothesis. Thresholding this probability yields an anytime-valid stopping rule. In areas such as clinical trials, stopping early and safely can ensure that subjects receive the best treatments and accelerate the development of effective therapies.
Figures
Reference graph
Works this paper leans on
-
[1]
David J. Aldous. Exchangeability and related topics. In \'Ecole d'\'et\'e de probabilit\'es de S aint- F lour, XIII ---1983 , volume 1117 of Lecture Notes in Math., pages 1--198. Springer, Berlin, 1985
1983
-
[2]
G.A. Barnard. Sequential tests in industrial statistics. Journal of the Royal Statistical Society, 8: 0 1--26, 1964
1964
-
[3]
J.M Bernardo and A.F.M. Smith. Bayesian Analysis. Wiley, 1994
1994
-
[4]
Bayesian adaptive methods for clinical trials
Scott M Berry, Bradley P Carlin, J Jack Lee, and Peter Muller. Bayesian adaptive methods for clinical trials. CRC press, 2010
2010
-
[5]
Probability and Measure
Patrick Billingsley. Probability and Measure. John Wiley & Sons, Inc., New York, third edition, 1995
1995
-
[6]
Ferguson distributions via P \'olya urn schemes
D Blackwell and J.B MacQueen. Ferguson distributions via P \'olya urn schemes. Annals of Mathematical Statistics, 1: 0 353--355, 1973
1973
-
[7]
La pr \'e vision: ses lois logiques, ses sources subjectives
Bruno de Finetti. La pr \'e vision: ses lois logiques, ses sources subjectives. In Annales de l'institut Henri Poincar \'e , volume 7, pages 1--68, 1937. [English translation in Studies in Subjective Probability (1980) (H. E. Kyburg and H. E. Smokler, eds.) 53-118. Krieger, Malabar, FL.]
1937
-
[8]
Gr \"u nwald
Rianne De Heide and Peter D. Gr \"u nwald. Why optional stopping can be a problem for B ayesians. Psychonomic Bulletin & Review, 28: 0 795--812, 2021
2021
-
[9]
J. L. Doob. Application of the theory of martingales. Actes du Colloque International Le Calcul des Probabilites et ses applications, Paris CNRS, pages 23--27, 1949 a
1949
-
[10]
J. L. Doob. Application of the theory of martingales. Actes du Colloque International Le Calcul des Probabilit\' e s et ses applications (Lyon, 28 Juin–3 Juillet 1948), Paris CNRS, 23–27 , 1949 b
1948
-
[11]
J.L. Doob. Stochastic Processes. J. Wiley & Sons, 1953
1953
-
[12]
Ferguson
T. Ferguson. A B ayesian analysis of some nonparametric problems. Annals of Mathematical Statistics, 1: 0 209--230, 1973
1973
-
[13]
R.A. Fisher. Statistical Methods for Researcher Workers. Oliver & Boyd, 1925
1925
-
[14]
E. Fong, C. Holmes, and S. G. Walker. Martingale posterior distributions. Journal of the Royal Statistical Society, Series B, 85: 0 1357--1391, 2023
2023
-
[15]
Alison L. Gibbs and Francis Edward Su. On choosing and bounding probability metrics. International Statistical Review, 70 0 (3): 0 419--435, 2002. doi:https://doi.org/10.1111/j.1751-5823.2002.tb00178.x
arXiv 2002
-
[16]
Gr \"u nwald, R
P. Gr \"u nwald, R. de Heide, and W. Koolen. Safe testing. Journal of the Royal Statistical Society, Series B, 86: 0 1091--1128, 2024
2024
-
[17]
Optional stopping with bayes factors: A categorization and extension of folklore results, with an application to invariant situations
A Hendriksen, R de Heide, and P Gr \"u nwald. Optional stopping with bayes factors: A categorization and extension of folklore results, with an application to invariant situations. Bayesian Analysis, 16 0 (3): 0 961--989, 2021
2021
-
[18]
Bruce M. Hill. Posterior distribution of percentiles: Bayes' theorem for sampling from a population. Journal of the American Statistical Association, 63 0 (322): 0 677--691, 1968
1968
-
[19]
Hirshleifer and J.G
J. Hirshleifer and J.G. Riley. The Analytics of Uncertainty and Information. Cambridge University Press, 2012
2012
-
[20]
M.M. Hoeper, D.B. Badesch, H. Ardeschir, and Stellar Trial Investigators. Phase 3 trial of sotatercept for treatment of pulmonary arterial hypertension. The New England Journal of Medicine, 388 0 (16): 0 1478--1490, 2023. doi:10.1056/NEJMoa2213558. URL https://www.nejm.org/doi/10.1056/NEJMoa2213558
-
[21]
Howard, A
S.R. Howard, A. Ramdas, J. McAuliffe, and J. Sekhon. Time-uniform, nonparametric, nonasymptotic confidence sequences. Annals of Statistics, 49: 0 1055--1080, 2021
2021
-
[22]
Jeffreys
H. Jeffreys. Scientific Inference. Cambridge University Press, 1931
1931
-
[23]
Jennison and B.W
C. Jennison and B.W. Turnbull. Interim analyses: the repeated confidence interval approach. Journal of the Royal Statistical Society, Series B, 51: 0 305--361, 1989
1989
-
[24]
Jennison and B.W
C. Jennison and B.W. Turnbull. Group sequential methods with application to clinical trials. CRC Press, 1999
1999
-
[25]
Johnstone and B
C. Johnstone and B. Cox. Conformal uncertainty sets for robust optimization. Proceedings of Machine Learning Research, 152: 0 1--19, 2021
2021
-
[26]
Kass and A.E
R.E. Kass and A.E. Raftery. Bayes factors. Journal of the American Statistical Association, 90: 0 773--795, 1995
1995
-
[27]
Lewis and H.A
R.J. Lewis and H.A. Bessen. Sequential clinical trials in emergency medicine. Annals of Emergency Medicine, 19: 0 1047--1053, 1990
1990
-
[28]
Neyman and E
J. Neyman and E. S. Pearson. On the problem of the most efficient tests of statistical hypotheses. Philosophical Transactions of the Royal Society A, 231: 0 289–337, 1933
1933
-
[29]
O'Brien and T.R
P.C. O'Brien and T.R. Fleming. A multiple testing procedure for clinical trials. Biometrics, 35: 0 549--556, 1979
1979
-
[30]
Papadopoulos, K
H. Papadopoulos, K. Proedrou, V. Vovk, and A. Gammerman. Inductive confidence machines for regression. In Machine Learning: European Conference on Machine Learning, pages 345--356, 2002
2002
-
[31]
K. Pearson. On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling. The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science, 5: 0 157–175, 1900
1900
-
[32]
S.J. Pocock. Group sequential methods in the design and analysis of clinical trials. Biometrika, 64: 0 191--199, 1977
1977
-
[33]
Ramdas, P
A. Ramdas, P. Gr \"u nwald, V. Vovk, and G. Shafer. Game-theoretic statistics and safe anytime-valid inference. Statistical Science, 38: 0 576--601, 2023 a
2023
-
[34]
Game-theoretic statistics and safe anytime-valid inference
A Ramdas, P Gr \"u nwald, V Vovk, and G Shafer. Game-theoretic statistics and safe anytime-valid inference. Statistical Science, 38 0 (4): 0 576--601, 2023 b
2023
-
[35]
H. Robbins. Statistical methods related to the law of the iterated logarithm. The Annals of Mathematical Statistics, 41: 0 1397--1409, 1970
1970
-
[36]
Robbins and D
H. Robbins and D. Siegmund. A convergence theorem for non negative almost supermartingales and some applications. In Jagdish S. Rustagi, editor, Optimizing Methods in Statistics, pages 233--257. Academic Press, 1971
1971
-
[37]
D.B. Rubin. The B ayesian bootstrap. Annals of Statistics, 9: 0 130--134, 1981
1981
-
[38]
Sandercock, M
P.A.G. Sandercock, M. Niewada, A. Czonkowska, and the international stroke trial collaborative group. The international stroke trial database. Trial, 12, 2011
2011
-
[39]
Schultzberg and S
M. Schultzberg and S. Ankargren. Choosing a sequential testing framework — comparisons and discussions. Spotify, 2023
2023
-
[40]
G. Shafer. Testing by betting: A strategy for statistical and scientific communcation. Journal of the Royal Statistical Society, Series A, 184: 0 407--431, 2021
2021
-
[41]
Shafer, A
G. Shafer, A. Shen, N. Vereshchagin, and V. Vovk. Test martingales, bayes factors and p values. Statistical Science, 26: 0 84--101, 2011
2011
-
[42]
Spiegelhalter, L.S
D.J. Spiegelhalter, L.S. Freedman, and M.K.B. Parmar. Bayesian approaches to randomized trials. Journal of the Royal Statistical Society, Series A, 157: 0 357--416, 1994
1994
-
[43]
J. Ville. Etude Critique de la Notion de Collectif. PhD, Paris, 1939
1939
-
[44]
Vovk and R
V. Vovk and R. Wang. E-values: Calibration, combination and applications. Annals of Statistics, 49: 0 1736--1754, 2021
2021
-
[45]
A. Wald. Sequential Analyis. Wiley & Sons, NY, 1947
1947
-
[46]
Wason and A.P Mander
J.M.S. Wason and A.P Mander. Minimizing the maximum expected sample size im two stage phase ii clinical trials with continuous outcomes. Journal of Biopharmaceutical Statistics, 22: 0 836--852, 2012
2012
-
[47]
Wasserman, A
L. Wasserman, A. Ramdas, and S. Balakrishnan. Universal inference. Proceedings of the National Academy of Sciences, 117: 0 16880--16890, 2020
2020
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.