Pith. sign in

REVIEW 4 major objections 6 minor 47 references

The paper claims that any fixed-sample hypothesis test can be made valid under arbitrary stopping by monitoring the probability that it would reject at its planned end, with Type I error controlled at the original α and near-optimal power.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 23:23 UTC pith:7FXXXAKX

load-bearing objection The simple-null version of this paper is a real contribution; the composite and censored-data extensions are not backed by the same martingale argument. the 4 major comments →

arxiv 2602.13872 v2 pith:7FXXXAKX submitted 2026-02-14 stat.ME stat.ML

Predicting fixed-sample test decisions enables anytime-valid inference

classification stat.ME stat.ML MSC 62L1062F03
keywords anytime-valid inferencesequential hypothesis testingpredictive rejection probabilitymartingaleoptional stoppingfixed-sample testclinical trialsnonparametric tests
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Classical hypothesis tests guarantee their false-positive rate only when the decision is made at a preplanned sample size; peeking at interim data breaks the guarantee. This paper claims the trade-off between flexible stopping and statistical efficiency is not fundamental. It constructs, for any fixed-sample test, a sequential rule: at each interim sample size compute the probability that the test would reject at the planned end if the unobserved future observations were drawn from the null distribution, and reject as soon as that probability crosses γ. Because this predicted rejection probability is a martingale under a correctly specified null, thresholding it keeps the overall Type I error at or below the original α, while the original test statistic and its interpretation are preserved rather than replaced by a likelihood-ratio or betting score. A sympathetic reader would care because the method promises safe continuous monitoring in clinical trials and streaming experiments with near-optimal power and meaningful sample savings.

Core claim

At n<N define Q_n = P_0(T_N^(n) ∈ C_{αγ} | F_n), where T_N^(n) is the fixed-sample statistic computed from the observed data plus future values imputed from the null, and C_{αγ} is a slightly tightened rejection region with level αγ. The rule stops and rejects when Q_n ≥ γ. Under the null, (Q_n) is a martingale, so the martingale maximal inequality bounds the probability of ever crossing γ by Q_0/γ = α. Thus any test of the form 'reject when T_N ∈ C_α' is adapted to an anytime-valid test by rejecting at stage n when Q_n ≥ γ. The authors argue this converts fixed-sample tests—parametric or nonparametric—into sequential tests without replacing them, matching fixed-sample power with about a 2%

What carries the argument

The predictive rejection probability Q_n — the null-probability that the completed fixed-sample test would reject at N given the data observed so far — is the central object. It is computed by treating future observations as missing data and imputing them from the null, either analytically or by Monte Carlo. Its martingale property under the null is the engine of error control: the maximal inequality converts a threshold crossing (Q_n ≥ γ) into a bound on the probability of any false rejection. The predictive critical region C_{αγ}, a level-αγ subset of the original region C_α, is the calibration device that makes the final Type I error equal to α rather than αγ.

Load-bearing premise

For the Type I error guarantee to hold, the distribution used to fill in the unobserved future data must be the true conditional distribution of those future data under the null; when the null is composite or nonparametric and the supplement substitutes plug-in estimates (S11.4.3) or imputes a death time for every unobserved patient (S10), that equality is an assumption, not a consequence of the theorem.

What would settle it

Generate data under a true null where the paper's imputation scheme is misspecified — for example, censored survival times with censoring dependent on covariates while imputing every future patient's outcome as a death time — run the sequential rule and count rejections before N in many simulations. If the empirical false-positive rate exceeds α, the claimed anytime-valid guarantee fails, locating the breakdown in the imputation model rather than in the martingale argument. A cleaner check: for any proposed imputation scheme, estimate E(Q_{n+1} | F_n) from simulations under the true null; if i

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A clinical trial or A/B test can be monitored continuously, with no prespecified interim-analysis schedule, and still claim the same Type I error as the original fixed-sample design.
  • Anytime-valid versions of nonparametric tests such as the two-sample Kolmogorov-Smirnov and log-rank tests follow from the same construction, without needing a likelihood ratio or a model for the alternative.
  • The original fixed-sample test is not discarded; its statistic, critical region, and scientific interpretation remain, with power matched by a maximum sample size roughly 2% larger.
  • Under the alternative, stopping typically occurs well before the planned end — mean stopping times around 384–388 for a 500-sample normal-mean design, and 7 months instead of 15 in the stroke-trial example — while the realized power can even exceed the fixed-sample power at the same effect size.
  • An analogous prediction under the design alternative gives a safe futility rule: stop without rejecting when the updated probability of failing to detect the design effect exceeds a threshold; this cannot inflate Type I error.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Going beyond the paper: the construction is a specific instance of a more general recipe — take any decision rule that is a deterministic function of a complete dataset and make it sequential by predicting its final output under a null or reference imputation model. Hypothesis tests are the first application; point estimation, classification, or ranking rules could be converted the same way.
  • Going beyond the paper: for composite nulls and nonparametric settings where the null does not fully determine the conditional law of future data, the Type I error guarantee is only as good as the chosen imputation scheme. The supplement itself substitutes plug-in predictive nulls and imputes a death time for every unobserved patient; a sensitivity analysis that varies the imputation scheme within
  • Going beyond the paper: the reported power and sample-savings figures come from calibrated normal-theory designs and Monte Carlo runs; small-sample, heavy-tailed, or mis-specified settings are likely to show different trade-offs, so simulation studies over a range of null and alternative distributions would be needed before using the numbers as design guarantees.
  • Going beyond the paper: the martingale property suggests a direct route to sequential confidence intervals and confidence sequences by inverting the test at every parameter value; the paper sketches one-sided and two-sided intervals for a normal mean, and this could extend to nonparametric quantities via the same predictive completion.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes a 'predictive anytime-valid testing' procedure. For a fixed-sample test of sample size N with rejection region C_α, at each interim n<N it computes Q_n = P_0(T_N^{(n)} ∈ C_{αγ} | F_n), the probability under the null that the completed-sample test statistic will fall in a slightly tightened rejection region, with future observations imputed from a null-conditional distribution. The stopping rule rejects H_0 at the first n with Q_n ≥ γ, and the authors set the tightened level to αγ so that Doob's inequality yields a Type I error bound of α. The paper develops closed forms for the one-sided normal-mean test, gives simulation results on power and sample savings, applies the method to a two-sample log-rank analysis of the International Stroke Trial, and sketches extensions to confidence intervals, composite nulls, nonparametric tests, and futility stopping. The central claim is that any fixed-sample hypothesis test can be transformed into an anytime-valid test with near-optimal power and substantial sample savings under alternatives.

Significance. If the claimed generality were established, this would be a practically significant contribution: it would permit continuous, unplanned monitoring of a classical fixed-sample test while retaining Type I error control and avoiding the power losses associated with e-process methods. The simple-null version is an elegant and correct contribution: Proposition 1 (SM S3) and the Doob/Ville bound (SM S4) are rigorous, the Gaussian example admits an exact closed-form Q_n, and the Monte Carlo scheme for Q_n is straightforward and parallelizable. The confidence-interval extension in SM S7 is also attractive. However, the paper's headline claims—'any fixed-sample hypothesis test', 'broad applicability' including nonparametric and composite null settings—are not supported by the proofs provided. In the composite and nonparametric extensions the imputation distributions are not the true null-conditional laws required by Proposition 1; consequently the claimed Type I error control, especially for the International Stroke Trial analysis, is unproved. The paper should either substantially narrow its claims to fully specified nulls or supply rigorous justifications for the extended settings.

major comments (4)
  1. [SM S11.4.3, Eq. (3)] The composite-null extension defines a sequential null hypothesis H_0: X_{n+1} ~ f(·|θ̂_n), with plug-in MLEs. This is not the conditional law under the original composite null H_0: θ∈Θ_0. Proposition 1 (SM S3) requires imputation from the true null-conditional law. A martingale under this artificial plug-in predictive null does not imply the Doob inequality bound uniformly over θ∈Θ_0. The statement in S11.4.3 that 'the sequence (Q_n) is a martingale under the null hypothesis' is therefore unsupported, and the claimed Type I error control for composite nulls does not follow without additional argument.
  2. [SM S10, IST analysis (Figure S11)] The International Stroke Trial analysis uses the two-sample log-rank test and imputes future outcomes by sampling a death time for every patient whose event/censoring time has not yet been observed. Under the null, the conditional law of future data includes censoring; imputing all unobserved patients as deaths is not the predictive distribution defined in SM S2.1. The paper asserts that the IST sequential procedure 'can be rejected safely after 7 months' with Type I error 0.05, but no proof is given that the resulting Q_t process is a martingale or supermartingale under the null. The distribution-free property of the KS statistic invoked in S10 does not establish conditional invariance of the imputed-completion statistic given partial observation.
  3. [SM S11.4.1] For the composite null H_0: θ≤θ_0, the paper claims that (Q_n) is a supermartingale and justifies it by an inequality asserted to hold 'regardless of which one this is'. This is not generally true for arbitrary composite nulls; it relies on a stochastic-ordering/monotone-likelihood property that happens to hold in the normal one-sided example but is not established for 'any fixed-sample hypothesis test'. Without a general proof or explicitly stated conditions, the abstract's 'any' claim is an overgeneralization.
  4. [Main text 'Near-optimal power' and Table S1] The claims of 'near-optimal power' and a '~2% sample size increase' to match fixed-sample power are supported only for the one-sided normal-mean test with a particular design alternative. Table S1 reports only that setting; no simulations or analytical results are provided for other tests. The comparison with an e-process method (Figure 4) uses one specific e-process with a standard normal prior, which may not be representative. The power/sample-saving claims should be stated as empirical findings for these examples, not as general guarantees.
minor comments (6)
  1. [SM S4 and SM S2.4] The notation is confusing: 'αγ' is used both as the product α×γ and as a subscript in C_{αγ}; the relationship α̃ = αγ is introduced twice. Please define once and use consistently.
  2. [SM S10] The text states the critical value for the log-rank test is '3.84' but later says 'using the critical value of 3.93'. Please reconcile.
  3. [Figure 5 caption and SM S10] The caption says 'death including censoring times', while the text says 'death and censoring times'. Please use consistent terminology.
  4. [SM S6] Typo: 'fized' should be 'fixed'.
  5. [SM S4] Typo: 'gauranteeing' should be 'guaranteeing'.
  6. [SM S10 and S11] The terms P_0, F_n, and Q_n are used both for the original full-sample objects and for the sequentially conditioned versions; this is a source of confusion in the nonparametric sections. A more explicit subscripting scheme would help.

Circularity Check

0 steps flagged

Anytime-validity proof is self-contained (martingale + Ville inequality); no prediction reduces to a fit. Composite/nonparametric gaps are correctness risks, not circularity.

full rationale

The paper's central result is Proposition 1 (SM Section S3): under the null, Q_n = P_0(T_N^{(n)} in C | F_n) is a martingale when future data are imputed from the null-conditional law. The proof uses the tower property directly and is self-contained. Section S4 derives P_0(max Q_n >= gamma) <= Q_0/gamma = (alpha*gamma)/gamma = alpha by Ville/Doob, giving Type I error control without fitting any data. The threshold gamma is user-set; alpha*gamma is a design choice, not a fitted parameter. Power and sample-size claims (Section S5, Table S1, Figures 3-4) are standard normal-power calculations and Monte Carlo evaluations of the defined procedure, not retro-fitted predictions. The only self-citations (Fong, Holmes, Walker 2023, in SM Section S1) support a philosophical remark about predictive modeling and are not load-bearing for the theorem. The composite-null and survival extensions (SM Sections S10, S11.4.3) replace the true null conditional law with plug-in or ad hoc imputations (e.g., 'all individuals will be sampled to provide a death time'); this is an assumption-match or proof gap for those extensions rather than a circular reduction, because the martingale claim for the redefined null is still a forward derivation. No fitted parameter is renamed as a prediction; no external uniqueness theorem is invoked; no known result is simply renamed. Hence no significant circularity.

Axiom & Free-Parameter Ledger

1 free parameters · 4 axioms · 0 invented entities

The construction rests on one clean standard assumption (a martingale inequality) and one strong domain assumption (the null gives a valid conditional imputation law for the missing future data). The composite-null and censored-data extensions add ad hoc predictive nulls and imputation schemes that are not proven to preserve the error guarantee.

free parameters (1)
  • γ (rejection threshold) = 0.95 (default)
    User-set threshold in the stopping rule τ = min{n: Q_n ≥ γ}; controls the conservatism/power trade-off. Not fitted to data, but it is a hand-chosen procedural parameter.
axioms (4)
  • standard math Doob/Ville maximal inequality for nonnegative supermartingales
    Used in Section S4 to bound P(max Q_n ≥ γ) by Q_0/γ.
  • domain assumption The imputation model P_0(X_{n+1:N}|X_{1:n}) is the true conditional law of the future observations under the null.
    Invoked in Definition S1 and Proposition 1; for i.i.d. simple nulls this is standard, but for composite/nonparametric nulls the paper substitutes constructed predictive distributions.
  • ad hoc to paper Data-dependent predictive nulls (plug-in MLEs in Eq. (3), or N(θ̄_n, S_n^2) in S11.1) define a valid null for Type I error purposes.
    S11.4.1 asserts 'the expectation is upper bounded by Q_n' without proof; the claimed supermartingale property over H_0: θ≤θ_0 is not demonstrated.
  • ad hoc to paper For the two-sample KS/log-rank procedures, imputing future values as uniform on (t,T), or as death times for all unobserved patients, gives the correct predictive distribution.
    S8.2 and S10 rely on invariance properties of the KS statistic; for the censored log-rank data, imputing death times for censored individuals does not condition on the censoring times, so the distributional claim is unsupported.

pith-pipeline@v1.3.0-alltime-deepseek · 29609 in / 19267 out tokens · 159896 ms · 2026-08-02T23:23:18.501899+00:00 · methodology

0 comments
read the original abstract

Statistical hypothesis tests typically use prespecified sample sizes, yet data often arrive sequentially. Interim analyses invalidate classical error guarantees, while existing sequential methods require rigid testing preschedules or incur substantial losses in statistical power. We introduce a simple procedure that transforms any fixed-sample hypothesis test into an anytime-valid test while ensuring Type-I error control and near-optimal power with substantial sample savings when the null hypothesis is false. At each step, the procedure predicts the probability that a classical test would reject the null hypothesis at its fixed-sample size, treating future observations as missing data under the null hypothesis. Thresholding this probability yields an anytime-valid stopping rule. In areas such as clinical trials, stopping early and safely can ensure that subjects receive the best treatments and accelerate the development of effective therapies.

Figures

Figures reproduced from arXiv: 2602.13872 by Chris Holmes, Stephen Walker.

Figure 1
Figure 1. Figure 1: Overview of the predictive procedure. Panel A: the classical fixed-sample test precludes early analysis. Panel B: the uncertainty in the outcome of the fixed-sample test at n < N is driven by the unobserved data Xn+1:N . Panel C: we can characterise the uncertainty in the test decision at N, assuming the null hypothesis to be true, by simulating the missing data under the null hypothesis. Repeated testing … view at source ↗
Figure 2
Figure 2. Figure 2: Predicting the fixed-sample decision enables anytime-valid testing. Shown is the evolution of the predicted rejection probability Qn, defined as the probability that the fixed￾sample test would reject at sample size N, conditional on data observed up to stage n and assuming the null hypothesis is true. The experiment is testing H0 : θ = 0 versus H1 : θ > 0 for a normal mean θ with variance assumed known at… view at source ↗
Figure 3
Figure 3. Figure 3: Distribution of early stopping times. Shown is the cumulative distribution of the early stopping times of the experiment of the same type as in [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Near-optimal power with minimal sample inflation. Power as a function of effect size is shown for a classical fixed-sample test (dotted line), the predictive anytime-valid test (bold line), and a representative anytime-valid likelihood-ratio–based method (dashed line); the latter two with a 2% sample size increase. The predictive test closely tracks the power of the fixed-sample test across effect sizes, r… view at source ↗
Figure 5
Figure 5. Figure 5: Predictive anytime-valid analysis of a clinical trial. Predicted rejection probability Qt for the International Stroke Trial (IST), comparing death including censoring times for patients assigned to aspirin for the first 14 days of the trial versus not assigned to aspirin. Here the t are represented by a discrete set of time points which are set over 15 months following the start of the trial. Full details… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

47 extracted references · 1 canonical work pages

  1. [1]

    David J. Aldous. Exchangeability and related topics. In \'Ecole d'\'et\'e de probabilit\'es de S aint- F lour, XIII ---1983 , volume 1117 of Lecture Notes in Math., pages 1--198. Springer, Berlin, 1985

  2. [2]

    G.A. Barnard. Sequential tests in industrial statistics. Journal of the Royal Statistical Society, 8: 0 1--26, 1964

  3. [3]

    J.M Bernardo and A.F.M. Smith. Bayesian Analysis. Wiley, 1994

  4. [4]

    Bayesian adaptive methods for clinical trials

    Scott M Berry, Bradley P Carlin, J Jack Lee, and Peter Muller. Bayesian adaptive methods for clinical trials. CRC press, 2010

  5. [5]

    Probability and Measure

    Patrick Billingsley. Probability and Measure. John Wiley & Sons, Inc., New York, third edition, 1995

  6. [6]

    Ferguson distributions via P \'olya urn schemes

    D Blackwell and J.B MacQueen. Ferguson distributions via P \'olya urn schemes. Annals of Mathematical Statistics, 1: 0 353--355, 1973

  7. [7]

    La pr \'e vision: ses lois logiques, ses sources subjectives

    Bruno de Finetti. La pr \'e vision: ses lois logiques, ses sources subjectives. In Annales de l'institut Henri Poincar \'e , volume 7, pages 1--68, 1937. [English translation in Studies in Subjective Probability (1980) (H. E. Kyburg and H. E. Smokler, eds.) 53-118. Krieger, Malabar, FL.]

  8. [8]

    Gr \"u nwald

    Rianne De Heide and Peter D. Gr \"u nwald. Why optional stopping can be a problem for B ayesians. Psychonomic Bulletin & Review, 28: 0 795--812, 2021

  9. [9]

    J. L. Doob. Application of the theory of martingales. Actes du Colloque International Le Calcul des Probabilites et ses applications, Paris CNRS, pages 23--27, 1949 a

  10. [10]

    J. L. Doob. Application of the theory of martingales. Actes du Colloque International Le Calcul des Probabilit\' e s et ses applications (Lyon, 28 Juin–3 Juillet 1948), Paris CNRS, 23–27 , 1949 b

  11. [11]

    J.L. Doob. Stochastic Processes. J. Wiley & Sons, 1953

  12. [12]

    Ferguson

    T. Ferguson. A B ayesian analysis of some nonparametric problems. Annals of Mathematical Statistics, 1: 0 209--230, 1973

  13. [13]

    R.A. Fisher. Statistical Methods for Researcher Workers. Oliver & Boyd, 1925

  14. [14]

    E. Fong, C. Holmes, and S. G. Walker. Martingale posterior distributions. Journal of the Royal Statistical Society, Series B, 85: 0 1357--1391, 2023

  15. [15]

    Gibbs and Francis Edward Su

    Alison L. Gibbs and Francis Edward Su. On choosing and bounding probability metrics. International Statistical Review, 70 0 (3): 0 419--435, 2002. doi:https://doi.org/10.1111/j.1751-5823.2002.tb00178.x

  16. [16]

    Gr \"u nwald, R

    P. Gr \"u nwald, R. de Heide, and W. Koolen. Safe testing. Journal of the Royal Statistical Society, Series B, 86: 0 1091--1128, 2024

  17. [17]

    Optional stopping with bayes factors: A categorization and extension of folklore results, with an application to invariant situations

    A Hendriksen, R de Heide, and P Gr \"u nwald. Optional stopping with bayes factors: A categorization and extension of folklore results, with an application to invariant situations. Bayesian Analysis, 16 0 (3): 0 961--989, 2021

  18. [18]

    Bruce M. Hill. Posterior distribution of percentiles: Bayes' theorem for sampling from a population. Journal of the American Statistical Association, 63 0 (322): 0 677--691, 1968

  19. [19]

    Hirshleifer and J.G

    J. Hirshleifer and J.G. Riley. The Analytics of Uncertainty and Information. Cambridge University Press, 2012

  20. [20]

    Hoeper, D.B

    M.M. Hoeper, D.B. Badesch, H. Ardeschir, and Stellar Trial Investigators. Phase 3 trial of sotatercept for treatment of pulmonary arterial hypertension. The New England Journal of Medicine, 388 0 (16): 0 1478--1490, 2023. doi:10.1056/NEJMoa2213558. URL https://www.nejm.org/doi/10.1056/NEJMoa2213558

  21. [21]

    Howard, A

    S.R. Howard, A. Ramdas, J. McAuliffe, and J. Sekhon. Time-uniform, nonparametric, nonasymptotic confidence sequences. Annals of Statistics, 49: 0 1055--1080, 2021

  22. [22]

    Jeffreys

    H. Jeffreys. Scientific Inference. Cambridge University Press, 1931

  23. [23]

    Jennison and B.W

    C. Jennison and B.W. Turnbull. Interim analyses: the repeated confidence interval approach. Journal of the Royal Statistical Society, Series B, 51: 0 305--361, 1989

  24. [24]

    Jennison and B.W

    C. Jennison and B.W. Turnbull. Group sequential methods with application to clinical trials. CRC Press, 1999

  25. [25]

    Johnstone and B

    C. Johnstone and B. Cox. Conformal uncertainty sets for robust optimization. Proceedings of Machine Learning Research, 152: 0 1--19, 2021

  26. [26]

    Kass and A.E

    R.E. Kass and A.E. Raftery. Bayes factors. Journal of the American Statistical Association, 90: 0 773--795, 1995

  27. [27]

    Lewis and H.A

    R.J. Lewis and H.A. Bessen. Sequential clinical trials in emergency medicine. Annals of Emergency Medicine, 19: 0 1047--1053, 1990

  28. [28]

    Neyman and E

    J. Neyman and E. S. Pearson. On the problem of the most efficient tests of statistical hypotheses. Philosophical Transactions of the Royal Society A, 231: 0 289–337, 1933

  29. [29]

    O'Brien and T.R

    P.C. O'Brien and T.R. Fleming. A multiple testing procedure for clinical trials. Biometrics, 35: 0 549--556, 1979

  30. [30]

    Papadopoulos, K

    H. Papadopoulos, K. Proedrou, V. Vovk, and A. Gammerman. Inductive confidence machines for regression. In Machine Learning: European Conference on Machine Learning, pages 345--356, 2002

  31. [31]

    K. Pearson. On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling. The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science, 5: 0 157–175, 1900

  32. [32]

    S.J. Pocock. Group sequential methods in the design and analysis of clinical trials. Biometrika, 64: 0 191--199, 1977

  33. [33]

    Ramdas, P

    A. Ramdas, P. Gr \"u nwald, V. Vovk, and G. Shafer. Game-theoretic statistics and safe anytime-valid inference. Statistical Science, 38: 0 576--601, 2023 a

  34. [34]

    Game-theoretic statistics and safe anytime-valid inference

    A Ramdas, P Gr \"u nwald, V Vovk, and G Shafer. Game-theoretic statistics and safe anytime-valid inference. Statistical Science, 38 0 (4): 0 576--601, 2023 b

  35. [35]

    H. Robbins. Statistical methods related to the law of the iterated logarithm. The Annals of Mathematical Statistics, 41: 0 1397--1409, 1970

  36. [36]

    Robbins and D

    H. Robbins and D. Siegmund. A convergence theorem for non negative almost supermartingales and some applications. In Jagdish S. Rustagi, editor, Optimizing Methods in Statistics, pages 233--257. Academic Press, 1971

  37. [37]

    D.B. Rubin. The B ayesian bootstrap. Annals of Statistics, 9: 0 130--134, 1981

  38. [38]

    Sandercock, M

    P.A.G. Sandercock, M. Niewada, A. Czonkowska, and the international stroke trial collaborative group. The international stroke trial database. Trial, 12, 2011

  39. [39]

    Schultzberg and S

    M. Schultzberg and S. Ankargren. Choosing a sequential testing framework — comparisons and discussions. Spotify, 2023

  40. [40]

    G. Shafer. Testing by betting: A strategy for statistical and scientific communcation. Journal of the Royal Statistical Society, Series A, 184: 0 407--431, 2021

  41. [41]

    Shafer, A

    G. Shafer, A. Shen, N. Vereshchagin, and V. Vovk. Test martingales, bayes factors and p values. Statistical Science, 26: 0 84--101, 2011

  42. [42]

    Spiegelhalter, L.S

    D.J. Spiegelhalter, L.S. Freedman, and M.K.B. Parmar. Bayesian approaches to randomized trials. Journal of the Royal Statistical Society, Series A, 157: 0 357--416, 1994

  43. [43]

    J. Ville. Etude Critique de la Notion de Collectif. PhD, Paris, 1939

  44. [44]

    Vovk and R

    V. Vovk and R. Wang. E-values: Calibration, combination and applications. Annals of Statistics, 49: 0 1736--1754, 2021

  45. [45]

    A. Wald. Sequential Analyis. Wiley & Sons, NY, 1947

  46. [46]

    Wason and A.P Mander

    J.M.S. Wason and A.P Mander. Minimizing the maximum expected sample size im two stage phase ii clinical trials with continuous outcomes. Journal of Biopharmaceutical Statistics, 22: 0 836--852, 2012

  47. [47]

    Wasserman, A

    L. Wasserman, A. Ramdas, and S. Balakrishnan. Universal inference. Proceedings of the National Academy of Sciences, 117: 0 16880--16890, 2020