REVIEW 2 major objections 4 minor 15 references
Supermartingales for One-Sided Tests: Sufficient Monotone Likelihood Ratios are Sufficient
T0 review · 2 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read One-sided t-tests are anytime-valid after all.
desk verdict A short, correct paper that resolves the Wang–Ramdas open question by proving a sufficient-statistic/MLR condition, with honest examples and a useful counterexample. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is Lemma 3, which shows that sufficiency forces the conditional likelihood ratio of the raw outcome to equal the conditional likelihood ratio of the sufficient statistic, $p^{U^n}_{\delta_+}(u^n \mid u^{n-1}) = p^{T_n}_{\delta_+}(t_n(u^n) \mid u^{n-1})$, and that the latter is non-decreasing in $t$ under the monotone likelihood ratio property. This lets the proof apply a known fixed-sample fact pointwise, conditional on the past: a monotone likelihood ratio likelihood ratio has expectation at most $1$ under every $\delta \leq \delta_0$. The resulting past-conditional e-variables multiply into a supermartingale. For the $t$-test, the sufficiency of the $t$-statistic and the monotone likelihood ratio property of noncentral $t$ densities supply the two ingredients.
What would settle it
Construct a coarsened model that has a one-dimensional sufficient statistic with the monotone likelihood ratio property at every $n$, simulate data from any $\delta < \delta_0$, and estimate $\mathbb{E}_\delta[p^{U^n}_{\delta_+}(U^n)]$ for increasing $n$; the theorem predicts each value is at most $1$, so any estimate reliably above $1$ would refute the central claim.
Extended reading notes
Core claim
Theorem 4 is the central claim: let $T_n = t_n(U^n)$ be a sufficient statistic for the coarsened model $\{P^{U^n}_{\delta} : \delta \in \Delta\}$, and suppose $T_n$ satisfies the monotone likelihood ratio property for every $n$. Then the process formed by multiplying the conditional likelihood ratios $p^{T_i}_{\delta_+}(T_i \mid U^{i-1})$ for $i = 1, \ldots, n$ is identical to the likelihood ratio process $(p^{U^n}_{\delta_+})_n$, and both are test supermartingales relative to the one-sided null $H_{\delta \leq \delta_0}$. For the anytime-valid $t$-test, the theorem resolves the open question about the one-sided null: because the $t$-statistic is sufficient for the maximal invariants and noncentral $t$ densities have monotone likelihood ratios, the $t$-likelihood ratio process is a supermartingale for $H_{\delta \leq \delta_0}$. The fixed-sample e-variable result then extends to the whole process by conditioning on the past and multiplying.
Load-bearing premise
The argument assumes that at every sample size $n$ the coarsened model admits a one-dimensional sufficient statistic $T_n$ whose likelihood ratio is monotone in $T_n$; if no such sufficient statistic exists, the paper does not establish the supermartingale property.
Editorial extensions
If this is right
- The anytime-valid $t$-test for the one-sided null $H_{\delta \leq \delta_0}$ is a genuine e-process: practitioners may stop at arbitrary data-dependent times and retain type-I error control.
- The same conclusion holds when a prior $W$ on the alternative replaces the point mass at $\delta_+$, giving log-optimal anytime-valid e-values for the one-sided null.
- The location-invariant $\chi^2$-test and the linear-regression likelihood ratio with nuisance covariates are supermartingales under their respective one-sided nulls.
- The argument is not a blank cheque: a zero-mean sub-Gaussian (symmetric Bernoulli) example shows the $t$-test likelihood ratio can have expectation greater than $1$ at fixed $n$, so the sufficient-statistic MLR condition is doing real work.
Reading between the lines
- For the designer of new anytime-valid tests, the proof suggests a practical recipe: find a one-dimensional sufficient statistic for the coarsened invariants and verify its MLR property, instead of trying to prove monotonicity of the raw conditional likelihood ratio, which often fails.
- The argument likely transfers to any group-invariant testing problem whose maximal invariants admit a one-dimensional sufficient statistic with MLR; spherical and elliptical families are natural candidates for new e-processes.
- A direct extension would replace the one-dimensional statistic by a vector sufficient statistic with a suitable stochastic ordering; success there would apply to multivariate analogues of the $t$-test.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies sequential testing of one-sided composite null hypotheses of the form H_{δ≤δ0} using likelihood ratio processes built from coarsened data (e.g., maximal invariants). The main result, Lemma 3 and Theorem 4, states that if for every sample size n there exists a one-dimensional sufficient statistic T_n for the coarsened model satisfying the monotone likelihood ratio (MLR) property, then the likelihood ratio process (p^{U^n}_{δ+}) is a test supermartingale relative to H_{δ≤δ0}. The key step shows that the conditional likelihood ratio of the next coarsened outcome given the past is equal to a conditional likelihood ratio of the sufficient statistic, which is increasing in the sufficient statistic by the MLR assumption. Applications are developed for the scale-invariant t-test (answering an open question of Wang and Ramdas 2024), a χ²-test for variances with translation nuisance, sequential linear regression with nuisance covariates, and a label-agnostic Bernoulli test. An appendix demonstrates that the theorem's conditions are not vacuous by giving a counterexample where the t-likelihood ratio fails to be an e-variable under Rademacher data.
Significance. The paper gives a simple, general, and checkable sufficient condition for converting one-sided parametric likelihood ratio tests into anytime-valid e-processes. The t-test application is particularly valuable, as it resolves a previously open question about whether the scale-invariant t-likelihood ratio process controls the one-sided composite null. The proof is elementary and transparent, and the authors are explicit about the condition being load-bearing. The applications cover several canonical settings, and each verifies the required MLR and sufficiency assumptions. The counterexample in Appendix C strengthens the paper by delimiting the scope of the theorem. The paper is honest about the simplicity of the proof and gives proper credit to prior work.
major comments (2)
- [Section 2, Theorem 4] The statement of Theorem 4 defines the two processes as (∏_{i=1}^n p^{T_i}_δ(T_i | U^{i-1}))_{n∈N} and (p^{U^n}_δ)_{n∈N}, but the supermartingale claim relative to H_{δ≤δ0} requires these processes to be built from an alternative δ+ ≥ δ0, not from the generic parameter δ. As written, the theorem asserts that the likelihood ratio process (p^{U^n}_δ) is a test supermartingale for H_{δ≤δ0}, which is false in general: for δ ≤ δ0, this process is typically not a supermartingale under the null distributions. The proof of Theorem 4 and all applications correctly use δ+; in particular, the product appearing after (8) has the same subscript error. Please correct the subscripts in the theorem statement and proof.
- [Section 1, Definition 1] The definition of the monotone likelihood ratio property is ambiguous because the notation p^T_{δ+}(t) was introduced as the density of P^T_{δ+} relative to the fixed baseline P_{δ0} of the testing problem. If the phrase 'for all δ0, δ+ ∈ Δ with δ0 ≤ δ+' is read with that fixed baseline, the condition is not the standard pairwise MLR property, and it would not generally imply the stochastic dominance used in the proof of Proposition 2 in Appendix A. Please restate Definition 1 as the usual pairwise MLR property (for all δ1 ≤ δ2, the density ratio f_{δ2}(t)/f_{δ1}(t) is increasing in t), or clarify that the ratios are taken with respect to the δ0 appearing in the quantifier.
minor comments (4)
- [Section 2, proof of Theorem 4] The product in the sentence after (8) uses p^{T_i}_δ(· | U^{i-1}) instead of p^{T_i}_{δ+}(· | U^{i-1}); this is the same subscript typo as in the theorem statement.
- [Section 1 and Lemma 3] The notation p^{U^n}_{δ+}(u^n|u^{n-1}) denotes the conditional density of the last outcome U_n given the past; using p^{U_n}_{δ+}(u_n|u^{n-1}) would avoid confusion with the density of the full vector U^n.
- [Appendix C] The Taylor expansion E_{Rad}[M^δ_n] = 1 + (n-1)/6 δ^4 + O(δ^6) is stated without derivation; a short justification would make the counterexample reproducible.
- [References] The reference to 'Wang (2024) Personal communication' for the numerical observation that the conditional likelihood ratio is not monotone in the t-test is not verifiable; the authors should either provide the details in the paper or cite a public version of this work.
Circularity Check
No significant circularity: the supermartingale theorem follows directly from sufficiency and the MLR property, with all load-bearing ingredients either proved in the paper or drawn from independent external results.
full rationale
The paper's central claim is a conditional theorem: if each T_n is sufficient and satisfies the MLR property, then the likelihood-ratio process is a test supermartingale for the one-sided null. The proof is self-contained. Lemma 3 derives p^{U^n}_{δ+}(u^n|u^{n-1}) = p^{T_n}_{δ+}(t_n(u^n)|u^{n-1}) directly from the sufficiency factorization (4), so no equality is transformed into the conclusion by construction. Theorem 4 invokes Proposition 2, which is stated as following from Grünwald et al. (2024) but is proved in Appendix A via Lehmann–Romano stochastic dominance and the cancellation identity h(δ0) = 1, so the self-citation is not load-bearing. The t-test application rests on external facts: the noncentral t MLR property (Kruskal, 1954) and sufficiency of the t-statistic, which the paper verifies directly and also attributes to Pérez-Ortiz et al. (2024). The 'No Free For All' appendix demonstrates non-triviality by exhibiting a sub-Gaussian distribution under which the same process is not an e-variable, confirming that the theorem is not a tautology. No parameter is fitted and later called a prediction; the sufficient-statistic assumption is explicitly stated and verified in each example rather than engineered to force the conclusion.
Assumptions & free parameters
assumptions (4)
- standard math The family of distributions Pδ and the coarsenings U^n admit regular conditional densities and are mutually absolutely continuous.
- domain assumption For each n, there exists a sufficient statistic T_n for the model {Pδ: δ∈Δ} such that T_n has the MLR property (Def. 1).
- standard math The t-statistic is sufficient for the coarsened data U^n and follows a noncentral t distribution (for the Gaussian scale model).
- standard math The noncentral t family has the monotone likelihood ratio property (Kruskal 1954).
Cite this review
Pith. "Pith review of Supermartingales for One-Sided Tests: Sufficient Monotone Likelihood Ratios are Sufficient." pith.science (2026). https://pith.science/paper/CKIL56GK
@misc{pith2026250204208,
author = {Pith},
title = {Pith review of: Supermartingales for One-Sided Tests: Sufficient Monotone Likelihood Ratios are Sufficient},
year = {2026},
howpublished = {\url{https://pith.science/paper/CKIL56GK}},
note = {Machine review of arXiv:2502.04208}
}
abstract
The t-statistic is a widely-used scale-invariant statistic for testing the null hypothesis that the mean is zero. Martingale methods enable sequential testing with the t-statistic at every sample size, while controlling the probability of falsely rejecting the null. For one-sided sequential tests, which reject when the t-statistic is too positive, a natural question is whether they also control false rejection when the true mean is negative. We prove that this is the case using monotone likelihood ratios and sufficient statistics. We develop applications to the scale-invariant t-test, the location-invariant $\chi^2$-test and sequential linear regression with nuisance covariates.
Reference graph
Works this paper leans on
-
[1]
J. Bhowmik and M. King. Maximal invariant likelihood based testing of semi-linear models. Statistical Papers, 48: 0 357--–383, 2007
work page 2007
-
[2]
L. D. Brown, I. M. Johnstone, and K. B. MacGibbon. Variation diminishing transformations: a direct approach to total positivity and its statistical applications. Journal of the American Statistical Association, 76 0 (376): 0 824--832, 1981
work page 1981
-
[3]
P. D. Gr\"unwald, R. de Heide, and W. M. Koolen. Safe testing. Journal of the Royal Statistical Society Series B: Statistical Methodology, 86 0 (5): 0 1091--1128, 2024. With Discussion
work page 2024
-
[4]
W. M. Koolen and P. Gr \"u nwald. Log-optimal anytime-valid e-values. International Journal of Approximate Reasoning, 2021. Festschrift for G. Shafer's 75th Birthday
work page 2021
-
[5]
W. Kruskal. The Monotonicity of the Ratio of Two Noncentral t-Density Functions . The Annals of Mathematical Statistics, 25 0 (1): 0 162--165, 1954
work page 1954
-
[6]
M. Larsson, A. Ramdas, and J. Ruf. The numeraire e-variable and reverse information projection. arXiv preprint arXiv:2402.18810, 2024
arXiv 2024
-
[7]
E. Lehmann and J. P. Romano. Testing statistical hypotheses, volume 3. Springer, 1986
work page 1986
- [8]
Show all 15 references
-
[9]
M. F. P \'e rez-Ortiz, T. Lardy, R. de Heide, and P. Gr \"u nwald. E-statistics, group invariance and anytime valid testing. The Annals of Statistics, 52 0 (4): 0 1410--1432, 2024
2024
-
[10]
Ramdas, P
A. Ramdas, P. Gr\"unwald, V. Vovk, and G. Shafer. Game-theoretic statistics and safe anytime-valid inference. Statist. Sci., 38 0 (4): 0 576--601, 2023. ISSN 0883-4237. doi:10.1214/23-STS894
2023 doi
-
[11]
Turner, A
R. Turner, A. Ly, and P. Gr \"u nwald. Generic e-variables for exact sequential k-sample tests that allow for optional stopping. Statistical Planning and Inference, 230: 0 106116, 2024
2024
-
[12]
Schure Ter ter Schure, M
J. Schure Ter ter Schure, M. Pérez-Ortiz, A. Ly, and P. Gr \"u nwald. The anytime-valid logrank test: Error control under continuous monitoring with unlimited horizon. New England Journal of Statistics in Data Science, 2 0 (2): 0 190--214, 2024
2024
-
[13]
H. Wang. Personal communication, 2024
2024
-
[14]
Wang and A
H. Wang and A. Ramdas. Anytime-valid t-tests and confidence sequences for G aussian means with unknown variance, 2024
2024
-
[15]
Williams
D. Williams. Probability with Martingales. Cambridge Mathematical Textbooks, 1991
1991
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.