REVIEW 4 major objections 4 minor 42 references
Sequential Scoring Rule Evaluation for Forecast Method Selection
T0 review · 4 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Sequential forecast selection via scoring-rule ratios becomes an anytime-valid e-value test, with finite-sample error control at every stopping time.
desk verdict The SSRE-to-e-value connection is a genuine new idea, but the advertised finite-sample error control is not yet proven because the omega-grid is only validated pointwise while the e-process claim needs a uniform bound. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the accumulated score-ratio product $C_n(Q,P)=\exp\{\sum_{m=1}^n [S(\varphi[Q],Y_m)-S(\varphi[P],Y_m)]\}$, combined with an exponential tilting change of measure whose tilting parameter $h_P$ is the nonzero root of $E_P[\exp(h\log C_n(Q,P))]=1$. Lemma 2 shows that $C_n^{\omega}(Q,P)=\exp(\omega\Delta_n(Q,P))$ is an e-process under $H_Q$ for every $\omega\in[0,|h_P|]$, and the analogous statistic is an e-process under $H_P$. An e-value is a nonnegative random variable with expectation at most one under the null; an e-process extends this property to every stopping time, so Ville's inequality converts it into a sequential error bound. The same e-process structure is then scaled and averaged, following the e-value false-discovery-rate construction, to control errors across repeated comparisons.
What would settle it
In a simple experiment with a known data-generating process where $|h_P|$ can be computed exactly by solving $E_P[\exp(h\Delta_n)]=1$, generate data under the null $H_Q$, set $\omega$ strictly larger than that computed value, and run the SSRE boundary rule. If the empirical rejection rate for falsely selecting method $P$ exceeds the claimed bound $\beta/(1-\beta)$ in finite samples while the same experiment with $\omega<|h_P|$ stays below the bound, the central e-process claim is falsified.
Extended reading notes
Core claim
The central claim is that the SSRE process $C_n(Q,P)=\prod_{m=1}^n \exp\{S(\varphi[Q],Y_m)-S(\varphi[P],Y_m)\}$, with boundaries $k_l$ and $k_u$, behaves like a sequential probability ratio test even though it is not a likelihood ratio. Under the null that method $Q$ is at least as good as $P$, each multiplicative factor is an e-variable, and the product is an e-process, so $\Pr(\sup_{n\le T} C_n^{\omega}(Q,P) \ge 1/\pi)\le \pi$ for every $\omega\in[0,|h_P|]$, where $h_P$ is the tilting parameter that solves the moment generating function equation. This gives finite-sample control over the probability of wrongly selecting method $P$ when $Q$ is better, and, with probability converging to one, the procedure eventually selects the better method when the alternative is correct. The paper also proves the SSRE stopping time is finite almost surely and has an exponentially decaying tail.
Load-bearing premise
The procedure requires the user to pick values of $\omega$ that do not exceed the unknown tilting parameter $|h_P|$; if $\omega>|h_P|$, the e-process property fails and the advertised finite-sample error control does not hold.
Editorial extensions
If this is right
- A practitioner can compare two forecasting methods in real time and stop as soon as the accumulated scoring-rule ratio crosses a prechosen boundary, with error probabilities bounded in finite samples rather than relying on asymptotic normality.
- The procedure terminates almost surely and has finite stopping-time moments, so expected stopping time can be budgeted in advance using Markov's inequality.
- Because the construction needs no stationarity or weak-dependence assumption on the score differences, it applies exactly in settings where Diebold-Mariano style tests are known to fail, such as estimated models and nonstationary series.
- The e-value formulation allows false-discovery-rate control across multiple hypothesis tests of forecast accuracy, via ordered and scaled e-values, without requiring independence among the tests.
- The framework naturally extends beyond a single benchmark versus a single alternative; the authors note it should support construction of sequential model confidence sets.
- The authors conclude that the exponential tilting structure can be represented as a universal generalized e-value, tying forecast selection to the wider anytime-valid inference literature.
Reading between the lines
- A direct extension, not pursued in the paper, is to use the same e-process to construct a sequential test of equal predictive ability rather than a two-way selection rule, with the indifference region $k_l < C_n < k_u$ interpreted as a zone of practical equivalence.
- The averaging over several $\omega$ values suggests a calibration recipe one could test empirically: estimate $|h_P|$ from the training sample by solving the moment generating equation, then verify that all chosen $\omega$ lie below the estimate; the paper's informal visual comparison over the training set is a lighter-weight version of this.
- The e-process property may allow a betting-style interpretation: the score-ratio product is a wealth process in a game where the forecaster wagers on method $Q$ versus method $P$, which could connect SSRE to portfolio and prediction-market formulations of forecast evaluation.
- One could stress-test the method by deliberately choosing $\omega>|h_P|$ in a simulation to confirm the advertised error control fails exactly when the key inequality is violated.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a sequential scoring-rule evaluation (SSRE) framework for choosing between two forecasting methods. The key statistic is C_n(Q,P)=∏ exp(S(φ[Q],Y)-S(φ[P],Y)), and a boundary-crossing rule on C_n defines the stopping time N. The authors prove (i) a large-deviations/tilting lemma giving a root h_P of the MGF of Δ_n; (ii) termination with probability one and existence of all moments under weak nondegeneracy conditions; (iii) that powered statistics C_n^{(ω)} are generalized universal e-values under a null HQ defined by E_P[R_n|B_n]≤1, with Ville-based finite-sample error control; and (iv) asymptotic power for detecting a false null. A practical protocol chooses ω∈{1/4,1/2,1} by an informal training-sample check, sets β=1/10, and simulates comparisons against Diebold-Mariano tests in three examples.
Significance. If established, this would provide a genuinely sequential, finite-sample alternative to asymptotic forecast-comparison tests, and the termination/moment results under composite hypotheses are of independent interest. The paper is also careful to compare against DM tests in regimes where the DM assumptions fail. However, the headline error-control guarantee is not fully established as stated: the tilting parameter h_P is n-dependent, the proof of Lemma 3 has filtration errors, and the proof of Proposition 1 uses an invalid product inequality. The e-process interpretation is valuable but, under the paper's own definition of HQ, the ω=1 case follows almost by construction; the independent content is the powered/ω-parameterized version and the asymptotic power results.
major comments (4)
- [Section 4.1, Lemma 2] The proof establishes E_P[exp(ω Δ_n(Q,P))]≤1 only pointwise in n, because h_P in Lemma 1 is the root of E_P[exp(h log C_n)]=1 for that fixed n and is n-dependent. The text suppresses this subscript when asserting that 'for any ω∈[0,|h_P|], C_n^(ω)(Q,P) is an e-process.' An e-process requires a single ω to satisfy the moment inequality for every n simultaneously, and the recommended grid {1/4,1/2,1} in Section 5.1 may contain values exceeding |h_P(n)| at some sample sizes if inf_n |h_P(n)| < 1/4. Since Lemma 3 and Theorem 1 rely on the uniform-in-n e-process property, this gap is load-bearing; a condition such as ω≤inf_{n,P∈HQ}|h_P(n)|, or a genuinely adaptive choice of ω, is needed.
- [Appendix, proof of Lemma 3] The proof states that each R_n(Q,P) is adapted to B_n and uses E_P[R_n|B_{n-1}]≤1, while Section 3.1 defines HQ through E_P[R_n|B_n]≤1 and R_n involving Y_{n+τ} is B_{n+1}-measurable when τ=1. The displayed tower-property equality also treats conditional expectations of the product as a product of conditional expectations, which is not valid. This is repairable by consistently shifting the filtration or by re-indexing R_n, but as written the supermartingale/e-process argument for Lemma 3 is incorrect.
- [Appendix, proof of Proposition 1] The proof asserts Pr(N>n) ≤ ∏_{m=1}^n Pr(k_l < C_m < k_u). For arbitrary events, the probability of an intersection is not bounded above by the product of the marginal probabilities, so this inequality is not available without additional independence or positive-correlation structure. The subsequent Cauchy-product argument depends on this bound, and therefore the claimed exponential tail bound, termination with probability one, and Corollary 1's moment result are not proven as written. A conditional contraction argument or another mechanism is required.
- [Section 5.2, simulation protocol] The text sets β=1/10 and states k_u≈11.11, but the formulas in Section 4.1 and Section 5.1 give k_u=(1−β)/β=9 (and β_Q=β/(1−β)=1/9≈0.111). If the simulations actually used 11.11, the bound controlled by Theorem 1 is not the one advertised, and the reported rejection frequencies in Table 1 should be compared with the correct threshold.
minor comments (4)
- [Section 4.2, Theorem 2 proof] The proof introduces G^(ω)_{n_j}(Q,P) without definition; this appears to be a typo for C^(ω)_{n_j}(Q,P).
- [Section 3.2 and Figure 1] The notation for the continuation event alternates between E^{Q∩P}_n and E^{QXP}_n; a single consistent symbol would improve readability.
- [Throughout] There are numerous typographical errors in the references and text, including 'Wolforwitz' for Wolfowitz, 'Mathemetical' in the Brown reference, 'Chpman' for Chapman in Wetherill, and 'Role's theorem' for Rolle's theorem.
- [Section 5.1] The informal training-sample check only compares the trajectories of exp(−ω sign(Δ_R) Δ_R) over one finite window; as the authors acknowledge, this cannot verify the uniform moment condition required by Lemma 2, and the paper should state this limitation explicitly in the main text rather than only in the surrounding discussion.
Circularity Check
No significant circularity: the central e-value and termination results have independent derivation content; the uniform-omega and filtration gaps are correctness risks rather than definitional loops.
full rationale
The paper's advertised error-control claim does not reduce by construction to the definition of HQ. HQ is written as E_P[R_n(Q,P)|B_n] ≤ 1, which makes each increment R_n an e-variable only after taking expectations; the product C_n, and hence the e-process property needed for Ville's inequality, requires an additional supermartingale/tower-property step. The Appendix attempts that step, and although the attempt has an adaptation and tower-property error, the flaw is a missing argument, not an identity between input and output. Lemma 2's bound E[exp(ω Δ_n)] ≤ 1 for ω ∈ [0,|hP|] is derived from convexity of the MGF and the tilting root in Lemma 1, so the 'generalized e-value' characterization rests on a genuine derivation rather than on relabeling. Proposition 1 and Corollary 1 (termination and existence of moments) are self-contained finite-sample arguments and are independent of the e-value section. Section 5.1's heuristic selection of ω ∈ {1/4,1/2,1} does not verify ω ≤ |hP(n)| uniformly in n, and the Appendix proof of Lemma 3 misuses the tower property because R_n is not adapted to B_n; these are correctness or robustness gaps, not circular steps. Self-citations (Frazier et al. 2023; Poskitt 1987; Poskitt and Sengarapillai 2010) are contextual or confined to simulation calibration and are not load-bearing for the main conclusions. The e-process error-control bound is standard Ville-type machinery once the e-process property is established, but the establishment of that property is where the paper's derivation has gaps, not where it loops back on itself. Score 2 reflects minor non-load-bearing self-citations and a mild definitional flavor in defining HQ so that R_n is an e-variable, not a central circularity.
Assumptions & free parameters
free parameters (4)
- learning rate omega =
{1/4, 1/2, 1} in simulations
- error tolerance beta =
1/10
- boundary constants k_l, k_u =
k_l = 1, k_u = 11.11 (reported)
- training and evaluation window sizes R, N =
R = 500, N = 100
assumptions (5)
- domain assumption The moment generating function of log C_n(Q,P) is finite for all real h (Lemma 1).
- domain assumption E_P[log C_n(Q,P)] != 0 and the distribution of C_n has mass on both sides of 1 (Lemma 1).
- ad hoc to paper HQ is defined as the set of P for which E_P[R_n(Q,P)|B_n] <= 1 for all n (Section 3.1).
- domain assumption A weak law of large numbers holds for n^{-1} Delta_n(Q,P) (Lemma 4).
- domain assumption For each m, the distribution of C_m is nondegenerate and at least one boundary-crossing probability is nonzero (Proposition 1).
Cite this review
Pith. "Pith review of Sequential Scoring Rule Evaluation for Forecast Method Selection." pith.science (2026). https://pith.science/paper/E7RHXYOO
@misc{pith2026250509090,
author = {Pith},
title = {Pith review of: Sequential Scoring Rule Evaluation for Forecast Method Selection},
year = {2026},
howpublished = {\url{https://pith.science/paper/E7RHXYOO}},
note = {Machine review of arXiv:2505.09090}
}
read the original abstract
This paper shows that sequential statistical analysis techniques can be generalised to the problem of selecting between alternative forecasting methods using scoring rules. A return to basic principles is necessary in order to show that ideas and concepts from sequential statistical methods can be adapted and applied to sequential scoring rule evaluation (SSRE). One key technical contribution of this paper is the development of a large deviations type result for SSRE schemes using a change of measure that parallels a traditional exponential tilting form. Further, we also show that SSRE will terminate in finite time with probability one, and that the moments of the SSRE stopping time exist. A second key contribution is to show that the exponential tilting form underlying our large deviations result allows us to cast SSRE within the framework of generalised e-values. Relying on this formulation, we devise sequential testing approaches that are both powerful and maintain control on error probabilities underlying the analysis. Through several simulated examples, we demonstrate that our e-values based SSRE approach delivers reliable results that are more powerful than more commonly applied testing methods precisely in the situations where these commonly applied methods can be expected to fail.
Figures
Reference graph
Works this paper leans on
-
[1]
barticle [author] Barnard , G. A. G. A. ( 1946 ). Sequential Tests in Industrial Statistics . Journal of the Royal Statistical Society Supplement 8 1-26 . Contains four pages of discussion . barticle
work page 1946
-
[2]
barticle [author] Brehmer , J. R. J. R. Gneiting , T. T. ( 2021 ). Scoring interval forecasts: Equal-tailed, shortest, and modal interval. Bernoulli 27 1993-2110 . barticle
work page 2021
-
[3]
barticle [author] Brown , B. M. B. M. ( 1969 ). Moments of a stopping rule related to the central limit theorem . Annals of Mathemetical Statistics 40 1236-1249 . barticle
work page 1969
-
[4]
bbook [author] DeGroot , M. H. M. H. ( 1970 ). Optimal Statistical Decisions . McGraw-Hill , New York . bbook
work page 1970
-
[5]
barticle [author] Dey , Neil N. , Martin , Ryan R. Williams , Jonathan P J. P. ( 2024 a). Anytime-Valid Generalized Universal Inference on Risk Minimizers . arXiv preprint arXiv:2402.00202 . barticle
arXiv 2024
-
[6]
Multiple Testing in Generalized Universal Inference
barticle [author] Dey , Neil N. , Martin , Ryan R. Williams , Jonathan P J. P. ( 2024 b). Multiple Testing in Generalized Universal Inference . arXiv preprint arXiv:2412.01008 . barticle
work page Pith review arXiv 2024
-
[7]
barticle [author] Diebold , Francis X F. X. ( 2015 ). Comparing predictive accuracy, twenty years later: A personal perspective on the use and abuse of Diebold--Mariano tests . Journal of Business & Economic Statistics 33 1--1 . barticle
work page 2015
-
[8]
barticle [author] Diebold , Francis X F. X. Mariano , Robert S R. S. ( 1995 ). Comparing predictive accuracy . Journal of Business & Economic Statistics 13 253--263 . barticle
work page 1995
Show all 42 references
-
[9]
bbook [author] Doob , J. L. J. L. ( 1953 ). Stochastic Processes . John Wiley , New York . bbook
1953
-
[10]
barticle [author] Dvoretzky , A. A. , Kiefer , J. J. Wolforwitz , J. J. ( 1953 ). Sequential decision problems for processess with continuous time parameter . Annals of Mathematical Statistics 24 254-264 . barticle
1953
-
[11]
bbook [author] Ferguson , Thomas S T. S. ( 1967 ). Mathematical Statistics: A Decision Theoretic Approach . Academic Press . bbook
1967
-
[12]
Ziegel , Johanna F J
barticle [author] Fissler , Tobias T. Ziegel , Johanna F J. F. ( 2016 ). Higher order elicitability and Osband’s principle . The Annals of Statistics 44 1680--1707 . barticle
2016
-
[13]
btechreport [author] Frazier , David T D. T. , Covey , Ryan R. , Martin , Gael M G. M. Poskitt , Donald S D. S. ( 2023 ). Solving the forecast combination puzzle Technical Report , arXiv preprint arXiv:2308.05263 . btechreport
2023 arXiv
-
[14]
White , Halbert H
barticle [author] Giacomini , Raffaella R. White , Halbert H. ( 2006 ). Tests of conditional predictive ability . Econometrica 74 1545--1578 . barticle
2006
-
[15]
( 2011 )
barticle [author] Gneiting , Tilmann T. ( 2011 ). Making and evaluating point forecasts . Journal of the American Statistical Association 106 746--762 . barticle
2011
-
[16]
, Balabdaoui , Fadoua F
barticle [author] Gneiting , Tilmann T. , Balabdaoui , Fadoua F. Raftery , Adrian E A. E. ( 2007 ). Probabilistic forecasts, calibration and sharpness . Journal of the Royal Statistical Society: Series B (Statistical Methodology) 69 243--268 . barticle
2007
-
[17]
Raftery , Adrian E A
barticle [author] Gneiting , Tilmann T. Raftery , Adrian E A. E. ( 2007 ). Strictly Proper Scoring Rules, Prediction, and Estimation . Journal of the American Statistical Association 102 359--378 . barticle
2007
-
[18]
Ranjan , Roopesh R
barticle [author] Gneiting , Tilmann T. Ranjan , Roopesh R. ( 2011 ). Comparing Density Forecasts Using Threshold- and Quantile-Weighted Scoring Rules . Journal of Business & Economic Statistics 29 411-422 . 10.1198/jbes.2010.08110 barticle
2011 arXiv
-
[19]
, de Heide , Rianne R
barticle [author] Gr\" u nwald , Peter P. , de Heide , Rianne R. Koolen , Wouter W. ( 2024 ). Safe Testing . Journal of the Royal Statistical Society Series B: Statistical Methodology 86 1091-1128 . barticle
2024
-
[20]
barticle [author] Hall , W. J. W. J. ( 1970 ). On Wald's equations in continuous time . Journal of Applied Probability 7 59-68 . barticle
1970
-
[21]
bbook [author] Halmos , P. R. P. R. ( 1950 ). Measure Theory . Van Nostrand Reinhold , New York . bbook
1950
-
[22]
barticle [author] Hansen , Peter Reinhard P. R. ( 2005 ). A test for superior predictive ability . Journal of Business & Economic Statistics 23 365--380 . barticle
2005
-
[23]
barticle [author] Hansen , Peter R P. R. , Lunde , Asger A. Nason , James M. J. M. ( 2011 ). The model confidence set . Econometrica 79 453--497 . barticle
2011
-
[24]
barticle [author] Lai , T. Z. T. Z. , Gross , S. T. S. T. Shen , D. B. D. B. ( 2011 ). Evaluating probability forecasts . Annals of Statistics 39 2356-2382 . barticle
2011
-
[25]
barticle [author] Lazarus , E. E. , Lewis , D. J. D. J. , Stock , J. H. J. H. Watson , M. W. M. W. ( 2018 ). HAR Inference: Recommendations for Practice . Journal of Business and Economic Statistics 36 541-559 . barticle
2018
-
[26]
barticle [author] Martin , Gael M G. M. , Loaiza-Maya , Rub \'e n R. , Maneesoonthorn , Worapree W. , Frazier , David T D. T. Ram \' rez-Hassan , Andr \'e s A. ( 2022 ). Optimal probabilistic forecasts: When do they work? International Journal of Forecasting 38 384--406 . barticle
2022
-
[27]
bbook [author] Nocedal , J. J. Wright , S. S. ( 2006 ). Numerical Optimization , 2nd ed. Springer . bbook
2006
-
[28]
barticle [author] Patton , Andrew J A. J. ( 2020 ). Comparing possibly misspecified forecasts . Journal of Business & Economic Statistics 38 796--809 . barticle
2020
-
[29]
barticle [author] Poskitt , D. S. D. S. ( 1987 ). Precision, complexity and B ayesian model determination . Journal of the Royal Statistical Society: Series B 49 199-208 . barticle
1987
-
[30]
barticle [author] Poskitt , D. S. D. S. Sengarapillai , A. A. ( 2010 ). Dual P-Values, Evidential Tension and Balanced Tests . Department of Econometrics and Business Statistics Working Paper 15/10 . barticle
2010
-
[31]
barticle [author] Robbins , H. H. Samuel , E. E. ( 1966 ). An extention of a lemma of Wald . Journal of Applied Probability 3 272-273 . barticle
1966
-
[32]
( 2021 )
barticle [author] Shafer , Glenn G. ( 2021 ). Testing by betting: A strategy for statistical and scientific communication . Journal of the Royal Statistical Society Series A 184 407-431 . barticle
2021
-
[33]
barticle [author] Tashman , L. . J. L. . J. ( 2000 ). Out-of-sample tests of forecasting accuracy: An analysis and review . International Journal of Forecasting 16 437-450 . barticle
2000
-
[34]
Wang , Ruodu R
barticle [author] Vovk , Vladimir V. Wang , Ruodu R. ( 2021 ). E-values: C alibration, combination and applications . Annals of Statistics 39 1736-1754 . barticle
2021
-
[35]
bbook [author] Wald , A. A. ( 1947 ). Sequential Analysis . John Wiley , New York . bbook
1947
-
[36]
Ramdas , Aaditya A
barticle [author] Wang , Ruodu R. Ramdas , Aaditya A. ( 2022 ). False discovery rate control with e-values . Journal of the Royal Statistical Society Series B 84 822-852 . barticle
2022
-
[37]
, Ramdas , Aaditya A
barticle [author] Wasserman , Larry L. , Ramdas , Aaditya A. Balakrishnan , Sivaraman S. ( 2020 ). Universal inference . Proceedings of the National Academy of Sciences 117 16880--16890 . barticle
2020
-
[38]
barticle [author] West , Kenneth D K. D. ( 1996 ). Asymptotic inference about predictive ability . Econometrica: Journal of the Econometric Society 64 1067--1084 . barticle
1996
-
[39]
bbook [author] Wetherill , G. B. G. B. ( 1975 ). Sequential Methods in Statistics , 2 ed. Chapmen and Hall , Cambridge . bbook
1975
-
[40]
( 2000 )
barticle [author] White , Halbert H. ( 2000 ). A reality check for data snooping . Econometrica 68 1097--1126 . barticle
2000
-
[41]
barticle [author] Yen , Y. Y. Yen , T. T. ( 2021 ). Testing forecast accuracy of expectiles and quantiles with the extremal consistent loss functions . International Journal of Forecasting 37 733-758 . barticle
2021
-
[42]
barticle [author] Zhu , Y. Y. Timmermann , A. A. ( 2020 ). Can Two Forecasts Have the Same Conditional Expected Accuracy? arXiv:2006.03238v2 [stat.ME] . barticle
2020 arXiv
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.