REVIEW 4 major objections 5 minor 23 references
Evaluating probabilistic forecasts of football matches: The case against the Ranked Probability Score
T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read This paper argues that the ranked probability score, despite its popularity for evaluating football forecasts, is less effective than the ignorance score at identifying the better forecasting system, and recommends the ignorance score for…
desk verdict A well-written and useful critique of the RPS, but its central simulation result is a likelihood-ratio theorem, not empirical evidence; the strong recommendation outruns the evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a controlled comparison of three proper scoring rules under the perfect-model scenario: the Brier score, the ranked probability score (which compares cumulative forecast distributions over ordered outcomes), and the ignorance score (the negative base-2 logarithm of the probability assigned to the occurred outcome). Each scoring rule is judged by the probability that its mean score over $n$ matches selects the perfect forecasting system over a single imperfect alternative. The concept of locality, whether a score uses only the probability at the occurred outcome, is the key discriminator; the paper argues that since the true probabilities of non-outcomes are unknowable even after the result, only the outcome probability is legitimately usable.
What would settle it
Find a pair of real forecasting systems where the RPS ranks the system that is measurably better out-of-sample, such as higher betting profit, while the ignorance score ranks the worse one; if such cases are common, the paper's recommendation fails in the imperfect-model context it claims to address. More directly, re-run the perfect-model experiment with a set of multiple imperfect alternatives and a non-local score designed to exploit differences in the non-outcome probabilities; if that score identifies the perfect system faster than the ignorance score, the claim that extra distributional information is useless is false.
Extended reading notes
Core claim
The paper's central claim is that sensitivity to distance and non-locality, the properties that motivate the RPS, provide no benefit in the practical task of comparing football forecasting systems. The author defines that task as selecting the forecasting system closer to the data-generating distribution; in a perfect-model experiment, the right score should require fewer observed matches to identify the perfect system. The ignorance score, the negative log probability of the observed outcome, outperforms both the non-local Brier score and the distance-sensitive RPS, with the RPS notably inefficient when the two candidate forecasts differ only slightly. The conclusion is that the RPS should not be recommended over the ignorance score for football forecasts.
Load-bearing premise
The whole comparison rests on the assumption that better means more likely to pick the perfect forecast when the only alternative is one fixed imperfect forecast; since the ignorance score is the log-likelihood, that test is mathematically guaranteed to win, so the paper's headline result is built into the experimental setup.
Editorial extensions
If this is right
- Forecast evaluation studies that currently report only the RPS for football should also report the ignorance score, since the two can rank imperfect forecast systems differently.
- The difference in mean ignorance between two forecasting systems has a direct bits-per-match interpretation, making the size of a skill difference meaningful in a way the RPS and Brier score do not offer.
- If a forecast is intended for a specific binary decision, such as a lay bet on the away win, the derived binary forecast should be evaluated on its own rather than through a single three-outcome score.
- The results imply that adding distance sensitivity to a score does not help identify better forecasts more quickly, contrary to the reasoning that motivated adoption of the RPS.
Reading between the lines
- Because the ignorance score is the logarithm of the forecast likelihood, the perfect-model selection experiment is a likelihood-ratio test, and its dominance over the other scores in that restricted setting is guaranteed by standard hypothesis-testing theory; the empirical result is therefore a demonstration of that theorem rather than a football-specific discovery.
- In the more realistic imperfect-model scenario, the best score may depend on the user's loss function; a bettor facing asymmetric payoffs should test forecast systems on betting profit rather than on a generic score, and the paper's comparison does not settle that choice.
- A natural extension is to test whether an outcome-dependent non-local score that uses the full forecast distribution to assess calibration can beat the ignorance score in imperfect-model settings where forecast types differ, such as overconfidence versus bias.
- The same simulation protocol could be applied to other ordered discrete sports outcomes, such as margin-of-victory bins, to see whether distance sensitivity has any value outside football's three-outcome space.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper challenges the widespread use of the ranked probability score (RPS) for evaluating probabilistic football forecasts. After defining the RPS, Brier, and ignorance scores and reviewing properties such as propriety, locality, and sensitivity to distance, the author argues that the reasoning of Constantinou and Fenton for preferring the RPS is flawed because a single observed outcome cannot determine which of two forecasts is better without knowing the data-generating distribution. Two simulation experiments are then used to compare the three scores' ability to identify a 'perfect' forecasting system against an 'imperfect' one: the first uses the five Constantinou–Fenton match examples, the second uses pairs of forecasts constructed from bookmaker odds. In both experiments the ignorance score selects the perfect system more often than the RPS or Brier score, and the paper concludes with a strong recommendation to use the ignorance score.
Significance. If the conclusions were supported, the paper would be a useful contribution to the ongoing debate about scoring rules for ordered categorical outcomes, with practical implications for sports forecasting and model comparison. The paper has clear strengths: it identifies a real weakness in judging forecasts by a single outcome, it carefully defines the scoring rules and their properties, and the simulations are transparent and reproducible in principle. However, the central empirical claim is not informative in the way the paper presents it, because the experiments are simple hypothesis tests in which the ignorance score is the likelihood-ratio statistic and is therefore guaranteed to be optimal by well-known statistical theory. The paper's broad recommendation extends beyond the perfect-model scenario that it actually tests. The contribution would be strengthened by citing the relevant optimality theory and by either softening or re-deriving the practical recommendations.
major comments (4)
- [§6–§7, Figs. 3–7] The simulation design makes the ignorance score's superiority a mathematical identity rather than an empirical discovery. In both experiments the outcome for each match is generated from one of two known candidate distributions, the 'perfect' system P_i and the 'imperfect' system Q_i, and the decision rule is to select the system with the lower mean score. For the ignorance score, the difference in mean scores is (1/n) Σ log[p_{P_i}(y_i)/p_{Q_i}(y_i)], the mean log-likelihood ratio. In §6 the true distribution is chosen with equal prior probability, so the zero threshold is the Bayes decision rule and maximizes the probability of selecting the perfect system by the Neyman–Pearson lemma; in §7 the same pairwise likelihood-ratio comparison is used. The Brier and RPS were not expected to beat this rule, and the empirical results illustrate a theorem rather than provide evidence about football forecasts. The paper should either present the result as a direct consequence of likelihood-ratio optimality and cite the relevant theory, or redesign the experiment so that the decision problem is not a simple hypothesis test.
- [§5 and §8, Abstract] The strong recommendation in §8 ('we strongly recommend the ignorance score') and the abstract's claim that the results 'cast doubt on the value of non-locality and sensitivity to distance' go beyond the evidence. Section 5 explicitly says that the imperfect-model scenario, which is the practically relevant one, 'is left as future work,' and all experiments only compare a perfect system with a single imperfect alternative. The simulations therefore cannot discriminate between scores for the practically relevant task of choosing among several imperfect forecast systems, and they do not establish that non-locality or sensitivity to distance is valueless. At most, the simulations show that non-local scores are not needed in a simple-hypothesis perfect-model selection problem. The recommendation and abstract should be tempered accordingly.
- [§5 and §7.1] The paper uses 'efficiency in selecting the perfect model in finite samples' as the criterion for comparing scoring rules, but it never justifies why this decision-theoretic quantity is the appropriate criterion for choosing a score for practical forecast evaluation. Propriety already ensures that each score prefers the perfect system in expectation; the finite-sample selection probability depends on the test statistic and the class of alternatives. Without a link between selection probability and the forecast user's loss function, the experiments measure a property that may not correspond to the aims described in §4. The author should either provide such a justification or present the finite-sample results as an illustration of a known property of the log score rather than as a basis for a general recommendation.
- [§8] The conceptual argument that probabilities placed on non-outcomes are irrelevant is asserted rather than established. The discussion says that knowing the outcome reveals little about the true probabilities of other outcomes, but this does not imply that a score should ignore the rest of the forecast distribution; proper scores can use the full distribution and still rank imperfect forecasts differently. The paper should either supply a decision-theoretic argument for why only the probability on the realized outcome should matter, or explicitly acknowledge that this is a normative position rather than a consequence of the experiments.
minor comments (5)
- [§6.1] The sentence 'The results for matches one to four are shown in figure 3' should refer to Figure 4, not Figure 3.
- [§3, §6] References to 'table 3' for the Constantinou–Fenton match examples and the colour scheme are inconsistent with the table numbering; the match examples are in Table 1 and the colour scheme is in Table 2.
- [§7.1] The sentence 'for the lowest level of imperfection, in which δ = 0.1' contradicts the definition of δ and the later classification of δ = 0.01 and 0.025 as the two lowest levels; this appears to be a typographical error that should be corrected.
- [§2.2, §5] There are several typographical slips, including 'An scoring rule' in §2.2 and 'truely' in §5, and the typesetting of 'Staël von Holstein' should be fixed.
- [§1] The statement that the ignorance score is 'the only local and proper scoring rule' should be qualified in the standard way, typically 'up to affine transformation,' to avoid overclaiming.
Circularity Check
The simulation 'evidence' for the ignorance score is the Neyman-Pearson lemma in disguise: the selection rule and the log-score definition make the experiment's outcome a theorem.
-
other
[Section 6 (Experiment one, selection protocol) and Section 8 (Discussion, recommendation)]
"For a given pair of forecasts, define the outcomes of a series of n matches by drawing from forecast α or forecast β with equal probability 0.5. ... A scoring rule is defined to 'select' a forecasting system if it is assigned the lowest mean score over n forecasts."
With this selection rule, the ignorance score's comparison is exactly the log-likelihood ratio: choose the perfect system iff (1/n)Σ[log p_i^P(y_i) − log p_i^Q(y_i)] > 0. In Experiment one the true distribution is α or β with equal probability, so threshold 0 is the Bayes decision rule; by Neyman-Pearson/Bayes decision theory this maximizes the probability of selecting the perfect system among all rules. The simulated 'outperformance' of the ignorance score is therefore a mathematical consequence of the definition of the score and the experimental protocol, not an empirical discovery about football forecasts. Experiment two uses the same pair-by-pair mean-score comparison, so the log-score rule remains the likelihood-ratio test.
full rationale
The paper's only load-bearing circular element is the interpretation of the simulations. The conceptual argument against the RPS (Section 3) and the locality argument (Section 8) are non-circular, though debatable. The self-citation to Wheatcroft (2019) in Section 4 is minor and not load-bearing. However, the headline empirical claim that the ignorance score outperforms the RPS and Brier in identifying a perfect forecasting system is not an independent result: the experiment defines 'selecting' a system as having the lower mean score, and for the ignorance score that decision rule is the likelihood-ratio test, which is optimal for the equal-prior problem in Experiment one. The simulation outcomes are therefore forced by construction, and presenting them as evidence for the recommendation (Section 8) is equivalent to citing an uncited theorem as empirical support. This warrants a score of 6: partial circularity, because the central empirical 'prediction' reduces by construction, while the philosophical arguments retain independent content.
Assumptions & free parameters
free parameters (1)
- δ (imperfection threshold) =
0.01, 0.025, 0.05, 0.1
assumptions (4)
- domain assumption The set of normalized bookmaker odds from football-data.co.uk is a representative distribution of realistic match outcome probabilities.
- domain assumption The probability of selecting a perfect forecasting system over a single imperfect alternative is a meaningful criterion for comparing scoring rules for practical forecast evaluation.
- standard math The ignorance score is the only proper and local scoring rule for discrete outcomes.
- domain assumption Scoring rules used for evaluation should be assessed by how efficiently they identify a perfect forecasting system in finite samples.
Cite this review
Pith. "Pith review of Evaluating probabilistic forecasts of football matches: The case against the Ranked Probability Score." pith.science (2026). https://pith.science/paper/5YAHQCYX
@misc{pith2026190808980,
author = {Pith},
title = {Pith review of: Evaluating probabilistic forecasts of football matches: The case against the Ranked Probability Score},
year = {2026},
howpublished = {\url{https://pith.science/paper/5YAHQCYX}},
note = {Machine review of arXiv:1908.08980}
}
read the original abstract
A scoring rule is a function of a probabilistic forecast and a corresponding outcome that is used to evaluate forecast performance. A wide range of scoring rules have been defined over time and there is some debate as to which are the most appropriate for evaluating the performance of forecasts of sporting events. This paper focuses on forecasts of the outcomes of football matches. The ranked probability score (RPS) is often recommended since it is `sensitive to distance', that is it takes into account the ordering in the outcomes (a home win is `closer' to a draw than it is to an away win, for example). In this paper, this reasoning is disputed on the basis that it adds nothing in terms of the actual aims of using scoring rules. A related property of scoring rules is locality. A scoring rule is local if it only takes the probability placed on the outcome into consideration. Two simulation experiments are carried out in the context of football matches to compare the performance of the RPS, which is non-local and sensitive to distance, the Brier score, which is non-local and insensitive to distance, and the ignorance score, which is local and insensitive to distance. The ignorance score is found to outperform both the RPS and the Brier score, casting doubt on the value of non-locality and sensitivity to distance as properties of scoring rules in this context.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Baboota, R., and H. Kaur, Predictive analysis and modelling football results using machine learning approach for english premier league, International Journal of Forecasting, 35 (2), 741–755, 2019
work page 2019
-
[2]
M., Expected information as expected utility, the Annals of Statistics, pp
Bernardo, J. M., Expected information as expected utility, the Annals of Statistics, pp. 686–690, 1979
work page 1979
-
[3]
Brier, G. W., Verification of forecasts expressed in terms of probability, Monthly weather review, 78 (1), 1–3, 1950. Br¨ ocker, J., and L. A. Smith, Scoring probabilistic forecasts: The importance of being proper, Weather and Forecasting, 22 (2), 382–388, 2007
work page 1950
- [4]
-
[5]
Diniz, M. A., R. Izbicki, D. Lopes, and L. E. Salasar, Comparing probabilistic predictive models applied to football, Journal of the Operational Research Society, 70 (5), 770–782, 2019
work page 2019
-
[6]
Epstein, E. S., A scoring system for probability forecasts of ranked categories, Journal of Applied Meteorology, 8 (6), 985–987, 1969
work page 1969
-
[7]
Forrest, D., J. Goddard, and R. Simmons, Odds-setters as forecasters: The case of english football, International journal of forecasting , 21 (3), 551– 564, 2005
work page 2005
-
[8]
Friedman, D., Effective scoring rules for probabilistic forecasts, Management Science, 29 (4), 447–454, 1983
work page 1983
Show all 23 references
-
[9]
Gneiting, T., and A. E. Raftery, Strictly proper scoring rules, prediction, and estimation, Journal of the American Statistical Association, 102 (477), 359–378, 2007
2007
-
[10]
J., Rational decisions, in Breakthroughs in statistics , pp
Good, I. J., Rational decisions, in Breakthroughs in statistics , pp. 365–377, Springer, 1992. 27
1992
-
[11]
Jose, V. R. R., R. F. Nau, and R. L. Winkler, Sensitivity to distance and baseline distributions in forecast evaluation, Management Science, 55 (4), 582–590, 2009
2009
-
[12]
Judd, K., and L. A. Smith, Indistinguishable states i: The perfect model scenario, Physica D: nonlinear phenomena , 151 (2-4), 125–141, 2001
2001
-
[13]
Judd, K., and L. A. Smith, Indistinguishable states ii: The imperfect model scenario, Physica D: nonlinear phenomena , 196 (3-4), 224–242, 2004
2004
-
[14]
J., and R
Koopman, S. J., and R. Lit, Forecasting football match results in national league competitions using score-driven time series models, International Journal of Forecasting, 35 (2), 797–809, 2019
2019
-
[15]
thesis, London School of Economics and Political Science, 2016
Maynard, T., Extreme insurance and the dynamics of risk, Ph.D. thesis, London School of Economics and Political Science, 2016
2016
-
[16]
H., The ranked probability score and the probability score: A comparison, weather, 81, 82, 1970
Murphy, A. H., The ranked probability score and the probability score: A comparison, weather, 81, 82, 1970
1970
-
[17]
Parry, M., A. P. Dawid, S. Lauritzen, et al., Proper local scoring rules, The Annals of Statistics , 40 (1), 561–592, 2012
2012
-
[18]
S., and L
Roulston, M. S., and L. A. Smith, Evaluating probabilistic forecasts using information theory, Monthly Weather Review, 130 (6), 1653–1660, 2002
2002
-
[19]
Groll, and G
Schauberger, G., A. Groll, and G. Tutz, Modeling football results in the german bundesliga using match-specific covariates, 2016
2016
-
[20]
Strobel, and H
Schmidt, C., M. Strobel, and H. O. Volkland, Accuracy, certainty and sur- prise: a prediction market on the outcome of the 2002 fifa world cup, 2008
2002
-
[21]
Selten, R., Axiomatic characterization of the quadratic scoring rule, Experi- mental Economics, 1 (1), 43–61, 1998
1998
-
[22]
Ng, One match to go!, Significance, 6 (4), 151– 153, 2009
Spiegelhalter, D., and Y.-L. Ng, One match to go!, Significance, 6 (4), 151– 153, 2009. Sta¨ el von Holstein, C.-A. S., A family of strictly proper scoring rules which are sensitive to distance, Journal of Applied Meteorology , 9 (3), 360–364, 1970. 28
2009
-
[23]
Wheatcroft, E., Interpreting the skill score form of forecast performance met- rics, International Journal of Forecasting, 35 (2), 573–579, 2019. 29
2019
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.