Pith. sign in

REVIEW 4 major objections 5 minor 23 references

Evaluating probabilistic forecasts of football matches: The case against the Ranked Probability Score

T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read This paper argues that the ranked probability score, despite its popularity for evaluating football forecasts, is less effective than the ignorance score at identifying the better forecasting system, and recommends the ignorance score for…

desk verdict A well-written and useful critique of the RPS, but its central simulation result is a likelihood-ratio theorem, not empirical evidence; the strong recommendation outruns the evidence. read the letter →

arxiv 1908.08980 v1 pith:5YAHQCYX submitted 2019-08-23 stat.AP

classification stat.AP
keywords probabilisticforecastingscoringrulesrankedprobabilityscoreignoranceBrierfootballmatchlocalitysensitivitytodistance
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the ranked probability score (RPS), despite its popularity for evaluating football match forecasts, is a worse choice than the ignorance score for selecting the better forecasting system. It disputes the reasoning that rewarding probability placed on outcomes close to the actual result is beneficial, contending that only the probability placed on the observed outcome should count. Two simulation experiments, one using five constructed forecast pairs and one using 39,343 real bookmaker-odds matches, compare how quickly the RPS, Brier score, and ignorance score identify a perfect forecasting system against an imperfect one. In both, the ignorance score is at least as good and often better, and it is the only proper local scoring rule. The paper recommends the ignorance score for football forecast evaluation.

What carries the argument

The machinery is a controlled comparison of three proper scoring rules under the perfect-model scenario: the Brier score, the ranked probability score (which compares cumulative forecast distributions over ordered outcomes), and the ignorance score (the negative base-2 logarithm of the probability assigned to the occurred outcome). Each scoring rule is judged by the probability that its mean score over $n$ matches selects the perfect forecasting system over a single imperfect alternative. The concept of locality, whether a score uses only the probability at the occurred outcome, is the key discriminator; the paper argues that since the true probabilities of non-outcomes are unknowable even after the result, only the outcome probability is legitimately usable.

What would settle it

Find a pair of real forecasting systems where the RPS ranks the system that is measurably better out-of-sample, such as higher betting profit, while the ignorance score ranks the worse one; if such cases are common, the paper's recommendation fails in the imperfect-model context it claims to address. More directly, re-run the perfect-model experiment with a set of multiple imperfect alternatives and a non-local score designed to exploit differences in the non-outcome probabilities; if that score identifies the perfect system faster than the ignorance score, the claim that extra distributional information is useless is false.

Watch

Extended reading notes

Core claim

The paper's central claim is that sensitivity to distance and non-locality, the properties that motivate the RPS, provide no benefit in the practical task of comparing football forecasting systems. The author defines that task as selecting the forecasting system closer to the data-generating distribution; in a perfect-model experiment, the right score should require fewer observed matches to identify the perfect system. The ignorance score, the negative log probability of the observed outcome, outperforms both the non-local Brier score and the distance-sensitive RPS, with the RPS notably inefficient when the two candidate forecasts differ only slightly. The conclusion is that the RPS should not be recommended over the ignorance score for football forecasts.

Load-bearing premise

The whole comparison rests on the assumption that better means more likely to pick the perfect forecast when the only alternative is one fixed imperfect forecast; since the ignorance score is the log-likelihood, that test is mathematically guaranteed to win, so the paper's headline result is built into the experimental setup.

Editorial extensions

If this is right

  • Forecast evaluation studies that currently report only the RPS for football should also report the ignorance score, since the two can rank imperfect forecast systems differently.
  • The difference in mean ignorance between two forecasting systems has a direct bits-per-match interpretation, making the size of a skill difference meaningful in a way the RPS and Brier score do not offer.
  • If a forecast is intended for a specific binary decision, such as a lay bet on the away win, the derived binary forecast should be evaluated on its own rather than through a single three-outcome score.
  • The results imply that adding distance sensitivity to a score does not help identify better forecasts more quickly, contrary to the reasoning that motivated adoption of the RPS.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the ignorance score is the logarithm of the forecast likelihood, the perfect-model selection experiment is a likelihood-ratio test, and its dominance over the other scores in that restricted setting is guaranteed by standard hypothesis-testing theory; the empirical result is therefore a demonstration of that theorem rather than a football-specific discovery.
  • In the more realistic imperfect-model scenario, the best score may depend on the user's loss function; a bettor facing asymmetric payoffs should test forecast systems on betting profit rather than on a generic score, and the paper's comparison does not settle that choice.
  • A natural extension is to test whether an outcome-dependent non-local score that uses the full forecast distribution to assess calibration can beat the ignorance score in imperfect-model settings where forecast types differ, such as overconfidence versus bias.
  • The same simulation protocol could be applied to other ordered discrete sports outcomes, such as margin-of-victory bins, to see whether distance sensitivity has any value outside football's three-outcome space.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper challenges the widespread use of the ranked probability score (RPS) for evaluating probabilistic football forecasts. After defining the RPS, Brier, and ignorance scores and reviewing properties such as propriety, locality, and sensitivity to distance, the author argues that the reasoning of Constantinou and Fenton for preferring the RPS is flawed because a single observed outcome cannot determine which of two forecasts is better without knowing the data-generating distribution. Two simulation experiments are then used to compare the three scores' ability to identify a 'perfect' forecasting system against an 'imperfect' one: the first uses the five Constantinou–Fenton match examples, the second uses pairs of forecasts constructed from bookmaker odds. In both experiments the ignorance score selects the perfect system more often than the RPS or Brier score, and the paper concludes with a strong recommendation to use the ignorance score.

Significance. If the conclusions were supported, the paper would be a useful contribution to the ongoing debate about scoring rules for ordered categorical outcomes, with practical implications for sports forecasting and model comparison. The paper has clear strengths: it identifies a real weakness in judging forecasts by a single outcome, it carefully defines the scoring rules and their properties, and the simulations are transparent and reproducible in principle. However, the central empirical claim is not informative in the way the paper presents it, because the experiments are simple hypothesis tests in which the ignorance score is the likelihood-ratio statistic and is therefore guaranteed to be optimal by well-known statistical theory. The paper's broad recommendation extends beyond the perfect-model scenario that it actually tests. The contribution would be strengthened by citing the relevant optimality theory and by either softening or re-deriving the practical recommendations.

major comments (4)
  1. [§6–§7, Figs. 3–7] The simulation design makes the ignorance score's superiority a mathematical identity rather than an empirical discovery. In both experiments the outcome for each match is generated from one of two known candidate distributions, the 'perfect' system P_i and the 'imperfect' system Q_i, and the decision rule is to select the system with the lower mean score. For the ignorance score, the difference in mean scores is (1/n) Σ log[p_{P_i}(y_i)/p_{Q_i}(y_i)], the mean log-likelihood ratio. In §6 the true distribution is chosen with equal prior probability, so the zero threshold is the Bayes decision rule and maximizes the probability of selecting the perfect system by the Neyman–Pearson lemma; in §7 the same pairwise likelihood-ratio comparison is used. The Brier and RPS were not expected to beat this rule, and the empirical results illustrate a theorem rather than provide evidence about football forecasts. The paper should either present the result as a direct consequence of likelihood-ratio optimality and cite the relevant theory, or redesign the experiment so that the decision problem is not a simple hypothesis test.
  2. [§5 and §8, Abstract] The strong recommendation in §8 ('we strongly recommend the ignorance score') and the abstract's claim that the results 'cast doubt on the value of non-locality and sensitivity to distance' go beyond the evidence. Section 5 explicitly says that the imperfect-model scenario, which is the practically relevant one, 'is left as future work,' and all experiments only compare a perfect system with a single imperfect alternative. The simulations therefore cannot discriminate between scores for the practically relevant task of choosing among several imperfect forecast systems, and they do not establish that non-locality or sensitivity to distance is valueless. At most, the simulations show that non-local scores are not needed in a simple-hypothesis perfect-model selection problem. The recommendation and abstract should be tempered accordingly.
  3. [§5 and §7.1] The paper uses 'efficiency in selecting the perfect model in finite samples' as the criterion for comparing scoring rules, but it never justifies why this decision-theoretic quantity is the appropriate criterion for choosing a score for practical forecast evaluation. Propriety already ensures that each score prefers the perfect system in expectation; the finite-sample selection probability depends on the test statistic and the class of alternatives. Without a link between selection probability and the forecast user's loss function, the experiments measure a property that may not correspond to the aims described in §4. The author should either provide such a justification or present the finite-sample results as an illustration of a known property of the log score rather than as a basis for a general recommendation.
  4. [§8] The conceptual argument that probabilities placed on non-outcomes are irrelevant is asserted rather than established. The discussion says that knowing the outcome reveals little about the true probabilities of other outcomes, but this does not imply that a score should ignore the rest of the forecast distribution; proper scores can use the full distribution and still rank imperfect forecasts differently. The paper should either supply a decision-theoretic argument for why only the probability on the realized outcome should matter, or explicitly acknowledge that this is a normative position rather than a consequence of the experiments.
minor comments (5)
  1. [§6.1] The sentence 'The results for matches one to four are shown in figure 3' should refer to Figure 4, not Figure 3.
  2. [§3, §6] References to 'table 3' for the Constantinou–Fenton match examples and the colour scheme are inconsistent with the table numbering; the match examples are in Table 1 and the colour scheme is in Table 2.
  3. [§7.1] The sentence 'for the lowest level of imperfection, in which δ = 0.1' contradicts the definition of δ and the later classification of δ = 0.01 and 0.025 as the two lowest levels; this appears to be a typographical error that should be corrected.
  4. [§2.2, §5] There are several typographical slips, including 'An scoring rule' in §2.2 and 'truely' in §5, and the typesetting of 'Staël von Holstein' should be fixed.
  5. [§1] The statement that the ignorance score is 'the only local and proper scoring rule' should be qualified in the standard way, typically 'up to affine transformation,' to avoid overclaiming.

Circularity Check

1 steps flagged · score 6.0 of 10

The simulation 'evidence' for the ignorance score is the Neyman-Pearson lemma in disguise: the selection rule and the log-score definition make the experiment's outcome a theorem.

  1. other [Section 6 (Experiment one, selection protocol) and Section 8 (Discussion, recommendation)]
    "For a given pair of forecasts, define the outcomes of a series of n matches by drawing from forecast α or forecast β with equal probability 0.5. ... A scoring rule is defined to 'select' a forecasting system if it is assigned the lowest mean score over n forecasts."

    With this selection rule, the ignorance score's comparison is exactly the log-likelihood ratio: choose the perfect system iff (1/n)Σ[log p_i^P(y_i) − log p_i^Q(y_i)] > 0. In Experiment one the true distribution is α or β with equal probability, so threshold 0 is the Bayes decision rule; by Neyman-Pearson/Bayes decision theory this maximizes the probability of selecting the perfect system among all rules. The simulated 'outperformance' of the ignorance score is therefore a mathematical consequence of the definition of the score and the experimental protocol, not an empirical discovery about football forecasts. Experiment two uses the same pair-by-pair mean-score comparison, so the log-score rule remains the likelihood-ratio test.

full rationale

The paper's only load-bearing circular element is the interpretation of the simulations. The conceptual argument against the RPS (Section 3) and the locality argument (Section 8) are non-circular, though debatable. The self-citation to Wheatcroft (2019) in Section 4 is minor and not load-bearing. However, the headline empirical claim that the ignorance score outperforms the RPS and Brier in identifying a perfect forecasting system is not an independent result: the experiment defines 'selecting' a system as having the lower mean score, and for the ignorance score that decision rule is the likelihood-ratio test, which is optimal for the equal-prior problem in Experiment one. The simulation outcomes are therefore forced by construction, and presenting them as evidence for the recommendation (Section 8) is equivalent to citing an uncited theorem as empirical support. This warrants a score of 6: partial circularity, because the central empirical 'prediction' reduces by construction, while the philosophical arguments retain independent content.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The paper's conclusion is not derived from a physical law or a fitted model; it relies on a chosen simulation protocol. The key unstated premise is that 'speed of selecting a perfect system in a two-hypothesis test' is the right criterion for choosing a scoring rule. No new entities are introduced.

free parameters (1)
  • δ (imperfection threshold) = 0.01, 0.025, 0.05, 0.1
    Hand-chosen thresholds in experiment two that control the maximum L1 distance between the perfect and imperfect forecasts. The qualitative result holds across these values, but no principled justification is given for their range.
assumptions (4)
  • domain assumption The set of normalized bookmaker odds from football-data.co.uk is a representative distribution of realistic match outcome probabilities.
    Used in experiment two to generate true and imperfect forecasts; if these distributions are not representative, the results may not generalize.
  • domain assumption The probability of selecting a perfect forecasting system over a single imperfect alternative is a meaningful criterion for comparing scoring rules for practical forecast evaluation.
    The paper's experiments are based on this criterion; it is not the only or obviously most relevant criterion, and the paper leaves the imperfect-model scenario to future work.
  • standard math The ignorance score is the only proper and local scoring rule for discrete outcomes.
    Cited from Bernardo (1979) and Bröcker and Smith (2007); used to argue the uniqueness of the local proper score.
  • domain assumption Scoring rules used for evaluation should be assessed by how efficiently they identify a perfect forecasting system in finite samples.
    Section 5; this is the premise of the two experiments.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Evaluating probabilistic forecasts of football matches: The case against the Ranked Probability Score." pith.science (2026). https://pith.science/paper/5YAHQCYX

@misc{pith2026190808980,
  author       = {Pith},
  title        = {Pith review of: Evaluating probabilistic forecasts of football matches: The case against the Ranked Probability Score},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5YAHQCYX}},
  note         = {Machine review of arXiv:1908.08980}
}
read the original abstract

A scoring rule is a function of a probabilistic forecast and a corresponding outcome that is used to evaluate forecast performance. A wide range of scoring rules have been defined over time and there is some debate as to which are the most appropriate for evaluating the performance of forecasts of sporting events. This paper focuses on forecasts of the outcomes of football matches. The ranked probability score (RPS) is often recommended since it is `sensitive to distance', that is it takes into account the ordering in the outcomes (a home win is `closer' to a draw than it is to an away win, for example). In this paper, this reasoning is disputed on the basis that it adds nothing in terms of the actual aims of using scoring rules. A related property of scoring rules is locality. A scoring rule is local if it only takes the probability placed on the outcome into consideration. Two simulation experiments are carried out in the context of football matches to compare the performance of the RPS, which is non-local and sensitive to distance, the Brier score, which is non-local and insensitive to distance, and the ignorance score, which is local and insensitive to distance. The ignorance score is found to outperform both the RPS and the Brier score, casting doubt on the value of non-locality and sensitivity to distance as properties of scoring rules in this context.

Figures

Figures reproduced from arXiv: 1908.08980 by the authors.

Figure 1
Figure 1. Randomly chosen probability distributions of match five coloured [PITH_FULL_IMAGE:figures/full_fig_p011_1.png] view at source ↗
Figure 2
Figure 2. Randomly chosen probability distributions of matches one to four [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗
Figure 3
Figure 3. Probability of each scoring rule selecting the perfect forecasting [PITH_FULL_IMAGE:figures/full_fig_p019_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Probability of each scoring rule selecting the perfect forecasting [PITH_FULL_IMAGE:figures/full_fig_p020_4.png]
Figure 5
Figure 5. Figure 5: Examples of forecast pairs for different levels of imperfection. The [PITH_FULL_IMAGE:figures/full_fig_p022_5.png]
Figure 6
Figure 6. Figure 6: The proportion of cases in which the perfect forecasting system [PITH_FULL_IMAGE:figures/full_fig_p023_6.png]
Figure 7
Figure 7. Figure 7: Pairwise differences in the proportion of cases in which the perfect [PITH_FULL_IMAGE:figures/full_fig_p024_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

23 extracted references · 23 canonical work pages

  1. [1]

    Kaur, Predictive analysis and modelling football results using machine learning approach for english premier league, International Journal of Forecasting, 35 (2), 741–755, 2019

    Baboota, R., and H. Kaur, Predictive analysis and modelling football results using machine learning approach for english premier league, International Journal of Forecasting, 35 (2), 741–755, 2019

  2. [2]

    M., Expected information as expected utility, the Annals of Statistics, pp

    Bernardo, J. M., Expected information as expected utility, the Annals of Statistics, pp. 686–690, 1979

  3. [3]

    W., Verification of forecasts expressed in terms of probability, Monthly weather review, 78 (1), 1–3, 1950

    Brier, G. W., Verification of forecasts expressed in terms of probability, Monthly weather review, 78 (1), 1–3, 1950. Br¨ ocker, J., and L. A. Smith, Scoring probabilistic forecasts: The importance of being proper, Weather and Forecasting, 22 (2), 382–388, 2007

  4. [4]

    C., and N

    Constantinou, A. C., and N. E. Fenton, Solving the problem of inadequate scoring rules for assessing probabilistic football forecast models, Journal of Quantitative Analysis in Sports , 8 (1), 2012

  5. [5]

    Diniz, M. A., R. Izbicki, D. Lopes, and L. E. Salasar, Comparing probabilistic predictive models applied to football, Journal of the Operational Research Society, 70 (5), 770–782, 2019

  6. [6]

    S., A scoring system for probability forecasts of ranked categories, Journal of Applied Meteorology, 8 (6), 985–987, 1969

    Epstein, E. S., A scoring system for probability forecasts of ranked categories, Journal of Applied Meteorology, 8 (6), 985–987, 1969

  7. [7]

    Goddard, and R

    Forrest, D., J. Goddard, and R. Simmons, Odds-setters as forecasters: The case of english football, International journal of forecasting , 21 (3), 551– 564, 2005

  8. [8]

    Friedman, D., Effective scoring rules for probabilistic forecasts, Management Science, 29 (4), 447–454, 1983

Show all 23 references
  1. [9]

    Gneiting, T., and A. E. Raftery, Strictly proper scoring rules, prediction, and estimation, Journal of the American Statistical Association, 102 (477), 359–378, 2007

  2. [10]

    J., Rational decisions, in Breakthroughs in statistics , pp

    Good, I. J., Rational decisions, in Breakthroughs in statistics , pp. 365–377, Springer, 1992. 27

  3. [11]

    Jose, V. R. R., R. F. Nau, and R. L. Winkler, Sensitivity to distance and baseline distributions in forecast evaluation, Management Science, 55 (4), 582–590, 2009

  4. [12]

    Judd, K., and L. A. Smith, Indistinguishable states i: The perfect model scenario, Physica D: nonlinear phenomena , 151 (2-4), 125–141, 2001

  5. [13]

    Judd, K., and L. A. Smith, Indistinguishable states ii: The imperfect model scenario, Physica D: nonlinear phenomena , 196 (3-4), 224–242, 2004

  6. [14]

    J., and R

    Koopman, S. J., and R. Lit, Forecasting football match results in national league competitions using score-driven time series models, International Journal of Forecasting, 35 (2), 797–809, 2019

  7. [15]

    thesis, London School of Economics and Political Science, 2016

    Maynard, T., Extreme insurance and the dynamics of risk, Ph.D. thesis, London School of Economics and Political Science, 2016

  8. [16]

    H., The ranked probability score and the probability score: A comparison, weather, 81, 82, 1970

    Murphy, A. H., The ranked probability score and the probability score: A comparison, weather, 81, 82, 1970

  9. [17]

    Parry, M., A. P. Dawid, S. Lauritzen, et al., Proper local scoring rules, The Annals of Statistics , 40 (1), 561–592, 2012

  10. [18]

    S., and L

    Roulston, M. S., and L. A. Smith, Evaluating probabilistic forecasts using information theory, Monthly Weather Review, 130 (6), 1653–1660, 2002

  11. [19]

    Groll, and G

    Schauberger, G., A. Groll, and G. Tutz, Modeling football results in the german bundesliga using match-specific covariates, 2016

  12. [20]

    Strobel, and H

    Schmidt, C., M. Strobel, and H. O. Volkland, Accuracy, certainty and sur- prise: a prediction market on the outcome of the 2002 fifa world cup, 2008

  13. [21]

    Selten, R., Axiomatic characterization of the quadratic scoring rule, Experi- mental Economics, 1 (1), 43–61, 1998

  14. [22]

    Ng, One match to go!, Significance, 6 (4), 151– 153, 2009

    Spiegelhalter, D., and Y.-L. Ng, One match to go!, Significance, 6 (4), 151– 153, 2009. Sta¨ el von Holstein, C.-A. S., A family of strictly proper scoring rules which are sensitive to distance, Journal of Applied Meteorology , 9 (3), 360–364, 1970. 28

  15. [23]

    Wheatcroft, E., Interpreting the skill score form of forecast performance met- rics, International Journal of Forecasting, 35 (2), 573–579, 2019. 29

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.