{"id":"1da7dc4f-b20c-40da-ab13-774473811a2f","arxiv_id":"1908.08980","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Simulations with real bookmaker odds show the ignorance score identifies the correct forecast distribution faster than the RPS or Brier score, but the result follows from known likelihood-ratio optimality.","lead":"This paper argues that the ranked probability score, often used to evaluate football match forecasts, is outperformed by the simpler ignorance score in simulation experiments. A specialist reader may care because this challenges a widely used evaluation metric and could change how sports forecasting models are compared.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The simulations in §§6–7 guarantee the ignorance score's superiority by construction: selecting the lower mean log score is a likelihood-ratio test, so the empirical results are a theorem of Neyman–Pearson theory, not evidence for the general recommendation.","rationale":"The reader's weakest assumption is exactly the point that the experiments reduce to a two-hypothesis likelihood-ratio test. My independent reading of §§5–8 confirms this and adds that the paper's own limitation statement makes the overreach explicit. The central claim is not that the ignorance score is better for ranking imperfect forecasts or for decision-making under distance-sensitive losses; it is that non-locality and distance-sensitivity have no value in football forecast evaluation. The simulations cannot establish this because the evaluation criterion is the probability of selecting the data-generating distribution, a criterion for which the log score is provably optimal. The correct response is to keep the paper as a conditional accept: the critique of Constantinou and Fenton's single-match examples is useful and the perfect-model results are a clean illustration of likelihood-ratio optimality, but the recommendation must be tempered and the imperfect-model scenario must be addressed before the broader claim is accepted. No verdict change from the reader's conditional is needed.","tokens_in":13311,"tokens_out":8771,"duration_ms":96339,"concrete_test":"Compute the §6 selection probability for the ignorance score not by simulation but from the likelihood-ratio statistic: for each match in Table 1, P_select(n) = E[ 1{ Σ_{i=1}^n log[p_{D_i}(y_i)/p_{A_i}(y_i)] > 0 } ], with D_i drawn uniformly from {α,β} and y_i drawn from D_i. Evaluate this exactly or by a large Monte Carlo run and overlay it on Figures 3–4. If the curves coincide, the reported outperformance is the Neyman–Pearson theorem, and the §8 recommendation cannot be inferred from these experiments.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central recommendation in §8 ('strongly recommend the ignorance score') is supported by the claim that the experiments in §§6–7 show the ignorance score outperforms the RPS and Brier in identifying a perfect forecasting system. That claim is true but uninformative, because the experiments are designed so that the log score is the optimal decision rule. In both experiments, each match's outcome is generated from one of two known candidate distributions P_i (the 'perfect' system) and Q_i (the 'imperfect' system). For any score S, the rule 'choose the system with the lower mean score' is equivalent to using the statistic (1/n)Σ[S(Q_i,y_i)-S(P_i,y_i)] with threshold 0. For the ignorance score S=-log p(y), this statistic is (1/n)Σ log[p_{P_i}(y_i)/p_{Q_i}(y_i)], the mean log-likelihood ratio. In §6 the true distribution is chosen from {α,β} with equal probability, so the threshold-zero rule is the Bayes rule and maximizes the chance of selecting the perfect system; in §7 the same likelihood-ratio comparison is applied to each pair (P_i,Q_i). Thus the observed advantage is a mathematical consequence of Neyman–Pearson optimality, not an empirical discovery about football forecasts. The paper itself acknowledges in §5 that the practically relevant imperfect-model scenario is left as future work, yet §8 issues a strong recommendation. The conceptual argument in §8 that probabilities on non-outcomes are irrelevant is asserted, not established, so the recommendation rests on a theorem rather than on empirical evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper challenges the widespread use of the ranked probability score (RPS) for evaluating probabilistic football forecasts. After defining the RPS, Brier, and ignorance scores and reviewing properties such as propriety, locality, and sensitivity to distance, the author argues that the reasoning of Constantinou and Fenton for preferring the RPS is flawed because a single observed outcome cannot determine which of two forecasts is better without knowing the data-generating distribution. Two simulation experiments are then used to compare the three scores' ability to identify a 'perfect' forecasting system against an 'imperfect' one: the first uses the five Constantinou–Fenton match examples, the second uses pairs of forecasts constructed from bookmaker odds. In both experiments the ignorance score selects the perfect system more often than the RPS or Brier score, and the paper concludes with a strong recommendation to use the ignorance score.","tokens_in":13592,"tokens_out":6844,"duration_ms":72169,"significance":"If the conclusions were supported, the paper would be a useful contribution to the ongoing debate about scoring rules for ordered categorical outcomes, with practical implications for sports forecasting and model comparison. The paper has clear strengths: it identifies a real weakness in judging forecasts by a single outcome, it carefully defines the scoring rules and their properties, and the simulations are transparent and reproducible in principle. However, the central empirical claim is not informative in the way the paper presents it, because the experiments are simple hypothesis tests in which the ignorance score is the likelihood-ratio statistic and is therefore guaranteed to be optimal by well-known statistical theory. The paper's broad recommendation extends beyond the perfect-model scenario that it actually tests. The contribution would be strengthened by citing the relevant optimality theory and by either softening or re-deriving the practical recommendations.","major_comments":[{"comment":"The simulation design makes the ignorance score's superiority a mathematical identity rather than an empirical discovery. In both experiments the outcome for each match is generated from one of two known candidate distributions, the 'perfect' system P_i and the 'imperfect' system Q_i, and the decision rule is to select the system with the lower mean score. For the ignorance score, the difference in mean scores is (1/n) Σ log[p_{P_i}(y_i)/p_{Q_i}(y_i)], the mean log-likelihood ratio. In §6 the true distribution is chosen with equal prior probability, so the zero threshold is the Bayes decision rule and maximizes the probability of selecting the perfect system by the Neyman–Pearson lemma; in §7 the same pairwise likelihood-ratio comparison is used. The Brier and RPS were not expected to beat this rule, and the empirical results illustrate a theorem rather than provide evidence about football forecasts. The paper should either present the result as a direct consequence of likelihood-ratio optimality and cite the relevant theory, or redesign the experiment so that the decision problem is not a simple hypothesis test.","section":"§6–§7, Figs. 3–7"},{"comment":"The strong recommendation in §8 ('we strongly recommend the ignorance score') and the abstract's claim that the results 'cast doubt on the value of non-locality and sensitivity to distance' go beyond the evidence. Section 5 explicitly says that the imperfect-model scenario, which is the practically relevant one, 'is left as future work,' and all experiments only compare a perfect system with a single imperfect alternative. The simulations therefore cannot discriminate between scores for the practically relevant task of choosing among several imperfect forecast systems, and they do not establish that non-locality or sensitivity to distance is valueless. At most, the simulations show that non-local scores are not needed in a simple-hypothesis perfect-model selection problem. The recommendation and abstract should be tempered accordingly.","section":"§5 and §8, Abstract"},{"comment":"The paper uses 'efficiency in selecting the perfect model in finite samples' as the criterion for comparing scoring rules, but it never justifies why this decision-theoretic quantity is the appropriate criterion for choosing a score for practical forecast evaluation. Propriety already ensures that each score prefers the perfect system in expectation; the finite-sample selection probability depends on the test statistic and the class of alternatives. Without a link between selection probability and the forecast user's loss function, the experiments measure a property that may not correspond to the aims described in §4. The author should either provide such a justification or present the finite-sample results as an illustration of a known property of the log score rather than as a basis for a general recommendation.","section":"§5 and §7.1"},{"comment":"The conceptual argument that probabilities placed on non-outcomes are irrelevant is asserted rather than established. The discussion says that knowing the outcome reveals little about the true probabilities of other outcomes, but this does not imply that a score should ignore the rest of the forecast distribution; proper scores can use the full distribution and still rank imperfect forecasts differently. The paper should either supply a decision-theoretic argument for why only the probability on the realized outcome should matter, or explicitly acknowledge that this is a normative position rather than a consequence of the experiments.","section":"§8"}],"minor_comments":[{"comment":"The sentence 'The results for matches one to four are shown in figure 3' should refer to Figure 4, not Figure 3.","section":"§6.1"},{"comment":"References to 'table 3' for the Constantinou–Fenton match examples and the colour scheme are inconsistent with the table numbering; the match examples are in Table 1 and the colour scheme is in Table 2.","section":"§3, §6"},{"comment":"The sentence 'for the lowest level of imperfection, in which δ = 0.1' contradicts the definition of δ and the later classification of δ = 0.01 and 0.025 as the two lowest levels; this appears to be a typographical error that should be corrected.","section":"§7.1"},{"comment":"There are several typographical slips, including 'An scoring rule' in §2.2 and 'truely' in §5, and the typesetting of 'Staël von Holstein' should be fixed.","section":"§2.2, §5"},{"comment":"The statement that the ignorance score is 'the only local and proper scoring rule' should be qualified in the standard way, typically 'up to affine transformation,' to avoid overclaiming.","section":"§1"}],"recommendation":"major_revision","confidential_remarks":"The paper is well written and the critique of single-outcome forecast comparisons is sound, but the empirical core is essentially an illustration of Neyman–Pearson optimality. The editor may wish to ensure that the revised version either engages directly with that theory or substantially narrows its claims. The strong recommendation in §8 should be reconsidered in light of the paper's own admission that the imperfect-model scenario is left for future work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, the conceptual attack on Constantinou and Fenton's single-match examples is solid and worth reading. The author shows that whether forecast α or β is better depends on the unknown underlying distribution, and the colored simplex plots make that point clearly. The match-five lay-bet argument is also neatly dispatched: you should evaluate the binary forecast derived from the ternary one, not a score that quietly rewards distance. So the paper does real work in clearing away a popular intuition.\n\nSecond, the simulation evidence is much weaker than the abstract and Section 8 suggest. Both experiments are two-hypothesis selection tasks: outcomes are drawn from either the perfect or imperfect system, and the decision rule is 'choose the system with the lower mean score.' For the ignorance score, that statistic is exactly the mean log-likelihood ratio, so the threshold-zero rule is the Neyman-Pearson optimal test. The observed advantage is a theorem, not a discovery about football forecasts. The paper never acknowledges this, and it cites neither Neyman-Pearson nor the optimality of the log score in simple hypothesis testing. That is the load-bearing flaw.\n\nThe author does flag, in Section 5, that the imperfect-model scenario is left as future work. That is honest. But Section 8 then says 'we strongly recommend the ignorance score for this purpose.' That recommendation depends on the conceptual claim that probabilities on non-outcomes are irrelevant, which is asserted rather than established. The simulations don't test that claim; they only re-derive likelihood-ratio optimality. The paper also has minor gaps: some curves lack sampling error bars, and the handling of candidate pairs in Experiment 2 is underspecified. These are fixable.\n\nWho gets value from this? Anyone working on football forecast evaluation or on scoring rules for ordered categories. The rebuttal of C&F is worth engaging with even if you don't accept the locality argument. The paper deserves serious peer review, but only after the claims are tempered, the likelihood-ratio connection is acknowledged, and the imperfect-model case is either studied or the recommendation is scaled back to 'deserves further investigation.'\n\nMy recommendation: send it to a good referee. It is not a major result, but it is a legitimate and clearly written contribution to a live debate. With revision, it could be a solid applied paper. Don't desk-reject it.","headline":"A well-written and useful critique of the RPS, but its central simulation result is a likelihood-ratio theorem, not empirical evidence; the strong recommendation outruns the evidence.","tokens_in":14129,"tokens_out":1661,"would_cite":true,"duration_ms":19855,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that the ranked probability score, despite its popularity for evaluating football forecasts, is less effective than the ignorance score at identifying the better forecasting system, and recommends the ignorance score for…","keywords":["probabilistic forecasting","scoring rules","ranked probability score","ignorance score","Brier score","football match forecasting","locality","sensitivity to distance"],"falsifier":"Find a pair of real forecasting systems where the RPS ranks the system that is measurably better out-of-sample, such as higher betting profit, while the ignorance score ranks the worse one; if such cases are common, the paper's recommendation fails in the imperfect-model context it claims to address. More directly, re-run the perfect-model experiment with a set of multiple imperfect alternatives and a non-local score designed to exploit differences in the non-outcome probabilities; if that score identifies the perfect system faster than the ignorance score, the claim that extra distributional information is useless is false.","tokens_in":63,"feed_emoji":"⚽","tokens_out":8607,"duration_ms":122411,"temperature":0.7,"pith_summary":"This paper argues that the ranked probability score (RPS), despite its popularity for evaluating football match forecasts, is a worse choice than the ignorance score for selecting the better forecasting system. It disputes the reasoning that rewarding probability placed on outcomes close to the actual result is beneficial, contending that only the probability placed on the observed outcome should count. Two simulation experiments, one using five constructed forecast pairs and one using 39,343 real bookmaker-odds matches, compare how quickly the RPS, Brier score, and ignorance score identify a perfect forecasting system against an imperfect one. In both, the ignorance score is at least as good and often better, and it is the only proper local scoring rule. The paper recommends the ignorance score for football forecast evaluation.","feed_headline":"Ignorance score beats the RPS at picking perfect football forecasts","feed_subtitle":"The log score identifies better forecasts faster than the distance-sensitive RPS in two football simulation experiments.","key_machinery":"The machinery is a controlled comparison of three proper scoring rules under the perfect-model scenario: the Brier score, the ranked probability score (which compares cumulative forecast distributions over ordered outcomes), and the ignorance score (the negative base-2 logarithm of the probability assigned to the occurred outcome). Each scoring rule is judged by the probability that its mean score over $n$ matches selects the perfect forecasting system over a single imperfect alternative. The concept of locality, whether a score uses only the probability at the occurred outcome, is the key discriminator; the paper argues that since the true probabilities of non-outcomes are unknowable even after the result, only the outcome probability is legitimately usable.","core_discovery":"The paper's central claim is that sensitivity to distance and non-locality, the properties that motivate the RPS, provide no benefit in the practical task of comparing football forecasting systems. The author defines that task as selecting the forecasting system closer to the data-generating distribution; in a perfect-model experiment, the right score should require fewer observed matches to identify the perfect system. The ignorance score, the negative log probability of the observed outcome, outperforms both the non-local Brier score and the distance-sensitive RPS, with the RPS notably inefficient when the two candidate forecasts differ only slightly. The conclusion is that the RPS should not be recommended over the ignorance score for football forecasts.","pith_inferences":["Because the ignorance score is the logarithm of the forecast likelihood, the perfect-model selection experiment is a likelihood-ratio test, and its dominance over the other scores in that restricted setting is guaranteed by standard hypothesis-testing theory; the empirical result is therefore a demonstration of that theorem rather than a football-specific discovery.","In the more realistic imperfect-model scenario, the best score may depend on the user's loss function; a bettor facing asymmetric payoffs should test forecast systems on betting profit rather than on a generic score, and the paper's comparison does not settle that choice.","A natural extension is to test whether an outcome-dependent non-local score that uses the full forecast distribution to assess calibration can beat the ignorance score in imperfect-model settings where forecast types differ, such as overconfidence versus bias.","The same simulation protocol could be applied to other ordered discrete sports outcomes, such as margin-of-victory bins, to see whether distance sensitivity has any value outside football's three-outcome space."],"forward_implications":["Forecast evaluation studies that currently report only the RPS for football should also report the ignorance score, since the two can rank imperfect forecast systems differently.","The difference in mean ignorance between two forecasting systems has a direct bits-per-match interpretation, making the size of a skill difference meaningful in a way the RPS and Brier score do not offer.","If a forecast is intended for a specific binary decision, such as a lay bet on the away win, the derived binary forecast should be evaluated on its own rather than through a single three-outcome score.","The results imply that adding distance sensitivity to a score does not help identify better forecasts more quickly, contrary to the reasoning that motivated adoption of the RPS."],"supporting_citations":[{"why":"Supplies the five hypothetical forecast pairs and the distance-sensitivity argument that the paper rebuts.","marker":"Constantinou and Fenton [2012]"},{"why":"Defines the Brier score used as one of the three scoring rules.","marker":"Brier [1950]"},{"why":"Defines the ranked probability score, the central target of the paper.","marker":"Epstein [1969]"},{"why":"Introduces the logarithmic ignorance score as an information-theoretic measure.","marker":"Good [1992]"},{"why":"Brings the ignorance score into forecast evaluation and links it to information theory.","marker":"Roulston and Smith [2002]"},{"why":"Establishes that the ignorance score is the only proper and local scoring rule, a key premise of the recommendation.","marker":"Bernardo [1979]"},{"why":"Represents the school of thought that recommends distance-sensitive scoring rules, which the paper's experiments challenge.","marker":"Jose et al. [2009]"},{"why":"Defines the perfect-model scenario used to frame the selection experiments.","marker":"Judd and Smith [2001]"},{"why":"Defines the imperfect-model scenario, which the paper acknowledges as the realistic setting left for future work.","marker":"Judd and Smith [2004]"}],"fun_headline_variants":["Log score dethrones distance-sensitive RPS for football forecasts","Ignorance score finds perfect football forecast fastest","Distance sensitivity adds nothing in football forecast scoring","Local log score beats non-local RPS in football simulations"],"cache_read_input_tokens":16256,"weakest_assumption_plain":"The whole comparison rests on the assumption that better means more likely to pick the perfect forecast when the only alternative is one fixed imperfect forecast; since the ignorance score is the log-likelihood, that test is mathematically guaranteed to win, so the paper's headline result is built into the experimental setup.","fun_headline_variants_meta":{"raw":{"variants":["Log score dethrones distance-sensitive RPS for football forecasts","Ignorance score finds perfect football forecast fastest","Distance sensitivity adds nothing in football forecast scoring","Local log score beats non-local RPS in football simulations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000574,"raw_usage":{"total_tokens":2700,"prompt_tokens":921,"completion_tokens":1779,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":537,"completion_tokens_details":{"reasoning_tokens":1716}},"tokens_in":537,"tokens_out":1779,"duration_ms":14029,"temperature":1.0,"reasoning_tokens":1716,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:24:45.441473+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Find a pair of real forecasting systems where the RPS ranks the system that is measurably better out-of-sample, such as higher betting profit, while the ignorance score ranks the worse one; if such cases are common, the paper's recommendation fails in the imperfect-model context it claims to address. More directly, re-run the perfect-model experiment with a set of multiple imperfect alternatives and a non-local score designed to exploit differences in the non-outcome probabilities; if that score identifies the perfect system faster than the ignorance score, the claim that extra distributional information is useless is false.","supporting_citations":[{"cited_title":"C., and N","cited_arxiv_id":null,"evidence_quote":"Supplies the five hypothetical forecast pairs and the distance-sensitivity argument that the paper rebuts."},{"cited_title":"W., Veriﬁcation of forecasts expressed in terms of probability, Monthly weather review, 78 (1), 1–3, 1950","cited_arxiv_id":null,"evidence_quote":"Defines the Brier score used as one of the three scoring rules."},{"cited_title":"S., A scoring system for probability forecasts of ranked categories, Journal of Applied Meteorology, 8 (6), 985–987, 1969","cited_arxiv_id":null,"evidence_quote":"Defines the ranked probability score, the central target of the paper."},{"cited_title":"J., Rational decisions, in Breakthroughs in statistics , pp","cited_arxiv_id":null,"evidence_quote":"Introduces the logarithmic ignorance score as an information-theoretic measure."},{"cited_title":"S., and L","cited_arxiv_id":null,"evidence_quote":"Brings the ignorance score into forecast evaluation and links it to information theory."},{"cited_title":"M., Expected information as expected utility, the Annals of Statistics, pp","cited_arxiv_id":null,"evidence_quote":"Establishes that the ignorance score is the only proper and local scoring rule, a key premise of the recommendation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Represents the school of thought that recommends distance-sensitive scoring rules, which the paper's experiments challenge."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the perfect-model scenario used to frame the selection experiments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the imperfect-model scenario, which the paper acknowledges as the realistic setting left for future work."}],"review_version":1}