REVIEW 3 major objections 5 minor 46 references
Bayesian model finds real dispersion in football team scoring
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 16:21 UTC pith:MHFUX4GB
load-bearing objection Solid, honest Bayesian extension of the Maher/Poisson goal model with a spike-and-slab on CMP dispersion; the machinery is sound, but the headline claim about systematic team-level non-equidispersion is not yet supported by the classifier's own noise floor. the 3 major comments →
Bayesian Conway-Maxwell-Poisson model with spike-and slab priors for dispersed count data with application to football scores
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Using a CMP likelihood with team-specific dispersion parameters ν_i and a spike-and-slab prior (ν_i = exp(Z_i η_i), where Z_i selects between Poisson and dispersed regimes), the model identifies systematic heterogeneity in EPL goal scoring. In the 2023/24 season, three over-dispersed teams (Fulham, Brighton, Wolves) and one under-dispersed team (Nottingham Forest) cross the 0.5 posterior-probability threshold. Across five seasons, the CMP-SAS model yields lower WAIC than the Poisson baseline in every season and lower Ignorance Scores in nearly every season and market (match outcome, over/under 2.5, goal difference). The authors interpret this as evidence that the Poisson assumption misses a
What carries the argument
The Conway-Maxwell-Poisson distribution, a two-parameter count distribution with a dispersion parameter ν (ν=1 recovers Poisson; ν<1 means over-dispersion; ν>1 means under-dispersion), paired with a spike-and-slab prior on each team's ν_i. The binary indicator Z_i selects between a point mass at ν_i=1 (Poisson regime) and a log-normal slab around 1. The likelihood is doubly intractable because of the CMP normalizing constant, so the MCMC uses an exchange algorithm with exact auxiliary draws from a rejection sampler, which cancels the constant and permits posterior sampling for the slab and the indicator.
Load-bearing premise
A single season of about 38 games per team is enough to detect mild-to-moderate dispersion, even though the paper's own simulations show detection rates below 50% for ν between 0.6 and 1.6.
What would settle it
Simulate 20-team seasons under a pure Poisson model with the same attack/defence structure, run the CMP-SAS MCMC on each, and count how often 3 or 4 teams cross the 0.5 posterior-probability threshold by chance. If the observed five-season rate matches this null rate, the detected dispersion is noise.
If this is right
- If team-level dispersion is real, Poisson-based goal models understate predictive uncertainty and mis-estimate attack parameters for dispersed teams; accounting for dispersion should sharpen match probabilities, especially for over/under and exact-score bets.
- The spike-and-slab mechanism provides a direct probabilistic test of whether a unit deviates from a baseline count model, transferable to any setting where equidispersion is the default assumption.
- The finding that non-equidispersion appears in only 3–4 teams per season suggests dispersion is not a league-wide property but a team-specific trait, potentially linked to playing style.
- The improvement in out-of-sample Ignorance Scores, though modest on average, can be large for specific matches involving highly dispersed teams (e.g., Fulham), so predictive gains concentrate where the Poisson model is most wrong.
- The under-dispersed examples (consistent goal totals) point to a behavioral interpretation: some teams play conservatively once leading, which a pure rate-based Poisson model cannot capture.
Where Pith is reading between the lines
- The paper's classification threshold of 0.5 ignores multiplicity: with a ~10% false-positive rate per team, a 20-team season implies roughly two spurious detections by chance, so the four teams flagged in 2023/24 may partly reflect noise rather than true heterogeneity.
- The simulation results show detection power below 50% for mild dispersion (ν between 0.6 and 1.6), meaning a single 38-game season is likely insufficient to reliably classify teams with modest departures; aggregating across seasons or more games per team would sharpen inference.
- A natural extension is to apply the same spike-and-slab CMP structure to other count data with unit-specific dispersion, such as player-level shot or injury counts, where baseline Poisson assumptions are common but rarely tested.
- The paper's admission that it under-predicts draws suggests the conditional independence assumption, not dispersion, is the main remaining flaw; combining CMP dispersion with a Dixon-Coles-style low-score correction could be a testable improvement.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a Bayesian Conway-Maxwell-Poisson model with a spike-and-slab prior on team-specific dispersion parameters for football score data. The model extends the standard Poisson goal model of Maher (1982) by treating each team's dispersion ν_i as either equal to 1 (equidispersed Poisson) or drawn from a log-normal slab (non-equidispersed CMP). Posterior inference uses a Metropolis-within-Gibbs sampler with the exchange algorithm and rejection sampling from Benson and Friel (2021) to handle the doubly intractable CMP likelihood. Simulation studies report recovery of extreme dispersion and WAIC comparisons against Poisson and a full CMP model. The application to five EPL seasons claims that several teams per season depart from equidispersion, with improved in-sample WAIC and out-of-sample Ignorance scores relative to Poisson. The paper provides reproducible code and data sources.
Significance. If the empirical claims were fully supported, the paper would offer a practically useful and interpretable way to identify unit-level dispersion in count data, with a natural football application and potential betting relevance. The methodological combination of CMP with spike-and-slab priors is a reasonable extension of existing machinery, and the MCMC implementation is sound. The simulation results convincingly show that extreme dispersion (ν ≤ 0.3 or ≥ 3.2) is recoverable, and the WAIC improvements in simulated extreme settings are clear. The paper also ships code and uses a published, independently established inference engine, which is a strength. However, the central applied claim of systematic team-level heterogeneity in EPL scoring rests on a classification rule whose false-positive rate is acknowledged to be about 10% per team, and the out-of-sample predictive gains are tiny and reported without uncertainty. Those gaps are load-bearing for the paper's headline conclusion.
major comments (3)
- [Section 5.3 and Table 1] The claim that football teams display systematic heterogeneity (Section 1) rests entirely on classifying teams as non-equidispersed when P(Z_i=1|Y)>0.5. Table 1 reports a false-positive rate of about 10% for equidispersed teams, so with 20 teams per season one expects about 2 false discoveries by chance. The paper reports 16 team-seasons above threshold across five seasons, versus about 10 expected under the global null, and provides no multiplicity correction or null calibration. For 2023/24, three of the four detected teams have posterior probabilities of 0.51, 0.53, and 0.58 (Table 3), exactly the region where the simulation's true-positive rates are near chance (TP ≤ 0.48 for ν_true in [0.6,1.6]). The paper's own simulation therefore shows that the classifier cannot reliably distinguish mild dispersion from noise at the sample sizes available (one season of ~38 games per team). This
- [Section 5.2.2 and Tables 5, 7] The out-of-sample Ignorance Score differences between CMP-SAS and Poisson are extremely small—e.g., 1.326 vs 1.327 for 2023/24 Outcome, 0.929 vs 0.974 for 2023/24 Over-Under 2.5, and 2.796 vs 2.801 for 2020/21 Goal Difference—and are reported as point estimates with no measure of uncertainty. The paper describes these as 'consistent' and 'systematic' improvements, but with 190 out-of-sample matches, differences of 0.001–0.03 in average IGN are plausibly within Monte Carlo and sampling noise. A paired comparison (e.g., per-match differences with standard errors or a Diebold-Mariano-type test) is needed to support the predictive-performance claim. Without it, the statement in Section 5.3 that the CMP-SAS model 'consistently outperforms the Poisson model in every season considered' is not established.
- [Section 5.2.1 and Table 6] The in-sample WAIC comparisons also lack uncertainty quantification. The WAIC advantage of CMP-SAS over Poisson is 30.3 for 2023/24 but only 5.8 for 2021/22 (Table 6). The latter is small relative to the effective number of parameters difference, and no estimate of Monte Carlo error from the likelihood estimator or MCMC chain is provided. Calling the WAIC evidence 'clear' (Section 5.3) is therefore not fully justified. Reporting standard errors or repeated computations with different posterior subsamples would strengthen the claim that the improvement is not noise.
minor comments (5)
- [Section 3.4, Eq. (13)] In the prior ratio for the joint α_i, η_i update, the denominator uses f_N(α_i | 0, σ_β^2) but should be σ_α^2. This appears to be a typo; the text and Appendix B use σ_α for α_i.
- [Section 3.4, Eq. (14)] The acceptance probability for β_i omits the prior ratio for β_N, although the Appendix version includes f_N(β_N | 0, σ_β^2). If the omission is intentional because the deterministic update cancels, this should be stated; otherwise it is an inconsistency.
- [Section 5.1, paragraph 2] The phrase '4 teams are significantly non-equidispersed' is stronger than the analysis supports. The paper's own threshold is P(Z_i=1|Y)>0.5, and the simulation section does not define 'significant' in a frequentist sense. Suggest 'classified as non-equidispersed'.
- [Section 4.1, Figure 3] The boxplot caption says each boxplot is formed by 50 posterior means; for clarity, state explicitly that these are 10 non-equidispersed teams × 5 replicates (and 10 equidispersed teams × 5 replicates), respectively.
- [Section 5.3] The discussion of Liverpool appearing in 3 of 5 seasons and Nottingham Forest in two is descriptive, but no attempt is made to quantify how often such persistence would occur under the null. This is related to the main calibration concern and could be briefly addressed.
Circularity Check
No significant circularity: the model is estimated from data and the self-citations supply independent computational tools, not the paper's conclusions.
full rationale
The derivation chain is self-contained in the relevant sense. The model is specified in Section 2.3 with the CMP likelihood (Eq. 1), team-level dispersion via ν_i = exp(Z_i η_i) (Eq. 5), and standard weakly informative priors; posterior inference in Section 3 uses the exchange algorithm and rejection sampler whose cancellation of normalizing constants is shown explicitly in Section 3.1 and Appendices A.1–A.3. The citations to Benson and Friel (2021) and Piancastelli et al. (2023) are self-citations through Friel, but they supply previously published computational tools and a literature pointer, not the paper's conclusion; the MCMC machinery is not used to define the team-dispersion estimand or to force posterior classifications. Simulations in Section 4 are generated-from-model checks, which is self-consistency rather than circularity. The empirical claims about EPL heterogeneity rest on posterior probabilities and on WAIC/IGN comparisons, including genuinely out-of-sample rolling forecasts (Section 5.2.2); these do not reduce by construction to fitted inputs. Concerns that the >0.5 classification rule may have a ~10% false-positive rate and no multiplicity correction are substantive statistical-evidence concerns, but they are not examples of a prediction being equivalent to its input by definition. Hence no circular step is identified.
Axiom & Free-Parameter Ledger
free parameters (3)
- σ_η (slab scale) =
1
- Classification threshold =
0.5 (and 0.4 as alternative)
- MCMC proposal scales (s_α,s_β,s_γ,s_η,ρ) =
0.10, 0.10, 0.08, 0.4, 0.85
axioms (5)
- standard math CMP likelihood can be sampled exactly via the Benson-Friel rejection sampler with geometric/Poisson envelopes
- domain assumption Conditional independence of home and away scores given team/attack parameters
- domain assumption Dispersion is a property of the attacking team and applies equally to home and away goals
- standard math Sum-to-zero constraint in Eq. (6) resolves attack/defence identifiability
- ad hoc to paper Beta(1,1) prior on p_i is non-informative
read the original abstract
Statistical modeling for goals scored in football is typically achieved using the Poisson distribution and its variants. Here we propose a Bayesian framework for modeling under- and over-dispersion in count data by combining the Conway-Maxwell-Poisson (CMP) likelihood with a spikeand-slab (SAS) prior on unit-specific dispersion parameters. The proposed methodology generalizes Poisson-based count data models by treating equidispersion as an explicit baseline, and offering probabilistic quantification of departures from this regime, while simultaneously estimating their magnitude. Posterior inference is performed through a tailored Metropolis-within-Gibbs sampler that handles the doubly-intractable likelihood and provides efficient posterior exploration. The new method is examined using simulated data to confirm its ability to capture non-equidispersion, and applied to English Premier League (EPL) data. Dispersion is modeled at the team level and linked to goal-scoring behavior, and allows for thresholding mechanisms to distinguish teams based on their posterior probability of non-equidispersion. The results reveal heterogeneities in team-specific dispersion in the EPL, and demonstrate improvements in both model fit and predictive performance with respect to the standard Poisson model.
Figures
Reference graph
Works this paper leans on
-
[1]
Dixon, M. J. and Coles, S. G. , title =. 1997 , volume =
1997
-
[2]
and Torelli, N
Egidi, L. and Torelli, N. , title =. 2021 , volume =
2021
-
[3]
1985 , volume =
Pollard, Richard , title =. 1985 , volume =
1985
-
[4]
Home advantage in football: A current review of an unsolved puzzle , author=
-
[5]
1986 , volume =
Pollard, Richard , title =. 1986 , volume =
1986
-
[6]
Nevill, A. M. and Holder, R. L. , title =. 1999 , volume =
1999
-
[7]
Courneya, K. S. and Carron, A. V. , title =. 1992 , volume =
1992
-
[8]
and Goffelt, Jeremy P
Guikema, Seth D. and Goffelt, Jeremy P. , title =. 2008 , volume =
2008
-
[9]
and Kadane, Joseph B
Shmueli, Galit and Minka, Thomas P. and Kadane, Joseph B. and Borle, Sharad and Boatwright, Peter , title =. 2005 , volume =
2005
-
[10]
Gareth and Titman, Andrew C
Ridall, P. Gareth and Titman, Andrew C. and Pettitt, Anthony N. , title =. 2025 , volume =
2025
-
[11]
and Silva, Ricardo and Edwards, Daniel and Kosmidis, Ioannis , title =
Whitaker, Gavin A. and Silva, Ricardo and Edwards, Daniel and Kosmidis, Ioannis , title =. 2021 , volume =
2021
-
[12]
, title =
Boshnakov, Georgi and Kharrat, Tarak and McHale, Ian G. , title =. 2017 , volume =
2017
-
[13]
Ruiz, Francisco J. R. and Perez-Cruz, Fernando , title =. 2014 , volume =
2014
-
[14]
Maher, M. J. , title =. 1982 , volume =
1982
-
[15]
McHale, I. G. and Scarf, P. A. , title =. 2011 , volume =
2011
-
[16]
2003 , volume =
Karlis, Dimitris and Ntzoufras, Ioannis , title =. 2003 , volume =
2003
-
[17]
2009 , volume =
Karlis, Dimitris and Ntzoufras, Ioannis , title =. 2009 , volume =
2009
-
[18]
and Blangiardo, M
Baio, G. and Blangiardo, M. , title =. 2010 , volume =
2010
-
[19]
Conway, R. W. and Maxwell, W. L. , title =. 1962 , volume =
1962
-
[20]
An efficient Markov chain Monte Carlo method for distributions with intractable normalising constants , journaltitle =
M. An efficient Markov chain Monte Carlo method for distributions with intractable normalising constants , journaltitle =. 2006 , volume =
2006
-
[21]
2012 , eprint=
MCMC for doubly-intractable distributions , author=. 2012 , eprint=
2012
-
[22]
Bayesian Analysis , number =
Alan Benson and Nial Friel , title =. Bayesian Analysis , number =. 2021 , doi =
2021
-
[23]
Epstein, E. S. , title =. 1969 , volume =
1969
-
[24]
Constantinou, A. C. and Fenton, N. E. , title =. 2012 , volume =
2012
-
[25]
Brier, G. W. Verification of forecasts expressed in terms of probability , journal =. 1950 , volume =
1950
-
[26]
On spike and slab empirical Bayes multiple testing , journaltitle =
Castillo, Isma. On spike and slab empirical Bayes multiple testing , journaltitle =. 2020 , volume =
2020
-
[27]
1951 , series =
Moroney, Maureen , title =. 1951 , series =
1951
-
[28]
2018 , volume =
Egidi, Luca and Pauli, Francesco and Torelli, Nicola , title =. 2018 , volume =
2018
-
[29]
The Annals of Statistics , number =
Gideon Schwarz , title =. The Annals of Statistics , number =. 1978 , doi =
1978
-
[30]
Evolution of soccer as a research topic , volume =
Kirkendall, Donald and Urbaniak, James , year =. Evolution of soccer as a research topic , volume =. Progress in Cardiovascular Diseases , doi =
-
[31]
Piancastelli, Luiza S. C. and Friel, Nial and Barreto-Souza, Wagner and Ombao, Hernando , title =. 2023 , volume =
2023
-
[32]
Watanabe, Sumio , title =. 2010 , volume =. doi:10.48550/arxiv.1004.2316 , url =
-
[33]
and Hwang, J
Gelman, A. and Hwang, J. and Vehtari, A. , title =. 2014 , volume =
2014
-
[34]
2021 , volume =
Wheatcroft, Edward , title =. 2021 , volume =
2021
-
[35]
2025 , volume =
Florez, Maria and Guindani, Michele and Vannucci, Marina , title =. 2025 , volume =
2025
-
[36]
Spiegelhalter, D. J. and Best, N. G. and Carlin, B. P. and Van Der Linde, A. , title =. 2002 , volume =
2002
-
[37]
Mitchell, T. J. and Beauchamp, J. J. , title =. 1988 , volume =
1988
-
[38]
George, E. I. and McCulloch, R. E. , title =. 1993 , volume =
1993
-
[39]
Approaches for Bayesian Variable Selection , volume =
George, Edward and McCulloch, Robert , year =. Approaches for Bayesian Variable Selection , volume =
-
[40]
and Rao, J
Ishwaran, H. and Rao, J. S. , title =. 2005 , volume =
2005
-
[41]
Handbook of Bayesian Variable Selection , isbn =
Tadesse, Mahlet and Vannucci, Marina , year =. Handbook of Bayesian Variable Selection , isbn =
-
[42]
and Haaf, Julia M
Rouder, Jeffrey N. and Haaf, Julia M. and Vandekerckhove, Joachim , title =. 2018 , volume =
2018
-
[43]
2009 , note =
Pang, Xiaodong and Gill, Jeff , title =. 2009 , note =
2009
-
[44]
Roberts, G. O. and Rosenthal, J. S. , title =. 2001 , volume =
2001
-
[45]
Roberts, G. O. and Gelman, A. and Gilks, W. R. , title =. 1997 , volume =
1997
-
[46]
and Roberts, Gareth O
Neal, Peter J. and Roberts, Gareth O. , title =. 2006 , volume =
2006
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.