Pith. sign in

REVIEW 2 major objections 4 minor 28 references

A Bayesian bivariate conditional Poisson regression for goal dependence in the English Premier League

T0 review · 2 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Home and away goals in the English Premier League are negatively associated after accounting for stadium attendance and fouls, and crowd presence is tied to home scoring but not away scoring.

desk verdict The negative dependence estimate is plausible but not yet identified; the paper needs a team-strength robustness check before I'd trust the headline claim. read the letter →

arxiv 2608.07168 v1 pith:F4PMJHFG submitted 2026-08-07 stat.AP stat.ME

classification stat.APstat.ME
keywords BayesianinferencebivariateconditionalPoissongoaldependencehomeadvantagestadiumattendanceposteriorpredictivechecksEnglishPremierLeague
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that, in the English Premier League, the number of goals scored by the home team and the away team are not independent but negatively associated once stadium attendance and fouls are accounted for. Using a Bayesian bivariate conditional Poisson regression on 1,140 matches across three seasons, the posterior for the dependence parameter $\phi$ is centered at $-0.107$ with its 95% credible interval entirely below zero. The paper also claims an asymmetry in crowd effects: attendance is positively associated with home scoring but shows no clear association with away scoring. This matters because standard bivariate Poisson models for football scores only permit non-negative correlation, so they cannot represent the tactical suppression of the opponent's scoring that a team protecting a lead may produce.

What carries the argument

The central object is the reparameterized bivariate conditional Poisson (BCP) distribution, which specifies one goal count marginally as Poisson($\lambda_1$) and the other conditionally as Poisson($\mu_2 e^{\phi y_1}$), with $\phi \in \mathbb{R}$ controlling the sign of dependence. This construction gives a closed-form correlation whose sign follows the sign of $(e^{\phi} - 1)$, allowing negative association, unlike the classic shared-component bivariate Poisson model, which forces covariance $\lambda_0 \geq 0$. The BCP regression places log-linear means on $\lambda_1$ and $\lambda_2$ with covariates log(attendance+1) and fouls suffered by each team, and the parameter $\phi$ carries the paper's main inferential weight.

What would settle it

Fit the same BCP regression with team-specific attack and defence random effects (or a ranking-strength covariate) on the same 1,140 matches. If the 95% credible interval for $\phi$ then straddles zero, the negative dependence conclusion is an artifact of omitted team quality rather than a match-level effect.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that the BCP H→A model fitted to the 2018-19, 2020-21, and 2023-24 Premier League seasons estimates $\phi = -0.107$ (95% credible interval $[-0.147, -0.066]$), a statistically significant negative dependence between home and away goals after adjusting for log attendance and fouls suffered. Because a one-goal increase in the home score multiplies the conditional expected away goals by $\exp(-0.107) \approx 0.899$, the association corresponds to roughly a 10% reduction in the conditional mean away scoring. The same model finds that log attendance has a positive effect on home goals (posterior mean 0.024, 95% CI [0.013, 0.034]) and an effect centered at zero on away goals, confirming an asymmetry consistent with home advantage. The paper interprets the negative dependence as a within-match suppressive effect, such as a team protecting a lead, while noting the estimate is an adjusted association, not a causal claim.

Load-bearing premise

The model has no team-strength terms, so the negative dependence $\phi$ is only a structural feature of goal dynamics if the residual association after controlling for attendance and fouls is not actually driven by mismatched team quality.

Editorial extensions

If this is right

  • Both BCP factorizations beat the independent Poisson model on ELPD-LOO (-3552.0 and -3552.1 versus -3565.0), so allowing dependence improves out-of-sample joint prediction.
  • Conditional ELPD prefers the home-to-away factorization (-1721.8 versus -1817.2), meaning home goals predict away goals better than the reverse, as a predictive factorization rather than a causal claim.
  • A one-goal increase in home score is associated with about a 10% reduction in the conditional expected number of away goals, holding modeled covariates fixed.
  • Posterior predictive checks show the BCP models reproduce the observed negative goal correlation and the home-loss proportion, while the independent Poisson model does not.
  • Attendance is positively associated with home goals (posterior mean 0.024, 95% CI [0.013, 0.034]) but not with away goals (95% CI [-0.011, 0.011]).

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the negative dependence is genuine, real-time prediction models for football should use joint score distributions that permit suppression rather than only positive-copula constructions.
  • Adding team-specific attack and defence random effects to the same data is a direct test: if the credible interval for $\phi$ then includes zero, the estimated negative dependence is largely an artifact of team-quality mismatch.
  • The attendance asymmetry predicts a concrete effect: in the two pandemic-affected seasons with low crowds, the home-goal attendance coefficient should shrink or vanish when estimated season-by-season.
  • The inconclusive foul effects suggest that granular match-event data, such as expected goals or cards, could separate attacking pressure from refereeing effects in the away-foul-to-home-goals pattern.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper proposes a Bayesian bivariate conditional Poisson (BCP) regression model for home and away goal counts in football, using a parameterization that allows both positive and negative dependence between the two counts. The model includes log attendance and fouls suffered as covariates, with separate coefficients for home and away scoring intensities. The authors fit the model to 1,140 English Premier League matches from the 2018–19, 2020–21, and 2023–24 seasons using Hamiltonian Monte Carlo in Stan with empirical Bayes priors. They compare two directional specifications (A→H and H→A) and a baseline independent Poisson model. The preferred H→A specification yields a posterior mean for the dependence parameter φ of −0.107 with a 95% credible interval [−0.147, −0.066], which the paper interprets as statistically significant negative dependence between home and away goals; attendance is positively associated with home goals but not away goals. The paper concludes that BCP models improve predictive performance over the independent Poisson model and better reproduce the observed goal correlation and home-loss proportion.

Significance. If the negative dependence estimate is trustworthy, the paper makes a worthwhile contribution by demonstrating that football goal counts can exhibit negative association after conditioning on few match-level covariates, going beyond the positive-only dependence of standard bivariate Poisson models. The BCP parameterization is analytically tractable (closed-form marginal moments and correlation) and the Bayesian implementation with HMC is clearly described. The paper is also transparent about several limitations, notably the absence of team-specific attack and defense strengths. However, the central empirical claim is conditional on the adequacy of a very simple covariate set, and the acknowledged omission of team strengths is directly load-bearing for the sign and significance of φ. The attendance asymmetry claim is likewise threatened by the near-collinearity of attendance with season and pandemic conditions. These identification concerns, unless addressed with robustness checks, undermine the strength of the paper's headline conclusions.

major comments (2)
  1. [§3, Eq. (2); §5] The model contains no team-specific attack or defense strength terms, as acknowledged in Section 5 ('the current specification does not explicitly account for team-specific attacking and defensive strengths'). In EPL data, strong teams frequently play at home against weak away teams, producing high home goal counts and low away goal counts, while the opposite pairing produces the reverse pattern. This between-match heterogeneity induces a negative association in the residuals even when there is no within-match dependence, so the negative posterior for φ in Table 2 may be an artifact of omitted team quality rather than evidence of match-level negative dependence. The central claim in Section 4 therefore requires a robustness check that adds team strength parameters (e.g., hierarchical attack and defense random effects) and re-examines whether the credible interval for φ still excludes zero.
  2. [§4.1; Table 2] The home-attendance coefficient β1,Att is identified almost entirely by the 2020–21 season: log(Attendance+1) is near zero in that season while taking values around 10.5 in the other two seasons. Consequently β1,Att functions largely as a season contrast rather than a within-season attendance effect, and the seasons also differ in team composition, substitution rules, and the broader scoring environment. The paper itself concedes in Section 4.1 that 'attendance, pandemic conditions, and season are closely intertwined in these data.' The claim that attendance is positively associated with home scoring needs a sensitivity analysis with season-specific intercepts or with the pandemic season excluded before it can be presented as a substantive finding.
minor comments (4)
  1. [§3, correlation formula] Please verify the printed formula for Corr(Y1,Y2). As typeset it appears dimensionally inconsistent with the stated covariance λ1λ2(e^φ−1) and variance λ2 + λ1λ2(e^φ−1)^2; the denominator should simplify to 1 + λ1(e^φ−1)^2 rather than the expression involving an exponential of λ1(e^φ−1)^2.
  2. [§4, model selection] The statement that 'results are robust to the prior choice as investigated by the authors but not reported here to save space' is not verifiable; please include a prior sensitivity analysis in an appendix or supplementary material, or remove the claim.
  3. [§5, Software and data availability] The Stan implementation is described as available 'upon request'; for a data-analysis paper, providing the code in a public repository and a reproduction script would substantially strengthen reproducibility.
  4. [§3.3] There is a minor formatting issue: 'W AIC' should be 'WAIC' in the running text, and the WAIC/LOO-CV comparison should report standard errors for the ELPD differences so that the claimed improvement over the independent Poisson model can be assessed.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the negative dependence estimate is a fitted posterior parameter, and the model is fully specified from the likelihood rather than defined in terms of its conclusions.

full rationale

The paper's derivation chain is self-contained: Eq. (2) defines the BCP likelihood from the stochastic representation in Eq. (1), Eq. (4) defines the posterior, and MCMC yields posterior summaries. The central result, the posterior mean of phi = -0.107 with 95% CI excluding zero, is an estimate from this posterior, not a prediction obtained by renaming a fitted input. The empirical Bayes prior uses MLE standard errors from the same data, but the prior is centered at zero and does not define phi in terms of the conclusion, so this is standard empirical Bayes rather than a circular reduction. The posterior predictive checks for the goal correlation are consistency checks that reuse the fitted phi; the paper does not present them as out-of-sample predictions, and the ELPD-LOO comparison is a legitimate predictive criterion. Self-citations to Piancastelli et al. (2023a, 2023b) supply the parameterization and moment formulas, but those formulas are written out in the paper and originate in Berkhout and Plug (2004), so the citations are not load-bearing. The acknowledged omission of team-specific attack and defense strengths (Section 5) is a potential confounding or robustness limitation, not a circularity in the derivation. The unstated prior-robustness check mentioned in Section 4 is a reporting gap, but it does not make any step reduce to its own input.

Assumptions & free parameters 7 free parameters · 5 assumptions · 0 invented entities

The central inference relies on the conditional Poisson structure, log-linear covariate assumptions, cross-match independence, and the omission of team strength terms. The ledger is dominated by fitted parameters: all regression coefficients and the dependence parameter are estimated from the data rather than derived from first principles, and the paper introduces no new latent entities.

free parameters (7)
  • beta_0,1 (home intercept) = 0.146
    Fitted via HMC in the BCP H to A model, reported in Table 2.
  • beta_1,Att (home attendance effect) = 0.024
    Fitted via HMC, reported in Table 2.
  • beta_1,AF (away fouls on home goals) = 0.012
    Fitted via HMC, reported in Table 2.
  • beta_0,2 (away intercept) = 0.232
    Fitted via HMC, reported in Table 2.
  • beta_2,Att (away attendance effect) = 0.000
    Fitted via HMC, reported in Table 2.
  • beta_2,HF (home fouls on away goals) = 0.007
    Fitted via HMC, reported in Table 2.
  • phi (dependence parameter) = -0.107
    Fitted via HMC, reported in Table 2; the central estimate that drives the negative dependence claim.
assumptions (5)
  • domain assumption Y1 follows a Poisson distribution and Y2 given Y1=y1 follows a Poisson distribution with mean µ2 exp(phi*y1).
    This is the defining structure of the BCP model, adopted from Berkhout and Plug (2004) and Piancastelli et al. (2023a), and is not derived in the paper.
  • domain assumption Matches are independent.
    The likelihood in Section 3 is a product over matches; no within-club or temporal correlation is modeled.
  • domain assumption Log-linear covariate specification: log lambda = beta0 + beta_Att log(Att+1) + beta_F FS.
    The mean structure in Section 3 assumes a log-linear relationship between the scoring intensities and the selected covariates.
  • domain assumption Team-specific attack and defense strengths are not confounding the dependence parameter.
    The model contains no team strength terms, so interpreting phi as match-level dependence requires that unmeasured team quality does not induce the residual negative correlation.
  • domain assumption Empirical Bayes priors centered at zero with standard deviations from frequentist standard errors.
    The prior specification in Section 3.1 uses the same data twice, through the MLEs that set prior scales, and is adopted rather than derived.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Bayesian bivariate conditional Poisson regression for goal dependence in the English Premier League." pith.science (2026). https://pith.science/paper/F4PMJHFG

@misc{pith2026260807168,
  author       = {Pith},
  title        = {Pith review of: A Bayesian bivariate conditional Poisson regression for goal dependence in the English Premier League},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/F4PMJHFG}},
  note         = {Machine review of arXiv:2608.07168}
}
read the original abstract

Understanding the relationship between home and away goal counts in football provides valuable insights into match-level dynamics. While the influence of home advantage is well-established, with historical records indicating roughly 50% of matches are won by home teams (versus about 30% by the away team), properly determining the joint goal distribution while accounting for key match factors remains under-explored. In this paper, we develop a Bayesian bivariate Conditional Poisson (BCP) regression model to explicitly capture the dependence between home and away goal counts, addressing a core limitation of traditional football scoring models that assume independence or only allow for positive correlation. The BCP regression is applied to the English Premier League (EPL) data spanning three seasons, incorporating stadium attendance and committed fouls by both teams as regressors. Inference is conducted within a Bayesian framework, enabling interpretable uncertainty quantification and model validation through posterior predictive checks. Our results reveal a negative correlation between home and away goal counts and highlight an asymmetry in how match attendance influences home versus away scoring.

Figures

Figures reproduced from arXiv: 2608.07168 by the authors.

Figure 1
Figure 1. Distributions of match outcomes (left) and home and away goals (right). [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Posterior intervals for covari￾ate effects. Thick bars represent 50% intervals, thin bars 90% intervals, and points indicate medians. 11 [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗
Figure 3
Figure 3. Posterior predictive densities for the marginal means and variances of goals. Teams [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: The observed dependency between goals is not captured by the independent Poisson [PITH_FULL_IMAGE:figures/full_fig_p012_4.png]
Figure 4
Figure 4. Figure 4: Posterior predictive densities for the correlation between home and away goals (left) [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

28 extracted references · 28 canonical work pages

  1. [1]

    Berkhout, P., and Plug, E. (2004). A bivariate Poisson count data model using conditional probabilities . Statistica Neerlandica 58, 349--364

  2. [2]

    J., Schreyer, D., and Singleton, C

    Bryson, A., Dolton, P., Reade, J. J., Schreyer, D., and Singleton, C. (2021). Causal effects of an absent crowd on performances and refereeing decisions during COVID-19 . Economics Letters 198, 109664

  3. [3]

    P., and Louis, T

    Carlin, B. P., and Louis, T. A. (2000). Empirical Bayes: Past, present and future . Journal of the American Statistical Association 95, 1286--1289

  4. [4]

    D., Lee, D., Goodrich, B., Betancourt, M., Brubaker, M., Guo, J., Li, P., and Riddell, A

    Carpenter, B., Gelman, A., Hoffman, M. D., Lee, D., Goodrich, B., Betancourt, M., Brubaker, M., Guo, J., Li, P., and Riddell, A. (2017). Stan: A probabilistic programming language . Journal of Statistical Software 76, 1--32

  5. [5]

    R., and Norman, J

    Clarke, S. R., and Norman, J. M. (1995). Home ground advantage of individual clubs in English soccer . Journal of the Royal Statistical Society: Series D (The Statistician) 44, 509--521

  6. [6]

    J., and Coles, S

    Dixon, M. J., and Coles, S. G. (1997). Modelling association football scores and inefficiencies in the football betting market . Journal of the Royal Statistical Society: Series C (Applied Statistics) 46, 265--280

  7. [7]

    Egidi, L., Pauli, F., and Torelli, N. (2018). Combining historical data and bookmakers' odds in modelling football scores . Statistical Modelling 18, 436--459

  8. [8]

    Fischer, K., and Haucap, J. (2021). Does crowd support drive the home advantage in professional football? Evidence from German ghost games during the COVID-19 pandemic . Journal of Sports Economics 22, 982--1008

Show all 28 references
  1. [9]

    Florez, M., Guindani, M., and Vannucci, M. (2025). Bayesian bivariate Conway--Maxwell--Poisson regression model for correlated count data in sports . Journal of Quantitative Analysis in Sports 21, 51--71

  2. [10]

    Gelman, A., Meng, X.-L., and Stern, H. S. (1996). Posterior predictive assessment of model fitness via realized discrepancies . Statistica Sinica 6, 733--760

  3. [11]

    B., Stern, H

    Gelman, A., Carlin, J. B., Stern, H. S., Dunson, D. B., Vehtari, A., and Rubin, D. B. (2013). Bayesian Data Analysis (3rd ed.). Chapman & Hall/CRC

  4. [12]

    D., and Gelman, A

    Hoffman, M. D., and Gelman, A. (2014). The no-u-turn sampler: Adaptively setting path lengths in Hamiltonian Monte Carlo . Journal of Machine Learning Research 15, 1593--1623

  5. [13]

    Holgate, P. (1964). Estimation for the bivariate Poisson distribution . Biometrika 51, 241--245

  6. [14]

    Karlis, D., and Ntzoufras, I. (2003). Analysis of sports data by using bivariate Poisson models . Journal of the Royal Statistical Society: Series D (The Statistician) 52, 381--393

  7. [15]

    Kawamura, K. (1984). Direct calculation of maximum likelihood estimator for the bivariate Poisson distribution . Kodai Mathematical Journal 7, 211--221

  8. [16]

    Koopman, S.J., and Lit, R. (2015). A dynamic bivariate Poisson model for analysing and forecasting match results in the English Premier League . Journal of the Royal Statistical Society: Series A (Statistics in Society) 178, 167--186

  9. [17]

    Maher, M. J. (1982). Modelling association football scores . Statistica Neerlandica 36, 109--118

  10. [18]

    McCarrick, D., Bilali\' c , M., Neave, N., and Wolfson, S. (2021). Home advantage during the COVID-19 pandemic: Analyses of European football leagues . Psychology of Sport and Exercise 56, 102013

  11. [19]

    Michels, R., Ötting, M., and Karlis, D. (2025). Extending the Dixon and Coles model: an application to women’s football data . Journal of the Royal Statistical Society: Series C (Applied Statistics) 74, 167--186

  12. [20]

    M., Balmer, N

    Nevill, A. M., Balmer, N. J., and Williams, A. M. (2002). The influence of crowd noise and experience upon refereeing decisions in football . Psychology of Sport and Exercise 3, 261--272

  13. [21]

    Petretta, M., Schiavon, L., and Diquigiovanni, J. (2025). Modelling dependence in football match outcomes: Traditional assumptions and an alternative proposal . Statistical Modelling 25, 255--269

  14. [22]

    Philipson, P. (2026). Yellow fever: an investigation into referee consistency in the ‘Big 5’ leagues of European football using a bivariate mean-parameterized Conway–Maxwell–Poisson copula model . Journal of the Royal Statistical Society Series A: Statistics in Society , 1--20...

  15. [23]

    Piancastelli, L., Friel, N., Barreto-Souza, W., and Ombao, H. (2023). Multivariate Conway-Maxwell-Poisson distribution: Sarmanov method and doubly-intractable Bayesian inference . Journal of Computational and Graphical Statistics 32, 483--500

  16. [24]

    Piancastelli, L., Barreto-Souza, W., and Ombao, H. (2023). Flexible bivariate INGARCH process with a broad range of contemporaneous correlation . Journal of Time Series Analysis 44, 206--222

  17. [25]

    Pollard, R. (2006). Worldwide regional variations in home advantage in association football . Journal of Sports Sciences 24, 231--240

  18. [26]

    Scoppa, V. (2021). Social pressure in the stadiums: Do agents change behavior without crowd support? Journal of Economic Psychology 82, 102344

  19. [27]

    Vehtari, A., Gelman, A., and Gabry, J. (2017). Practical bayesian model evaluation using leave-one-out cross-validation and WAIC . Statistics and Computing 27, 1413--1432

  20. [28]

    Zhang, N., Rastelli, R., and Friel, N. (2026). Bayesian Conway-Maxwell-Poisson model with spike-and-slab priors for dispersed count data with application to football scores . arXiv preprint arXiv:2607.18009v1

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.