{"id":"ee1e4591-6414-45b2-befd-04a05b2cb1dd","arxiv_id":"1908.08991","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Match outcomes in 11 major European football leagues became more predictable between 1993 and 2019, while team inequality rose and home advantage declined.","lead":"This paper measures whether top European football leagues have become more predictable over 26 years, using a simple past-results network model benchmarked against betting odds. It finds that predictability generally increased, team inequality grew, and home-field advantage declined across the leagues studied.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Excluding draws may drive the reported AUC trend; the paper never tests whether changing draw rates distort the predictability measure.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the paper's explicit exclusion of drawn matches. This is the most fundamental threat to the central claim because it concerns the validity of the dependent variable itself. If the draw rate has changed over time, the subset of non-draw matches is not a random sample of all matches, and the AUC trend on that subset may not reflect the predictability of football as a whole. The paper provides no analysis of draw frequencies and no justification for why the binary decisive-match problem is the right operationalization of 'predictability.' This is more central than the lack of confidence intervals (which would only affect the precision of an estimate that may be measuring the wrong construct) or the benchmark inconsistency (which concerns a secondary validation claim). The proposed test—recomputing the trend with a three-outcome model and/or controlling for draw share—directly addresses whether the conclusion survives a more faithful operationalization. I agree with the reader's conditional verdict: the paper presents an interesting and plausible empirical observation, but the central claim is not fully supported until the draw-exclusion issue is resolved. Independent strengths include the large dataset, the self-contained network model, and the home-advantage analysis, but these do not compensate for the potential selection artifact. The verdict should remain CONDITIONAL, pending the additional analysis.","tokens_in":14292,"tokens_out":6463,"duration_ms":64176,"concrete_test":"For each league-season, recompute predictability including draws: fit a multinomial logistic regression on the same dyadic/network score difference with three outcomes (home win, draw, away win) and compute the multiclass AUC (or three-class Brier score) on the same matches. Compare the time trend of this three-outcome measure with the reported two-class AUC trend. If the three-outcome trend is flat or downward while the two-class trend is upward, the central claim is an artifact of excluding draws. As a sensitivity check, regress the reported two-class AUC on season with the league-season draw share as a covariate; if the season coefficient loses significance after controlling for draw share, the same conclusion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that 'football is becoming more predictable' rests entirely on the AUC of a binary home-win/away-win classifier trained and evaluated on a subset of matches: 'we limit our analysis to the matches that have a winner and we eliminate the ties from the entirety of this study' (Modelling predictability section). This is not a harmless preprocessing step. Football draws are frequent (roughly a quarter of matches) and are known to occur disproportionately between closely matched teams. If the share of draws has declined over 26 years—as would be expected if team strengths have become more unequal—then the non-draw subset becomes progressively more lopsided: the matches retained in the analysis are increasingly those with a clear winner. A classifier's AUC on such a subset can rise even if the underlying predictability of a full match (home win/draw/away win) is unchanged, because the removal of boundary cases mechanically increases the separation between the two retained classes. The paper reports an upward AUC trend in most leagues and convergence near 0.75 AUC, but it never reports the draw rate over time, never models the three-outcome match, and never controls for draw share in the trend. The reported trend may therefore be an artifact of selection rather than evidence that matches have become inherently more predictable. This is the weakest link in the causal chain from data to conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper uses 87,816 matches from eleven top European leagues between the 1993/94 and 2018/19 seasons to ask whether football outcomes have become more predictable. The authors build two self-contained models—a simple points-share difference and a network eigenvector-centrality score computed from past matches—and benchmark them against Bet365 odds. After excluding drawn matches, they fit a logistic regression of home-win/away-win on the score difference and measure predictive performance with the Brier score and AUC. They report a general upward trend in per-season AUC (with Greece and Turkey as exceptions), convergence toward AUC around 0.75, a positive correlation between AUC and the Gini coefficient of season points, and a decreasing home-field advantage across all leagues, which they interpret as consistent with a 'gentrification' of football.","tokens_in":14513,"tokens_out":8632,"duration_ms":83589,"significance":"The study is a valuable first large-scale, historically consistent attempt to quantify predictability in football with a simple network-based model, and the public data source and fixed model specification make the approach transparent and reproducible in principle. The decline in home-field advantage is a robust and interesting side result. However, the central trend claim currently rests on an unconditioned two-class analysis of non-draw matches and on visually assessed time trends without inferential support; these issues must be resolved before the main conclusion can be accepted.","major_comments":[{"comment":"The decisive modelling choice is to 'limit our analysis to the matches that have a winner and eliminate the ties from the entirety of this study.' Because draws are a large, non-random fraction of football matches, removing them can change the apparent difficulty of the classification problem even if the full three-outcome distribution is unchanged: if the draw rate fell over time, the retained subset would contain progressively more lopsided matches, and AUC on that subset could rise mechanically. The paper never reports draw rates by league and season, never estimates a three-outcome (home/draw/away) model, and never checks whether the reported AUC trend survives adjustment for draw frequency. The authors should add these analyses or explicitly demonstrate that the non-draw AUC trend is not an artefact of selection.","section":"Modelling predictability"},{"comment":"The central 'becoming more predictable' claim is supported only by lowess curves and visual inspection in Figure 4. AUC for a single league-season is estimated from a few hundred matches and will carry substantial sampling noise, yet no confidence intervals, trend coefficients, significance tests, or adjustments for temporal autocorrelation are reported. The authors should supply formal trend estimates for each league (for example, regression or rank-based trend tests with season as the covariate) and, given the autocorrelated series, a time-series-aware inference procedure. The statement that 'all leagues tend to converge towards 0.75 AUC' likewise needs quantitative support.","section":"Results and Discussion: Predictability over Time"},{"comment":"The text says the models and the betting-market benchmark are 'statistically indistinguishable at the 2% significance level for the majority of year-leagues,' but Tables S1 and S2 contradict this. In Table S2, eight of eleven AUC comparisons between the Network model and the Market are significant at p<0.02 (England, Germany, Spain, Italy, Portugal, Netherlands, Belgium, France), all favouring the market, and several Brier-score comparisons are also significant at the 2% level. This is an internal inconsistency in a passage used to justify the model as a reliable measurement tool. The comparison should be re-run or, at minimum, described accurately in the text.","section":"Model Performance"}],"minor_comments":[{"comment":"Logistic-regression parameters are not obtained by 'ordinary least squares methods'; the standard estimation is maximum likelihood via iteratively reweighted least squares. Please correct the wording.","section":"Modelling predictability, Eq. (1)"},{"comment":"The caption says the network is shown 'after 240 matches have been played for n = 0.5' but also says the centrality is calculated from 'the last 190 matches'; for a 380-match season, n = 0.5 gives 190 matches, so one of these numbers is wrong.","section":"Fig. 1 caption"},{"comment":"Please state the smoothing span used for the lowess fits and clarify whether the plotted points are raw per-season values or smoothed values; the dual-axis display currently makes it hard to assess the data behind the curves.","section":"Fig. 4"},{"comment":"The AUC–Gini correlations need a stated method (Pearson or Spearman; raw or smoothed series) and confidence intervals or p-values; the value 0.413 for Belgium should be described as weak-to-moderate, not 'high.'","section":"Table 1"},{"comment":"Because draws are excluded, the market probabilities in Eq. (2) are conditional on no draw; please state this explicitly, since the raw Bet365 odds include a third outcome and the normalization is not the usual three-outcome calibration.","section":"Eq. (2)"},{"comment":"The limitations paragraph acknowledges data volume and model sophistication but omits the draw-exclusion choice, which is the most consequential modelling decision for the main claim; this should be discussed and, ideally, checked.","section":"Conclusion"},{"comment":"There are numerous typos and OCR-like artifacts (for example, 'predictibility' in the Results section, '88 thousands' in the title/abstract, 'Ho e Points' in Figure S8, 'T urkey' in Table 1) that a careful proofreading pass should remove.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The main risk is the draw-exclusion issue; if the robustness checks requested in Major Comment 1 show that the trend persists in a three-outcome setting and the trend tests are added, I would support publication. I do not see evidence of circularity or inappropriate handling of the data."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper gives a transparent, large-scale measurement of predictability in 11 European leagues over 26 years, using a self-contained network model based only on past results and benchmarking it against betting odds. That longitudinal cross-league comparison is genuinely new; the literature mostly studies single seasons or single leagues, and the authors made their data source and model explicit enough to reproduce. Credit where due: the paper does a lot of things right—fixed training window, external benchmark, robustness checks across n values, and a clear presentation of trends per league.\n\nThe soft spots are not hard to find, and one is load-bearing. The paper excludes draws from the entire analysis ('we eliminate the ties from the entirety of this study'). Draws are roughly a quarter of matches, and they occur disproportionately between evenly matched teams. If the share of draws declined over 26 years—which you would expect if stronger teams pulled away—then the remaining home-win/away-win subset becomes more lopsided, and a binary classifier's AUC can rise even if the full three-outcome distribution is unchanged. The authors never report draw rates over time, never model the three-outcome match, and never test whether the AUC trend survives controlling for draw share. This is not a minor preprocessing choice; it is the main evidence behind 'football is becoming more predictable.'\n\nThe second issue is statistical: the time trends are shown with smooth fits but no confidence intervals or significance tests. Some of the trends look clear, but Greece and Turkey go the other way, and the 'convergence near 0.75' claim needs more than a visual. The correlation between AUC and Gini is computed across time without removing shared trends, so the high r values are partly a time-series artifact. Finally, the text says the models are 'statistically indistinguishable' from the betting-market benchmark at the 2% level, but the supplementary tables show t-tests that are significant at 5% for most leagues and at 2% for several. That sentence should be corrected.\n\nNone of this kills the underlying question—it is plausible that football has become more predictable—but the paper as written does not establish it. The fix is feasible: report draw shares, fit an ordinal or multinomial model, add confidence bands, detrend the Gini correlations, and correct the benchmark claims.\n\nMy recommendation: send it to peer review, but with the expectation of major revision. A serious referee can help the authors turn a suggestive dataset into a defensible measurement.","headline":"A useful panel measurement of football predictability, but the headline trend may be an artifact of the paper's decision to drop draws; worth refereeing, not taking at face value.","tokens_in":15021,"tokens_out":2742,"would_cite":false,"duration_ms":29368,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Top-division European football has become more predictable over the 26 seasons from 1993-94 to 2018-19.","keywords":["football","predictability","eigenvector centrality","network analysis","competitive balance","home-field advantage","Gini coefficient","sports forecasting"],"falsifier":"Re-run the same centrality-based model on all 87,816 matches using three outcomes—home win, draw, and away win—and compare the AUC trend with the paper's two-outcome trend; if the rise in predictability disappears once draws are included, the central claim is specific to decisive matches, not to football generally.","tokens_in":14089,"feed_emoji":"⚽","tokens_out":12305,"duration_ms":120085,"temperature":0.7,"pith_summary":"The paper tries to establish that top-division European football has become more predictable over the 26 seasons from 1993-94 to 2018-19. It measures predictability with a deliberately self-contained model that uses only past results: matches are turned into a directed network with edges from loser to winner weighted by points, team strength is read from eigenvector centrality, and the home-minus-away score difference is fed into a logistic regression that is scored by area under the ROC curve. Across 87,816 matches and 11 major leagues, the paper finds rising predictability in most leagues, with all of them tending toward an AUC near 0.75. It reports two supporting trends—greater inequality in final points and a shrinking home-field advantage—and argues they are consistent with money concentrating success in the same clubs, while stating explicitly that the causal link to monetization is not directly tested.","feed_headline":"Top European football more predictable over 26 seasons","feed_subtitle":"A study of 87,816 matches finds rising inequality and fading home advantage as results grow easier to forecast.","key_machinery":"The central mechanism is a directed match network whose eigenvector centrality serves as a self-contained team-strength rating. In the network, each edge points from the loser to the winner and is weighted by the points the winner earned, and a team's eigenvector centrality reflects not only how often it wins but how strong its beaten opponents are—a team is central if it beats teams that are themselves central. The difference between the home and away teams' centrality scores is the single input to a logistic regression, and the area under that regression's ROC curve on matches outside the training window is the paper's operational measure of predictability. The model is intentionally simple and time-consistent: it uses only past results and fixed parameters, so any historical trend in its AUC reflects changes in the game itself rather than improvements in the predictor.","core_discovery":"The central claim is that football's results have become easier to anticipate, not just that some teams are stronger. The authors define predictability as the out-of-sample AUC of a fixed, past-results-only model, so an upward trend in AUC means the same simple predictor extracts more information from recent results now than it did in the 1990s. They find that seven of the eleven leagues studied show increasing AUC over the sample, Greece and Turkey show the opposite but with recent increases, Belgium and Italy are stable, and the leagues tend to converge near 0.75 AUC. The paper also shows that the Gini coefficient of end-of-season points is positively correlated with AUC in every league, and that home-field advantage, measured both by the model's logistic offset and by the historical share of home points, has declined in all eleven leagues. It explicitly does not claim to have established the monetary cause of these trends.","pith_inferences":["A direct test of the tie-removal assumption would be to include draws as a third outcome; if the AUC trend disappears, the paper's result holds only for matches with a winner, not for football as a whole.","The paper's gentrification feedback loop can be tested with club finances: within a league, seasons with larger gaps in wage bills or transfer spending between top and bottom clubs should show higher AUC than seasons with smaller gaps.","Because the methodology is sport-agnostic, applying the same network-AUC measure to leagues with salary caps would disentangle whether the predictability trend comes from unrestricted spending or from something intrinsic to football; the paper names this comparison as future work.","The model's own home-advantage parameter offers a way to separate the two trends the paper reports: a regression of league-season AUC on league-season home advantage would show whether the shrinking home boost is actually one of the mechanisms making results easier to predict."],"forward_implications":["Because the prediction model uses only past results and its parameters are fixed, the rising AUC cannot be attributed to smarter forecasting; the same simple model has become a better predictor of future results over time.","The positive correlation between AUC and Gini coefficient in all 11 leagues indicates that inequality in final points and predictability move together, so leagues with more concentrated results are also the leagues where outcomes are easier to forecast.","The universal decline in home-field advantage means that home status contributes less to match outcomes than it used to, which reduces the number of matches a weaker team can win simply by playing at home.","The observed convergence of league AUC values near 0.75 means that differences between leagues in how predictable their results are have shrunk even while the overall level of predictability has risen."],"supporting_citations":[{"why":"Defines eigenvector centrality, the recursive score used to turn the directed match network into team-strength ratings.","marker":"(27)"},{"why":"Supplies the network theory treatment of centrality that the match network is built on.","marker":"(26)"},{"why":"Motivates using eigenvector centrality to quantify success from relational data, the core modelling move of the paper.","marker":"(16)"},{"why":"Defines ROC curves and AUC, which the paper uses as its operational measure of predictability.","marker":"(37)"},{"why":"Provides the justification for treating AUC as a coherent measure of aggregated classification performance.","marker":"(42)"},{"why":"Supplies the Brier score, the complementary loss function used to validate both prediction models.","marker":"(36)"}],"fun_headline_variants":["Football predictability rises as home advantage fades","Top leagues get easier to predict, data on 88k matches","Rising inequality makes top-flight football more predictable","Predictability up in major leagues, home edge down","Football results increasingly foreseeable across 11 leagues"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that removing drawn matches does not distort the historical comparison; if draws changed in frequency or predictability over 26 years, the upward trend measured on matches with a winner may not describe football as a whole.","fun_headline_variants_meta":{"raw":{"variants":["Football predictability rises as home advantage fades","Top leagues get easier to predict, data on 88k matches","Rising inequality makes top-flight football more predictable","Predictability up in major leagues, home edge down","Football results increasingly foreseeable across 11 leagues"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000194,"raw_usage":{"total_tokens":1326,"prompt_tokens":893,"completion_tokens":433,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":509,"completion_tokens_details":{"reasoning_tokens":359}},"tokens_in":509,"tokens_out":433,"duration_ms":4543,"temperature":1.0,"reasoning_tokens":359,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:24:33.495396+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same centrality-based model on all 87,816 matches using three outcomes—home win, draw, and away win—and compare the AUC trend with the paper's two-outcome trend; if the rise in predictability disappears once draws are included, the central claim is specific to decisive matches, not to football generally.","supporting_citations":[],"review_version":1}