REVIEW 3 major objections 4 minor 25 references
Formula One racing has become more predictable than ever, according to a ranking-based measure of all 75 seasons.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-07-31 23:48 UTC pith:CPN4AMEH
load-bearing objection Careful application of weighted ranking distances to F1's longest panel; the empirical headline is plausible but entangled with retirement churn, and the three-break claim is oversold. the 3 major comments →
Ranking-based competitive balance measures in Formula One
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that the weighted distance between two rankings — a generalization of the Kemeny distance in which overtakes at the front of the grid count for more than overtakes at the back — provides a valid and robust measure of competitive balance in a racing championship. Using every Formula One season from 1950 to 2024 and comparing start-to-finish orders within races, the authors show that the evolution of competitive balance is nearly the same under three different weight vectors, that competition is more intense when the top positions are weighted more heavily, and that competitive balance has been more unfavourable in the last two decades than ever before: the average
What carries the argument
The central object is the weighted distance between rankings, a generalization of the Kemeny distance (the minimum number of adjacent swaps needed to turn one ranking into another). Instead of counting every swap equally, it assigns a weight to a swap between positions k and k+1, and the paper uses three weightings: uniform (which reduces to the Kemeny distance), top-heavy (w_k = 1/k), and inverse-square-root (w_k = 1/sqrt(k)). The distance is normalized by dividing by its theoretical maximum — the distance between two opposite rankings — so each season's value lies between 0 and 1, and then averaged over all races or all pairs of races in a season. This lets the analysis separate how much s
Load-bearing premise
The load-bearing premise is that the normalized ranking distances are comparable across seasons even though grid size varies from about 20 drivers in recent years to more than 100 in the early 1950s; if dividing by the theoretical maximum does not fully neutralize the effect of grid size, the record-low competitive balance of the last two decades could be an artifact of smaller grids rather than a genuine loss of overtaking. A secondary premise is that drivers who fail to qua
What would settle it
Take the seasons with the lowest measured competitiveness (roughly 2005–2024) and several seasons from the 1950s, truncate each season's grid to the same top 20 or top 25 qualifiers, recompute the normalized weighted start–finish distance, and check whether the modern era still shows a record-low value. If the modern values rise to or above the old ones, the paper's headline claim is false. A second, independent check would use telemetry or video data to count actual overtakes per race across a sample of seasons and compare that count to the ranking-distance series; if overtake counts move cou
If this is right
- If the last two decades really have the lowest start–finish shuffling of any era since 1950, then the succession of rule changes aimed at improving overtaking — DRS, hybrid engines, ground-effect cars — have not restored the unpredictability that fans saw in earlier decades.
- Because all three weighting schemes produce the same long-run trend and the same three break dates, the conclusion that regulatory and technical changes have shaped competitive balance is not an artifact of choosing one weighting over another.
- Finish positions are more volatile than starting positions, but the two series have converged since the mid-1990s, implying that qualifying has become as unpredictable as race day — a change in the type of uncertainty the sport offers.
- The structural breaks around 1965, 1981, and 1992 suggest that major technical and institutional interventions have historically shifted the level of competitiveness, but no single recent change explains the current low; the authors attribute it to the continuous decline in mechanical failures.
- The framework is not limited to Formula One and can be applied to any racing competition that produces full rankings, offering a standardized way to compare outcome uncertainty across sports.
Where Pith is reading between the lines
- An implied but untested consequence of this normalization is that seasons with very different grid sizes (about 20 drivers today versus more than 100 in the 1950s) are assumed to be directly comparable. A natural extension would recompute every season's distance after capping the grid at the fastest 20 qualifiers; if the 2020s still rank as the least competitive, the main conclusion is hardened, a
- The paper works with driver rankings and excludes team-level analysis. Since fan interest in Formula One often tracks team battles and the constructors' championship, a team-level weighted distance could behave differently and might connect more directly to the audience-demand literature the paper cites.
- The measure is computed retrospectively, but it can be used prospectively as a natural experiment: any future mid-season rule change could be tested by checking whether the normalized start–finish distance moves in the direction that the paper's historical break dates would predict.
- The weighting schemes treat a swap between positions k and k+1, but in the winners' decomposition the effective weight of a swap also depends on the reference ranking. An alternative that assigns a purely position-dependent weight directly to each pair could yield different absolute levels and possibly different break dates.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes weighted ranking distances (Can 2014) as measures of competitive balance in Formula One, applying them to start–finish, start–start, and finish–finish rankings for all seasons 1950–2024. Three weight schemes are compared: the uniform Kemeny distance, w_k=1/k, and w_k=1/sqrt(k). The headline findings are that the long-run evolution is robust to the weighting, competition is more intense at the top of the field, competitive balance has been more unfavourable in the last two decades than ever before, and Bai–Perron tests detect three structural breaks that are linked to regulatory changes.
Significance. If the headline result is valid, the paper makes a substantial contribution to the sports-economics literature on competitive balance in racing. The use of weighted distances is methodologically attractive, the dataset is the longest considered so far, and the explicit comparison of three weighting schemes is a genuine robustness exercise. The Bai–Perron structural-break analysis, including the sensitivity table, is a useful addition relative to the previous ranking-based work of Peeters and Wesselbaum (2023). However, the central claim that the last two decades are uniquely uncompetitive depends on the comparability of normalized distances across seasons with very different numbers of drivers. The paper does not establish this comparability, and the sensitivity analysis in Table 1 weakens the three-break claim. These issues make the current evidence conditional.
major comments (3)
- [Section 3.2 (normalized weighted distance) and Table A.1] The normalization by the theoretical maximum does not make the measure comparable across field sizes. For a fixed overtaking event (e.g., the race leader finishing last), the normalized Kemeny distance is 2/n, and for the w_k=1/k weighting it is H_{n-1}/(n-1); both decline as n grows. Early seasons had 62–108 drivers, while modern seasons have 20–24 (Table A.1). This mechanically deflates early-season values, so the conclusion that competitive balance is more unfavourable in the last two decades than ever before may be an artifact of grid-size changes. The paper neither tests invariance to n nor controls for field size. A concrete robustness check is needed: simulate races with a fixed overtaking propensity at different n to show the normalized measure is stable, or re-estimate the series and Bai–Perron breaks on a restricted sample of races with comparable field sizes (e.g., n between 2
- [Table 1 and Section 4 (three structural breaks)] The text states that 'the existence of three structural breaks is difficult to deny.' Table 1 shows that with a 20% trimming parameter, the w_k=1/k weighting yields only one break (1967), and the other two weightings yield two breaks. Only at the default 15% trimming do all weightings produce three breaks. This direct contradiction should be acknowledged and the claim qualified. The sensitivity of the number of breaks to the trimming parameter is an important caveat for the paper's structural-break narrative.
- [Section 3.1 (non-qualifiers and pit-lane starters)] Assigning drivers who did not qualify or started from the pit lane to the last starting position, with ties broken by finish order, artificially increases the concordance between start and finish rankings. Such cases were more prevalent in the high-field seasons of earlier decades, which compounds the field-size comparability problem. The paper should provide a robustness check that excludes these drivers or handles them as missing, rather than coding them as last-place starters, to show that the headline result is not driven by this data-cleaning rule.
minor comments (4)
- [Throughout] Spelling of the distance is inconsistent: 'Kemény' appears in the text while 'Kemeny' appears in the reference list and in some locations. Use one English rendering consistently.
- [References] The Groot (2008) reference has a typo: 'Kigndom' should be 'Kingdom.'
- [Section 4] The sentence 'competition turns out to be more balanced if one focuses on the top positions' is ambiguous given that a higher indicator value is earlier defined as more intense competition. Rephrase to avoid apparent contradiction with the abstract's claim that competition is more intense at the top.
- [Figures 6 and 7] The paper does not discuss whether the start–start and finish–finish rankings, restricted to drivers who competed in both races, are subject to the same field-size bias; a brief comment on this would be useful.
Circularity Check
No circularity: exogenous weights, data-driven breaks, independent external benchmarks.
full rationale
The derivation chain is self-contained. The competitive balance indicators are computed from external race data (Kaggle/ERGAST) using a distance defined by Can (2014); the three weight vectors are fixed a priori (Kemeny; Csató 2017; Ausloos 2024), not fitted to the F1 series, and the headline conclusion is robust across all three, so none of the weights is a fitted input renamed as a prediction. The Bai-Perron breaks are estimated endogenously from the annual CB series rather than imposed, so the break dates are outputs, not inputs. The only author self-citations (Csató 2017 for the w_k=1/k scheme; Csató-Petróczy 2025 as an analogy in football) are not load-bearing: changing or removing them would not alter the central result because the analysis also relies on the independent Kemeny and Ausloos weights and on the external benchmark Peeters-Wesselbaum. The practical concern that the normalization by the theoretical maximum may not make the measure invariant to changing grid size is a measurement/identification question about cross-era comparability, not circularity: it does not involve defining the conclusion in terms of the data, fitting parameters to the target, or importing a uniqueness theorem from the authors. Thus no step reduces to its input by construction; score 2 only reflects the presence of minor, non-load-bearing self-citations.
Axiom & Free-Parameter Ledger
free parameters (2)
- Weight vector w_k =
w_k ∈ {1, 1/k, 1/√k}
- Bai–Perron trimming parameter h =
0.15
axioms (4)
- domain assumption Position-specific weights are monotonically decreasing from the top to the bottom of the ranking.
- domain assumption Normalizing by the theoretical maximum of the distance makes indicators comparable across seasons with different numbers of drivers.
- domain assumption Drivers who do not qualify or start from the pit lane can be assigned the last starting positions without materially affecting the distance measures.
- domain assumption Within each regime, the annual competitive balance follows a linear trend with stationary noise (Bai–Perron model).
read the original abstract
Competitiveness in racing sports can be measured by comparing the start and finish rankings within races, as well as the start and finish rankings across races in a season. Since the importance of position changes is non-uniform and variance at the top of the ranking is more interesting than at the bottom of the ranking, we propose using weighted distances for this purpose. Therefore, two weighting schemes are applied and compared to the standard Kemeny distance to analyse the Formula One between 1950 and 2024. The evolution of competitive balance is unexpectedly robust to the weights, but competition is more intense if one focuses on the top positions. Competitive balance has been more unfavourable in the last two decades than ever before. Statistical tests uncover three structural breaks in competitive balance that are closely related to regulatory changes, highlighting the role of decision-makers in the evolution of competitiveness.
Figures
Reference graph
Works this paper leans on
-
[1]
Ausloos, M. (2024). Hierarchy selection: New team ranking indicators for cyclist multi-stage races.European Journal of Operational Research, 314(2):807–816
2024
-
[2]
Baecker, N., Ansari, P., and Schreyer, D. (2024). Formula 1 Grands Prix demand across different distribution channels.Managing Sport and Leisure, 29(6):869–882
2024
-
[3]
and Perron, P
Bai, J. and Perron, P. (1998). Estimating and testing linear models with multiple structural changes.Econometrica, 66(1):47–78
1998
-
[4]
and Perron, P
Bai, J. and Perron, P. (2003). Computation and analysis of multiple structural change models.Journal of Applied Econometrics, 18(1):1–22
2003
-
[5]
and Feddersen, A
Budzinski, O. and Feddersen, A. (2020). Measuring competitive balance in Formula One racing. In Rodríguez, P., Kesenne, S., and Humphreys, B. R., editors,Outcome Uncertainty in Sporting Events. Edward Elgar Publishing, Cheltenham, United Kingdom
2020
-
[6]
Can, B. (2014). Weighted distances between preferences.Journal of Mathematical Economics, 51:109–115. Csató, L. (2017). On the ranking of a Swiss system chess team tournament.Annals of Operations Research, 254(1-2):17–36. Csató, L. (2023). A comparative study of scoring systems by simulations.Journal of Sports Economics, 24(4):526–545. Csató, L. and Petró...
2014
-
[7]
Fahy, R., Butler, D., and Butler, R. (2026). Broadcasting demand for Formula One: Viewer preferences for outcome uncertainty in the United States.Managing Sport and Leisure, 31(1):181–192. Garcia-del Barrio, P. and Reade, J. J. (2022). Does certainty on the winner diminish the interest in sport competitions? The case of Formula One.Empirical Economics, 63...
2026
-
[8]
Gasparetto, T., Orlova, M., and Vernikovskiy, A. (2024). Same, same but different: analyzing uncertainty of outcome in Formula One races.Managing Sport and Leisure, 29(4):651–665
2024
-
[9]
(2008).Economics, Uncertainty and European Football: Trends in Competitive Balance
Groot, L. (2008).Economics, Uncertainty and European Football: Trends in Competitive Balance. New Horizons in the Economics of Sport. Edward Elgar Publishing, Cheltenham, United Kigndom. 14
2008
-
[10]
Haigh, J. (2009). Uses and limitations of mathematics in sport.IMA Journal of Manage- ment Mathematics, 20(2):97–108
2009
-
[11]
Judde, C., Booth, R., and Brooks, R. (2013). Second place is first of the losers: An analysis of competitive balance in Formula One.Journal of Sports Economics, 14(4):411–439
2013
-
[12]
Kemeny, J. G. (1959). Mathematics without numbers.Daedalus, 88(4):577–591
1959
-
[13]
Kemeny, J. G. and Snell, L. J. (1962). Preference ranking: an axiomatic approach. In Mathematical Models in the Social Sciences, pages 9–23. Ginn, New York
1962
-
[14]
Krauskopf, T., Langen, M., and Bünger, B. (2010). The search for optimal competitive balance in Formula One. CAWM Discussion Paper 38, Westfälische Wilhelms-Universität Münster, Centrum für Angewandte Wirtschaftsforschung (CAWM)
2010
-
[15]
Lee, J. M., Prince, C. S., and Zeager, L. A. (2026). Lap-time dispersion as an aspect of within-race competitive balance in Formula one.Applied Economics, in press. DOI: 10.1080/00036846.2026.2683698
arXiv 2026
-
[16]
Lee, Y. H. and Fort, R. (2005). Structural change in MLB competitive balance: The depression, team location, and integration.Economic Inquiry, 43(1):158–169
2005
-
[17]
Lee, Y. H. and Fort, R. (2012). Competitive balance: Time series lessons from the English Premier League.Scottish Journal of Political Economy, 59(3):266–282
2012
-
[18]
and Ntzoufras, I
Manasis, V. and Ntzoufras, I. (2014). Between-seasons competitive balance in European football: review of existing and development of specially designed indices.Journal of Quantitative Analysis in Sports, 10(2):139–152
2014
-
[19]
and Runkel, M
Mastromarco, C. and Runkel, M. (2009). Rule changes and competitive balance in Formula One motor racing.Applied Economics, 41(23):3003–3014
2009
-
[20]
Mills, B. M. and Salaga, S. (2015). Historical time series perspectives on competitive balance in NCAA Division I basketball.Journal of Sports Economics, 16(6):614–646
2015
-
[21]
Pedroche, F. (2024). Competitiveness of Formula 1 championship from 2012 to 2022 as measured by Kendall corrected evolutive coefficient. Manuscript. DOI: 10.48550/arXiv.2501.00126
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2501.00126 2024
-
[22]
and Wesselbaum, D
Peeters, R. and Wesselbaum, D. (2023). Competitiveness in Formula One.Sports Economics Review, 2:100007
2023
-
[23]
and Torgler, B
Schreyer, D. and Torgler, B. (2018). On the role of race outcome uncertainty in the TV demand for Formula 1 Grands Prix.Journal of Sports Economics, 19(2):211–229
2018
-
[24]
Szymanski, S. (2003). The economic design of sporting contests.Journal of Economic Literature, 41(4):1137–1187
2003
-
[25]
Zeileis, A., Leisch, F., Hornik, K., and Kleiber, C. (2002). strucchange: An R package for testing for structural change in linear regression models.Journal of Statistical Software, 7(3-4):1–38. 15 Appendix Table A.1: The number of drivers and races in Formula One seasons Season Drivers Races Season Drivers Races 1950 81 7 1988 36 16 1951 84 8 1989 47 16 ...
2002
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.