Pith. sign in

REVIEW 3 major objections 4 minor 25 references

Formula One racing has become more predictable than ever, according to a ranking-based measure of all 75 seasons.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-07-31 23:48 UTC pith:CPN4AMEH

load-bearing objection Careful application of weighted ranking distances to F1's longest panel; the empirical headline is plausible but entangled with retirement churn, and the three-break claim is oversold. the 3 major comments →

arxiv 2607.23303 v1 pith:CPN4AMEH submitted 2026-07-25 econ.GN physics.soc-phq-fin.ECstat.AP

Ranking-based competitive balance measures in Formula One

classification econ.GN physics.soc-phq-fin.ECstat.AP MSC 62P2091-1091B14
keywords competitive balanceFormula Onerankingweighted distanceKemeny distanceoutcome uncertaintystructural breakssports economics
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper proposes measuring competitive balance in Formula One by comparing the starting and finishing orders of every race, and by comparing rankings across races within a season, using a weighted distance that values position swaps at the top of the grid more than swaps at the back. Applying this to all seasons from 1950 to 2024, the authors find that the long-run trend in competitiveness is nearly identical under three different weighting schemes, but that the normalized distance between start and finish rankings has fallen to its lowest level in the last two decades — meaning racing has become more predictable than ever. They also detect three structural breaks in the series, around 1965, 1981, and 1992, which line up with major technical and regulatory changes. The payoff is a simple, parameter-free tool for judging whether rule changes actually affect how uneven the racing is, and one that transfers to any other ranking-based sport.

Core claim

The paper's central claim is that the weighted distance between two rankings — a generalization of the Kemeny distance in which overtakes at the front of the grid count for more than overtakes at the back — provides a valid and robust measure of competitive balance in a racing championship. Using every Formula One season from 1950 to 2024 and comparing start-to-finish orders within races, the authors show that the evolution of competitive balance is nearly the same under three different weight vectors, that competition is more intense when the top positions are weighted more heavily, and that competitive balance has been more unfavourable in the last two decades than ever before: the average

What carries the argument

The central object is the weighted distance between rankings, a generalization of the Kemeny distance (the minimum number of adjacent swaps needed to turn one ranking into another). Instead of counting every swap equally, it assigns a weight to a swap between positions k and k+1, and the paper uses three weightings: uniform (which reduces to the Kemeny distance), top-heavy (w_k = 1/k), and inverse-square-root (w_k = 1/sqrt(k)). The distance is normalized by dividing by its theoretical maximum — the distance between two opposite rankings — so each season's value lies between 0 and 1, and then averaged over all races or all pairs of races in a season. This lets the analysis separate how much s

Load-bearing premise

The load-bearing premise is that the normalized ranking distances are comparable across seasons even though grid size varies from about 20 drivers in recent years to more than 100 in the early 1950s; if dividing by the theoretical maximum does not fully neutralize the effect of grid size, the record-low competitive balance of the last two decades could be an artifact of smaller grids rather than a genuine loss of overtaking. A secondary premise is that drivers who fail to qua

What would settle it

Take the seasons with the lowest measured competitiveness (roughly 2005–2024) and several seasons from the 1950s, truncate each season's grid to the same top 20 or top 25 qualifiers, recompute the normalized weighted start–finish distance, and check whether the modern era still shows a record-low value. If the modern values rise to or above the old ones, the paper's headline claim is false. A second, independent check would use telemetry or video data to count actual overtakes per race across a sample of seasons and compare that count to the ranking-distance series; if overtake counts move cou

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the last two decades really have the lowest start–finish shuffling of any era since 1950, then the succession of rule changes aimed at improving overtaking — DRS, hybrid engines, ground-effect cars — have not restored the unpredictability that fans saw in earlier decades.
  • Because all three weighting schemes produce the same long-run trend and the same three break dates, the conclusion that regulatory and technical changes have shaped competitive balance is not an artifact of choosing one weighting over another.
  • Finish positions are more volatile than starting positions, but the two series have converged since the mid-1990s, implying that qualifying has become as unpredictable as race day — a change in the type of uncertainty the sport offers.
  • The structural breaks around 1965, 1981, and 1992 suggest that major technical and institutional interventions have historically shifted the level of competitiveness, but no single recent change explains the current low; the authors attribute it to the continuous decline in mechanical failures.
  • The framework is not limited to Formula One and can be applied to any racing competition that produces full rankings, offering a standardized way to compare outcome uncertainty across sports.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • An implied but untested consequence of this normalization is that seasons with very different grid sizes (about 20 drivers today versus more than 100 in the 1950s) are assumed to be directly comparable. A natural extension would recompute every season's distance after capping the grid at the fastest 20 qualifiers; if the 2020s still rank as the least competitive, the main conclusion is hardened, a
  • The paper works with driver rankings and excludes team-level analysis. Since fan interest in Formula One often tracks team battles and the constructors' championship, a team-level weighted distance could behave differently and might connect more directly to the audience-demand literature the paper cites.
  • The measure is computed retrospectively, but it can be used prospectively as a natural experiment: any future mid-season rule change could be tested by checking whether the normalized start–finish distance moves in the direction that the paper's historical break dates would predict.
  • The weighting schemes treat a swap between positions k and k+1, but in the winners' decomposition the effective weight of a swap also depends on the reference ranking. An alternative that assigns a purely position-dependent weight directly to each pair could yield different absolute levels and possibly different break dates.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes weighted ranking distances (Can 2014) as measures of competitive balance in Formula One, applying them to start–finish, start–start, and finish–finish rankings for all seasons 1950–2024. Three weight schemes are compared: the uniform Kemeny distance, w_k=1/k, and w_k=1/sqrt(k). The headline findings are that the long-run evolution is robust to the weighting, competition is more intense at the top of the field, competitive balance has been more unfavourable in the last two decades than ever before, and Bai–Perron tests detect three structural breaks that are linked to regulatory changes.

Significance. If the headline result is valid, the paper makes a substantial contribution to the sports-economics literature on competitive balance in racing. The use of weighted distances is methodologically attractive, the dataset is the longest considered so far, and the explicit comparison of three weighting schemes is a genuine robustness exercise. The Bai–Perron structural-break analysis, including the sensitivity table, is a useful addition relative to the previous ranking-based work of Peeters and Wesselbaum (2023). However, the central claim that the last two decades are uniquely uncompetitive depends on the comparability of normalized distances across seasons with very different numbers of drivers. The paper does not establish this comparability, and the sensitivity analysis in Table 1 weakens the three-break claim. These issues make the current evidence conditional.

major comments (3)
  1. [Section 3.2 (normalized weighted distance) and Table A.1] The normalization by the theoretical maximum does not make the measure comparable across field sizes. For a fixed overtaking event (e.g., the race leader finishing last), the normalized Kemeny distance is 2/n, and for the w_k=1/k weighting it is H_{n-1}/(n-1); both decline as n grows. Early seasons had 62–108 drivers, while modern seasons have 20–24 (Table A.1). This mechanically deflates early-season values, so the conclusion that competitive balance is more unfavourable in the last two decades than ever before may be an artifact of grid-size changes. The paper neither tests invariance to n nor controls for field size. A concrete robustness check is needed: simulate races with a fixed overtaking propensity at different n to show the normalized measure is stable, or re-estimate the series and Bai–Perron breaks on a restricted sample of races with comparable field sizes (e.g., n between 2
  2. [Table 1 and Section 4 (three structural breaks)] The text states that 'the existence of three structural breaks is difficult to deny.' Table 1 shows that with a 20% trimming parameter, the w_k=1/k weighting yields only one break (1967), and the other two weightings yield two breaks. Only at the default 15% trimming do all weightings produce three breaks. This direct contradiction should be acknowledged and the claim qualified. The sensitivity of the number of breaks to the trimming parameter is an important caveat for the paper's structural-break narrative.
  3. [Section 3.1 (non-qualifiers and pit-lane starters)] Assigning drivers who did not qualify or started from the pit lane to the last starting position, with ties broken by finish order, artificially increases the concordance between start and finish rankings. Such cases were more prevalent in the high-field seasons of earlier decades, which compounds the field-size comparability problem. The paper should provide a robustness check that excludes these drivers or handles them as missing, rather than coding them as last-place starters, to show that the headline result is not driven by this data-cleaning rule.
minor comments (4)
  1. [Throughout] Spelling of the distance is inconsistent: 'Kemény' appears in the text while 'Kemeny' appears in the reference list and in some locations. Use one English rendering consistently.
  2. [References] The Groot (2008) reference has a typo: 'Kigndom' should be 'Kingdom.'
  3. [Section 4] The sentence 'competition turns out to be more balanced if one focuses on the top positions' is ambiguous given that a higher indicator value is earlier defined as more intense competition. Rephrase to avoid apparent contradiction with the abstract's claim that competition is more intense at the top.
  4. [Figures 6 and 7] The paper does not discuss whether the start–start and finish–finish rankings, restricted to drivers who competed in both races, are subject to the same field-size bias; a brief comment on this would be useful.

Circularity Check

0 steps flagged

No circularity: exogenous weights, data-driven breaks, independent external benchmarks.

full rationale

The derivation chain is self-contained. The competitive balance indicators are computed from external race data (Kaggle/ERGAST) using a distance defined by Can (2014); the three weight vectors are fixed a priori (Kemeny; Csató 2017; Ausloos 2024), not fitted to the F1 series, and the headline conclusion is robust across all three, so none of the weights is a fitted input renamed as a prediction. The Bai-Perron breaks are estimated endogenously from the annual CB series rather than imposed, so the break dates are outputs, not inputs. The only author self-citations (Csató 2017 for the w_k=1/k scheme; Csató-Petróczy 2025 as an analogy in football) are not load-bearing: changing or removing them would not alter the central result because the analysis also relies on the independent Kemeny and Ausloos weights and on the external benchmark Peeters-Wesselbaum. The practical concern that the normalization by the theoretical maximum may not make the measure invariant to changing grid size is a measurement/identification question about cross-era comparability, not circularity: it does not involve defining the conclusion in terms of the data, fitting parameters to the target, or importing a uniqueness theorem from the authors. Thus no step reduces to its input by construction; score 2 only reflects the presence of minor, non-load-bearing self-citations.

Axiom & Free-Parameter Ledger

2 free parameters · 4 axioms · 0 invented entities

The central measure depends on a hand-chosen weight vector and a normalization convention; neither is fitted to F1 data. The structural-break analysis depends on the trimming parameter. No new entities are introduced.

free parameters (2)
  • Weight vector w_k = w_k ∈ {1, 1/k, 1/√k}
    The weighted distance d_C(w) depends on the chosen weights; they are not estimated from F1 data but taken from prior literature (Kemeny; Csató 2017; Ausloos 2024) and assumed monotonically decreasing.
  • Bai–Perron trimming parameter h = 0.15
    Default minimum segment size in strucchange; Table 1 shows the number and dates of breakpoints vary with h, e.g. three breaks at h=0.15 but only one break for w=1/k at h=0.20.
axioms (4)
  • domain assumption Position-specific weights are monotonically decreasing from the top to the bottom of the ranking.
    Stated in Sections 1 and 3.2; not empirically estimated for Formula One but consistent with point-scoring systems.
  • domain assumption Normalizing by the theoretical maximum of the distance makes indicators comparable across seasons with different numbers of drivers.
    Relied on throughout Section 4 (e.g., Figure 2, 'ever before' claim); no invariance test is provided in Section 3.2.
  • domain assumption Drivers who do not qualify or start from the pit lane can be assigned the last starting positions without materially affecting the distance measures.
    Data-cleaning rule in Section 3.1; creates artificial ties in the start ranking when multiple drivers are moved to the back.
  • domain assumption Within each regime, the annual competitive balance follows a linear trend with stationary noise (Bai–Perron model).
    Equation in Section 3.3; the number of breaks is selected by the strucchange package defaults.

pith-pipeline@v1.3.0-alltime-deepseek · 9561 in / 18827 out tokens · 163870 ms · 2026-07-31T23:48:26.598019+00:00 · methodology

0 comments
read the original abstract

Competitiveness in racing sports can be measured by comparing the start and finish rankings within races, as well as the start and finish rankings across races in a season. Since the importance of position changes is non-uniform and variance at the top of the ranking is more interesting than at the bottom of the ranking, we propose using weighted distances for this purpose. Therefore, two weighting schemes are applied and compared to the standard Kemeny distance to analyse the Formula One between 1950 and 2024. The evolution of competitive balance is unexpectedly robust to the weights, but competition is more intense if one focuses on the top positions. Competitive balance has been more unfavourable in the last two decades than ever before. Statistical tests uncover three structural breaks in competitive balance that are closely related to regulatory changes, highlighting the role of decision-makers in the evolution of competitiveness.

Figures

Figures reproduced from arXiv: 2607.23303 by D\'ora Gr\'eta Petr\'oczy, L\'aszl\'o Csat\'o.

Figure 1
Figure 1. Figure 1: Normalised Kemény distance of start–finish rankings, 1993–2019 [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Normalised Kemény distance of start–finish [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Normalised Kemény distance of start–finish rankings with structural breaks [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Normalised weighted distance of start–finish rankings with structural breaks [PITH_FULL_IMAGE:figures/full_fig_p010_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Evolution of competitive balance based on start–finish rankings [PITH_FULL_IMAGE:figures/full_fig_p011_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Normalised Kemény distance of start–start [PITH_FULL_IMAGE:figures/full_fig_p011_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Normalised weighted distance of start–start [PITH_FULL_IMAGE:figures/full_fig_p012_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

25 extracted references · 1 canonical work pages · 1 internal anchor

  1. [1]

    Ausloos, M. (2024). Hierarchy selection: New team ranking indicators for cyclist multi-stage races.European Journal of Operational Research, 314(2):807–816

  2. [2]

    Baecker, N., Ansari, P., and Schreyer, D. (2024). Formula 1 Grands Prix demand across different distribution channels.Managing Sport and Leisure, 29(6):869–882

  3. [3]

    and Perron, P

    Bai, J. and Perron, P. (1998). Estimating and testing linear models with multiple structural changes.Econometrica, 66(1):47–78

  4. [4]

    and Perron, P

    Bai, J. and Perron, P. (2003). Computation and analysis of multiple structural change models.Journal of Applied Econometrics, 18(1):1–22

  5. [5]

    and Feddersen, A

    Budzinski, O. and Feddersen, A. (2020). Measuring competitive balance in Formula One racing. In Rodríguez, P., Kesenne, S., and Humphreys, B. R., editors,Outcome Uncertainty in Sporting Events. Edward Elgar Publishing, Cheltenham, United Kingdom

  6. [6]

    Can, B. (2014). Weighted distances between preferences.Journal of Mathematical Economics, 51:109–115. Csató, L. (2017). On the ranking of a Swiss system chess team tournament.Annals of Operations Research, 254(1-2):17–36. Csató, L. (2023). A comparative study of scoring systems by simulations.Journal of Sports Economics, 24(4):526–545. Csató, L. and Petró...

  7. [7]

    Fahy, R., Butler, D., and Butler, R. (2026). Broadcasting demand for Formula One: Viewer preferences for outcome uncertainty in the United States.Managing Sport and Leisure, 31(1):181–192. Garcia-del Barrio, P. and Reade, J. J. (2022). Does certainty on the winner diminish the interest in sport competitions? The case of Formula One.Empirical Economics, 63...

  8. [8]

    Gasparetto, T., Orlova, M., and Vernikovskiy, A. (2024). Same, same but different: analyzing uncertainty of outcome in Formula One races.Managing Sport and Leisure, 29(4):651–665

  9. [9]

    (2008).Economics, Uncertainty and European Football: Trends in Competitive Balance

    Groot, L. (2008).Economics, Uncertainty and European Football: Trends in Competitive Balance. New Horizons in the Economics of Sport. Edward Elgar Publishing, Cheltenham, United Kigndom. 14

  10. [10]

    Haigh, J. (2009). Uses and limitations of mathematics in sport.IMA Journal of Manage- ment Mathematics, 20(2):97–108

  11. [11]

    Judde, C., Booth, R., and Brooks, R. (2013). Second place is first of the losers: An analysis of competitive balance in Formula One.Journal of Sports Economics, 14(4):411–439

  12. [12]

    Kemeny, J. G. (1959). Mathematics without numbers.Daedalus, 88(4):577–591

  13. [13]

    Kemeny, J. G. and Snell, L. J. (1962). Preference ranking: an axiomatic approach. In Mathematical Models in the Social Sciences, pages 9–23. Ginn, New York

  14. [14]

    Krauskopf, T., Langen, M., and Bünger, B. (2010). The search for optimal competitive balance in Formula One. CAWM Discussion Paper 38, Westfälische Wilhelms-Universität Münster, Centrum für Angewandte Wirtschaftsforschung (CAWM)

  15. [15]

    M., Prince, C

    Lee, J. M., Prince, C. S., and Zeager, L. A. (2026). Lap-time dispersion as an aspect of within-race competitive balance in Formula one.Applied Economics, in press. DOI: 10.1080/00036846.2026.2683698

  16. [16]

    Lee, Y. H. and Fort, R. (2005). Structural change in MLB competitive balance: The depression, team location, and integration.Economic Inquiry, 43(1):158–169

  17. [17]

    Lee, Y. H. and Fort, R. (2012). Competitive balance: Time series lessons from the English Premier League.Scottish Journal of Political Economy, 59(3):266–282

  18. [18]

    and Ntzoufras, I

    Manasis, V. and Ntzoufras, I. (2014). Between-seasons competitive balance in European football: review of existing and development of specially designed indices.Journal of Quantitative Analysis in Sports, 10(2):139–152

  19. [19]

    and Runkel, M

    Mastromarco, C. and Runkel, M. (2009). Rule changes and competitive balance in Formula One motor racing.Applied Economics, 41(23):3003–3014

  20. [20]

    Mills, B. M. and Salaga, S. (2015). Historical time series perspectives on competitive balance in NCAA Division I basketball.Journal of Sports Economics, 16(6):614–646

  21. [21]

    Pedroche, F. (2024). Competitiveness of Formula 1 championship from 2012 to 2022 as measured by Kendall corrected evolutive coefficient. Manuscript. DOI: 10.48550/arXiv.2501.00126

  22. [22]

    and Wesselbaum, D

    Peeters, R. and Wesselbaum, D. (2023). Competitiveness in Formula One.Sports Economics Review, 2:100007

  23. [23]

    and Torgler, B

    Schreyer, D. and Torgler, B. (2018). On the role of race outcome uncertainty in the TV demand for Formula 1 Grands Prix.Journal of Sports Economics, 19(2):211–229

  24. [24]

    Szymanski, S. (2003). The economic design of sporting contests.Journal of Economic Literature, 41(4):1137–1187

  25. [25]

    Zeileis, A., Leisch, F., Hornik, K., and Kleiber, C. (2002). strucchange: An R package for testing for structural change in linear regression models.Journal of Statistical Software, 7(3-4):1–38. 15 Appendix Table A.1: The number of drivers and races in Formula One seasons Season Drivers Races Season Drivers Races 1950 81 7 1988 36 16 1951 84 8 1989 47 16 ...