Pith. sign in

REVIEW 2 major objections 4 minor 2 cited by

Paired comparison models with strength-dependent ties and order effects

T0 review · 2 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper proposes that in chess-like paired comparisons, both draw probability and the white-pieces order advantage grow with the average strength of the two players, and that a model encoding this beats constant-effect alternatives on…

desk verdict A solid extension of David's paired comparison model with a well-supported strength-dependent draw effect, but the strength-dependent order effect is not backed by the paper's own posterior interval. read the letter →

arxiv 2505.24783 v1 pith:KQYPEZNP submitted 2025-05-30 stat.ME

classification stat.ME MSC 62F0762F1562J12
keywords pairedcomparisonmodelsBradley-Terrymodelchesstournamentsordereffectstieoutcomesstrength-dependentparametersBayesianDICUSOpen
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes that in head-to-head competitions with draws, ties are not equally likely at every level of skill, and neither is the order advantage (playing white in chess). It claims that both the probability of a draw and the size of the color advantage grow with the average strength of the two players, and builds this into a paired-comparison model by replacing the constant tie and order parameters with strength-dependent versions. On 24,888 US Chess Open games from 2006 to 2019, the full model fits substantially better than the standard constant-effect model (DIC 43,821.3 vs 44,658.4 under an informative prior), with the draw effect clearly positive and the order-effect coefficient positive but with a 95% interval touching zero. If the paper is right, analyses that use constant tie and home-field parameters are misspecified in chess-like settings, and rankings built on them would be biased in a strength-dependent way.

What carries the argument

The machine is equation (6), a three-outcome multinomial logit. For a match between $i$ and $j$, the win and loss numerators carry the order effect $\alpha_0+\alpha_1(\theta_i+\theta_j)/2$ inside terms of the form $\exp(\theta_i \pm x_{ij}(\alpha_0+\alpha_1(\theta_i+\theta_j)/2)/4)$, and the draw numerator is $\exp(\beta_0+(1+\beta_1)(\theta_i+\theta_j)/2)$, all normalized by their sum $D_{ij}$. The object doing the work is the pair-average strength $(\theta_i+\theta_j)/2$, which enters both the tie log-odds and the order effect; setting $\alpha_1=\beta_1=0$ recovers the constant-parameter model, making that model a nested special case and giving the comparison a direct likelihood-based interpretation.

What would settle it

Re-estimate the model on a large chess dataset, or on another tie-heavy sport, with smooth splines replacing the linear average-strength terms and compare DIC; if the spline version fits better, or if the posterior interval for $\alpha_1$ is not positive, then the strength-dependent order-effect claim is unsupported even though the draw side may survive.

Watch

Extended reading notes

Core claim

The central claim is that strength-dependence belongs in the model itself, not in the residual. Specifically, the paper's equation (6) replaces the fixed tie log-odds with $\exp(\beta_0+(1+\beta_1)(\theta_i+\theta_j)/2)$, so that $\beta_1 > 0$ means stronger pairs draw more often, and replaces the fixed order effect with $\alpha_0+\alpha_1(\theta_i+\theta_j)/2$, so that $\alpha_1>0$ means the white advantage grows with average strength. The full model has $\beta_1=0.120$ with 95% interval (0.103, 0.138) and $\alpha_1=0.037$ with 95% interval (0.000, 0.074), and beats all six variants tried on the US Open data, including the constant-parameter model, by a DIC margin of more than 800. The win-versus-loss odds remain independent of the tie parameters, so the extension preserves the standard property of the tie model while letting draw and order effects move with skill.

Load-bearing premise

The load-bearing premise is that the order effect grows linearly with average player strength, a premise that is fragile because the posterior interval for $\alpha_1$ includes zero, so if the true coefficient is not positive or the growth is not linear the order-effect half of the central claim collapses even though the model still fits better through the draw parameter $\beta_1$.

Editorial extensions

If this is right

  • If the central claim is correct, the constant tie and order-effect model that has been standard for paired comparisons is misspecified for chess-like data, with a DIC gap of roughly 835 points.
  • The model predicts that two strong evenly matched players draw about 49% of the time, versus about 24% for two average evenly matched players, a pattern testable in any large chess database.
  • The white advantage in win-to-loss odds rises from about 1.2 for average pairs to about 1.244 for strong pairs, so rankings that ignore this will slightly overstate weaker players' results when they play white.
  • The same three-outcome structure can be fitted with standard Bayesian software to any paired-comparison sport with draws and home advantage, such as soccer or hockey.
  • Because the model nests the constant-parameter model, existing likelihood-ratio and DIC comparisons can be used to test for strength dependence in new datasets without changing the estimation setup.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The draw-side of the claim is well supported; the order-effect side is fragile because $\alpha_1$'s posterior interval includes zero, so a conservative reading would treat only the strength-dependent tie rate as established.
  • In soccer-like settings where home advantage and draws both matter, the same structure transfers directly: one would expect draw rates to rise when two top teams meet, and a nonlinear version could be tested with splines.
  • Because player-years are treated as independent, the strength-dependent effects could be confounded with rating inflation or age trends; a longitudinal version would separate these dynamics.
  • If the true tie effect is nonlinear, the linear term $\beta_1$ may still win by DIC while mispredicting at very high strengths; fitting a monotone spline version would be a sharper test of the mechanism.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper proposes an extension of David (1988) paired comparison models in which both the tie probability and the order effect depend on the average strength of the two players (Eq. (6)). The model is motivated by GAM analyses on FIDE data, fitted to 24,888 US Chess Open games under six model variants and two priors, and compared by DIC. The strength-dependent tie parameter is estimated as β1 = 0.120 with interval (0.103, 0.138), supporting the first phenomenon. The strength-dependent order-effect parameter is estimated as α1 = 0.037 with interval (0.000, 0.074), which includes zero, and the corresponding GAM is explicitly described as less conclusive.

Significance. If taken as a modeling proposal, the paper contributes a clean, interpretable extension of a standard paired comparison framework, and the first empirical phenomenon (stronger pairs draw more often) is well supported by a tight posterior interval. The design is also a strength: the exploratory GAM analysis uses FIDE data while the confirmatory model fitting uses US Open data, reducing circularity concerns. However, the second headline phenomenon is not established by the reported posterior, and the in-sample DIC comparison does not by itself justify the language of predictive performance. The result is a sound model with one strong empirical finding and one suggestive but unconfirmed finding.

major comments (2)
  1. [Section 3, Eq. (6); Section 4, Table 4] The strength-dependent order effect is not empirically established. The posterior mean α1 = 0.037 has a 95% central interval (0.000, 0.074), so the data are consistent with α1 = 0; the GAM motivating this effect (Figure 2) is itself described in Section 2 as 'less conclusive'. The statement in Section 4 that 'The value of α1 being inferred to be positive indicates that the advantage to playing white increases slightly as a function of the players' strengths' is therefore an overstatement. Since the abstract and introduction present strength-dependent order effects as one of the two phenomena the model addresses, this is load-bearing for the central claim. The paper should either soften the claim to a model assumption with suggestive support or provide additional evidence that discriminates α1 > 0, such as out-of-sample prediction, prior sensitivity analysis, or a smoother estimate of the average-strength effect.
  2. [Section 4, Table 3; Section 5] The claim of 'clear improvements in predictive performance' is based on DIC values computed on the same data used for fitting. DIC is an in-sample approximation to out-of-sample predictive error, not a validation exercise, and the DIC ordering across nested models is not obviously stable: under the informative prior, Model 2 (no color effect) has lower DIC than Model 3 (constant color effect), despite the posterior for α0 being positive. To support the predictive language, the paper should add out-of-sample or cross-validated predictive comparisons, or at minimum change the wording to 'better fit' and discuss the instability.
minor comments (4)
  1. [Section 3, Eq. (8)] The notation in Eq. (8) uses K for the number of games and then refers to θ = (θ1, ..., θK), although θ has n player components; use n for the number of players to avoid confusion.
  2. [Section 4, prior specification] The invariance claim 'our model is invariant to additive shifts to the strength parameters' is correct only as a joint transformation: under θ_i → θ_i + c, invariance requires α0 → α0 − α1 c and β0 → β0 − β1 c. Stating this explicitly would clarify the role of the rating-scale origin.
  3. [Table 3] The DIC comparison shows that Model 2 (no color effect) outperforms Model 3 (constant color effect) under the informative prior, which is surprising in light of the posterior for α0; a sentence of explanation would help readers interpret the model-selection results.
  4. [Section 2, GAM details] For the GAM analyses, reporting the effective degrees of freedom and the basis used for the smoothing splines would improve reproducibility; the current text only says 'cubic smoothing spline'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central parameters are estimated from data, not defined into existence.

full rationale

Walking the claimed derivation chain, the paper proposes an explicit parametric extension of David's model in equation (6), fits it by MCMC to US Chess Open game outcomes, and compares DIC values across special cases (Table 3) and posterior intervals (Table 4). The two substantive claims, that stronger pairs draw more often and that order effects grow with strength, are empirical inferences from the free parameters beta_1 and alpha_1, not consequences of the model definition. beta_1 is identified by the likelihood, and its posterior interval (0.103, 0.138) is data-driven; alpha_1's interval includes 0, which is a statistical-evidence weakness, not circularity. The exploratory GAM on FIDE data (Figures 1 and 2) is independent motivation, and the confirmatory fit on a separate US Open dataset is an appropriate validation design rather than a circular reuse of the same fitted values. The only self-citation, Glickman and Jones (2024), supports the standard Elo-to-Bradley-Terry rating conversion used to set prior means; it is not load-bearing for the central inference, and the conversion is a conventional external result. The model's lack of invariance to additive shifts when alpha_1 and beta_1 are nonzero is a modeling and identifiability concern, but it is not a case of a fitted input being renamed as a prediction or of an equation reducing to its own input. No step in the paper exhibits the required pattern of circularity.

Assumptions & free parameters 9 free parameters · 4 assumptions · 0 invented entities

The model has seven fitted parameters plus two hand-set prior hyperparameters. The central claim depends most on α1 and β1: β1 is well identified, while α1 is not significantly different from zero. The axioms reflect the IIA structure, the rating-to-strength scale mapping, the Bayesian prior assumptions, and the independence of repeated player appearances.

free parameters (9)
  • α0 = 0.363 (0.289, 0.434)
    Baseline white advantage for an average-strength pair, fitted to US Open data, Table 4.
  • α1 = 0.037 (0.000, 0.074)
    Slope of the color advantage with average pair strength. The posterior interval includes 0, so it is not clearly nonzero.
  • β0 = -0.471 (-0.505, -0.437)
    Baseline log draw propensity for an average-strength pair, fitted to US Open data, Table 4.
  • β1 = 0.120 (0.103, 0.138)
    Increase in log draw propensity per unit average strength. Tightly estimated and well supported.
  • σ = 0.645 (0.604, 0.687)
    Standard deviation of the informative strength prior around the pre-tournament rating transform, Table 4.
  • μ_miss = -3.399 (-4.110, -2.703)
    Prior mean strength for unrated players, estimated from data, Table 4.
  • σ_miss = 2.636 (2.085, 3.310)
    Prior standard deviation for unrated players, estimated from data, Table 4.
  • Strength prior hyperparameters
    Inverse-Gamma(0.01, 0.1) chosen by hand as diffuse priors on variance parameters.
  • Non-strength parameter prior variances
    Normal priors with mean 0 and variance 100 chosen by hand for α0, α1, β0, β1.
assumptions (4)
  • domain assumption Luce's choice axiom / IIA: outcome probabilities are proportional to exponentials of linear terms, with the draw treated as a third multinomial category.
    Used to specify equation (6); the model inherits IIA from Davidson (1970) and David (1988), and this structure is assumed rather than derived.
  • domain assumption US Chess and FIDE ratings are linear transformations of Bradley-Terry log-strength parameters, so μ_i=(R_i-1500)/(400/log10) is the prior mean for θ_i.
    Equation (10) and Section 2 rely on this mapping; if the rating scale is not exactly this transform, the informative priors are miscalibrated.
  • domain assumption Independent normal priors on θ, α0, α1, β0, β1 and diffuse inverse-Gamma priors on variances yield a proper posterior and converged MCMC.
    Section 4 describes the priors; convergence is only checked via Rhat below 1.01, not by formal assessment of posterior propriety.
  • domain assumption Players who appear in multiple US Opens can be treated as distinct players, so repeated participation does not create within-player correlation.
    Section 4 explicitly states this choice and calls it conservative, but it is still an assumption about the data-generating process.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Paired comparison models with strength-dependent ties and order effects." pith.science (2026). https://pith.science/paper/KQYPEZNP

@misc{pith2026250524783,
  author       = {Pith},
  title        = {Pith review of: Paired comparison models with strength-dependent ties and order effects},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KQYPEZNP}},
  note         = {Machine review of arXiv:2505.24783}
}
read the original abstract

Paired comparison models, such as the Bradley-Terry (1952) model and its variants, are commonly used to measure competitor strength in games and sports. Extensions have been proposed to account for order effects (e.g., home-field advantage) as well as the possibility of a tie as a separate outcome, but such models are rarely adopted in practice due to poor fit with actual data. We propose a novel paired comparison model that accounts not only for ties and order effects, but recognizes two phenomena that are not addressed with commonly used models. First, the probability of a tie may be greater for stronger pairs of competitors. Second, order effects may be more pronounced for stronger competitors. This model is motivated in the context of tournament chess game outcomes. The models are demonstrated on the results of US Chess Open game outcomes from 2006 to 2019, large tournaments consisting of players of wide-ranging strengths.

Figures

Figures reproduced from arXiv: 2505.24783 by the authors.

Figure 1
Figure 1. Log-odds of white drawing as a function of the average rating between two players, [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Log-odds of white winning as a function of the average rating between two players, [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Each panel shows outcomes probabilities for a fixed [PITH_FULL_IMAGE:figures/full_fig_p025_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Each panel shows outcomes probabilities for a fixed [PITH_FULL_IMAGE:figures/full_fig_p026_4.png]
Figure 5
Figure 5. Figure 5: Each panel shows outcomes probabilities for a fixed [PITH_FULL_IMAGE:figures/full_fig_p027_5.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tied Pools and Drawn Games

    math.ST 2025-07 conditional novelty 6.0 of 10

    Draws should count as half-wins when estimating chess strength, with draw frequency estimated afterward; this justifies Glickman's rule and is applied to an 1821 chess pool.

  2. Rating competitors in games with strength-dependent tie probabilities

    stat.ME 2025-06 conditional novelty 5.0 of 10

    A Bayesian rating system that models draw probability as increasing with player strength, with fast one-step posterior updates, is developed and deployed on ICCF chess data.

Reference graph

Works this paper leans on

4 extracted references · 4 canonical work pages · cited by 2 Pith papers

  1. [1]

    Baker, R. D. & Scarf, P. A. (2020), ‘Modifying Bradley–Terry and other ranking models to allow ties’,IMA Journal of Management Mathematics32(4), 451–463. Bezdek, J. & Hathaway, R. (2002), Some notes on alternating optimization,inN. R. Pal & M. Sugeno, eds, ‘Advances in Soft Computing —AFSS 2002’, Springer, Berlin, Heidelberg, pp. 288–300. 22 Bradley, R. A...

  2. [9]

    Plummer, M. (2003), JAGS: A program for analysis of Bayesian graphical models using Gibbs sampling,in‘Proceedings of the 3rd International Workshop on Distributed Statistical Computing (DSC 2003). March’, pp. 20–22. 23 Spiegelhalter, D. J., Best, N. G., Carlin, B. P. & Van Der Linde, A. (2002), ‘Bayesian measures of model complexity and fit’,Journal of th...

  3. [333]

    & Tibshirani, R

    Hastie, T. & Tibshirani, R. (1986), ‘Generalized additive models’,Statistical science 1(3), 297–310. Ma, S. (2015), ‘Alternating proximal gradient method for convex minimization’,Journal of Scientific Computing68, 546 –

  4. [572]

    (1951), ‘Remarks on the method of paired comparisons: I

    Mosteller, F. (1951), ‘Remarks on the method of paired comparisons: I. The least squares so- lution assuming equal standard deviations and equal correlations’,Psychometrika16(1), 3–

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.