Pith. sign in

REVIEW 4 major objections 3 minor 16 references

Tied Pools and Drawn Games

T0 review · 4 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper argues that two draws carry the same strength information as a win and a loss, so strength should be estimated with the draw parameter fixed at two before draw propensity is estimated.

desk verdict Careful algebra and a useful connection to Glickman, but the central premise—two draws equal a win and a loss—is asserted, not proven. read the letter →

arxiv 2507.03894 v1 pith:CEZK54WK submitted 2025-07-05 math.ST stat.MEstat.TH

classification math.STstat.MEstat.TH MSC 62F0762J1562F10
keywords pairedcomparisonswithtiesthree-waymaximum-likelihoodestimationBradley-TerrymodelDavidsondrawpropensitychesspoolshalf-winscoring
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that when games can end in draws, strength ratings should be computed in two stages rather than all at once: first fix the draw parameter at the value that treats two draws as equivalent to a win and a loss, estimate player strengths, and only then estimate how draw-prone the players are. The target is the standard Davidson extension of the Bradley-Terry model, which estimates strengths and draw propensity simultaneously; the paper claims this lets draw frequency distort strength estimates, sometimes forcing infinite strength gaps when draws dominate. The proposed Constrained Alternative model keeps the ordering of strengths consistent with scores in balanced tournaments while avoiding the counterexamples that sink the unconstrained version. Applied to the incompletely recorded three-player chess pools of Paris 1821, the two-stage procedure gives relative ratings of about 238 and 140 for de la Bourdonnais and the handicapped Deschapelles above Cochrane.

What carries the argument

The load-bearing identity is the $\nu=2$ special case of the Alternative model: with $\sigma$ the strength parameters, $p_{ij} = (\sigma_i/(\sigma_i+\sigma_j))^2$ and $d_{ij} = 2\sigma_i\sigma_j/(\sigma_i+\sigma_j)^2 = 2\sqrt{p_{ij}p_{ji}}$. This makes a single game behave like two independent Bradley-Terry comparisons, with a draw the mixed outcome (one win, one loss), so draw frequency is exactly the information that should not move strength ratios. The two-stage maximum-likelihood procedure, iterating $\sigma$ with $\nu$ fixed at 2 and then updating $\nu$ with $\sigma$ fixed, is the mechanism that carries the argument, and it also lets incomplete pool data be completed by borrowing $\nu$ from matches with full draw information.

What would settle it

Fit the unconstrained and constrained models to complete game-by-game records from a balanced three-player pool competition, and compare each model's predicted draw counts ($t_{ij} = \nu\sqrt{s_{ij}s_{ji}}$ under the constrained procedure) with the actual draws; systematic divergence, or disagreement between the constrained strengths and an independent rating, would refute the $\nu=2$ equivalence.

Watch

Extended reading notes

Core claim

The central claim is that Davidson's extension of the Bradley-Terry model is being used in the wrong order: estimating strength parameters and draw propensity simultaneously lets draw frequency distort strength estimates, producing infinite strength gaps when one player never wins and, in the unconstrained Alternative model, score-strength inversions. The paper's proposed Constrained Alternative model fixes the draw parameter at $\nu=2$ while estimating strengths, then fixes those strengths and estimates $\nu$. At $\nu=2$, the Alternative model collapses to win probability $p_{ij} = (\sigma_i/(\sigma_i+\sigma_j))^2$, so a draw contributes exactly the likelihood of a win followed by a loss; two draws carry the same strength information as a win and a loss. The paper argues this is the most correct use of Davidson's method, that it satisfies the Consistency property, and that it gives Glickman's draw-as-half-win rule a theoretical basis. For the 1821 Paris pools the procedure yields relative ratings of about 238 for de la Bourdonnais and 140 for the handicapped Deschapelles above Cochrane.

Load-bearing premise

The load-bearing premise is that two draws carry the same strength information as a win and a loss, so draw propensity should not affect strength estimates; if draws instead reflect caution, risk aversion, or collusion, the constrained procedure will bias strengths.

Editorial extensions

If this is right

  • Strength ratings can be computed from half-point scores alone, before any draw-propensity estimate is made, because fixing $\nu=2$ makes draws contribute exactly half a win to the likelihood.
  • Davidson-style simultaneous estimation is replaced: the constrained procedure keeps scores and strengths in the same order in balanced three-player designs, which the unconstrained Alternative model fails to do.
  • Glickman's half-win scoring is a consequence of a draw model, not an ad hoc rule, and it still permits a draw-propensity parameter to be estimated in a second stage.
  • In the 1821 pools, the two-stage procedure rates de la Bourdonnais about 238 points above Cochrane and the handicapped Deschapelles about 140 points above Cochrane.
  • When pool results omit draws, expected draw counts between pairs can be recovered from encounter counts by $t_{ij} = \nu\sqrt{s^e_{ij}s^e_{ji}}$, using a draw propensity estimated from complete historical matches.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension the paper leaves implicit is to fit the unconstrained Alternative model to complete modern tournament data and ask whether maximum-likelihood $\nu$ clusters near 2; values far from 2 would not overturn the half-win strength rule but would say the draw submodel itself needs enrichment.
  • The two-stage principle suggests a generic estimation recipe for other tie models such as Rao-Kupper: fix the tie parameter at a benchmark value for structural estimation, then profile it afterward, because simultaneous estimation is the generic source of the distortion identified here.
  • The 1821 ratings inherit the borrowed $\nu=0.4815$ from 1821-1836 matches; because the paper itself notes draw rates vary over time and with player strength, a sensitivity analysis of the 238 and 140 gaps across the roughly 0.37 to 0.53 range of plausible $\nu$ would quantify that dependence.
  • The 'game as two comparisons' interpretation has a concrete variance prediction: with draws, the effective number of comparisons per game doubles relative to no-draw Bradley-Terry, halving the uncertainty of strength estimates, which is checkable in rating-system calibration.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. This paper addresses the estimation of strength parameters in three-way comparison experiments, using 1821 chess pools as the running example. It proposes an Alternative model for paired comparisons with draws in which strength ratios equal expected score ratios rather than win ratios (Property 3), and recommends a two-stage 'Constrained Alternative' procedure: fix the draw-propensity parameter at ν=2 when estimating strengths, then estimate ν with strengths fixed. This is argued to be more sensible than Davidson's simultaneous estimation, to fix flaws of the unconstrained models, and to provide a theoretical justification for Glickman's draw-as-half-win rule. The paper derives MLE equations for all four models, proves the unconstrained Alternative model can violate score-consistency in balanced three-player designs, and applies the recommended procedure to estimate draw counts and strength ratings for the 1821 pools. The main algebraic derivations are internally consistent, but the central normative claim is an argued axiom rather than a theorem, and the illustrative application in Section 10 contains a presentation/implementation mismatch and a circularity that weaken its status as validation.

Significance. If the arguments are accepted, the paper contributes a principled motivation for a draw-as-half-win strength model and a practical recipe for handling incomplete three-way comparison data. The work is clearly written, the derivations are transparent, and the historical example is engaging. The central contribution—separating strength estimation from draw-propensity estimation—is an interesting alternative to Davidson's model, and the identification of the unconstrained Alternative model's order-reversal in balanced designs is a useful cautionary result. However, the paper's central recommendation rests on an asserted equivalence (two draws = a win and a loss) that is never derived or empirically tested, and the Section 10 illustration is a consistency check rather than an independent validation. These issues do not undermine the internal algebra but do affect how strongly the conclusions can be stated.

major comments (4)
  1. [Section 10, Eq. (39)] The sentence 'Running the Constrained Alternative model (29)' is misleading because Eq. (29) is the unconstrained Alternative iteration from Section 5.2 (with the Δ-acceleration factor 3sk/(Gk+2sk/σk)), not the Constrained Alternative procedure defined in Section 7.1, which updates σ using Gk(σ,2) and only then estimates ν. The reported value ν=0.4814882 agrees with the unconstrained Alternative estimate 0.4814897 to seven digits, whereas the Constrained Davidson estimate is 0.532327; this suggests the numbers were actually produced with the unconstrained model. The illustration should be re-run with the correct constrained iteration, and the discrepancy should be explained.
  2. [Section 10, Eqs. (39)-(40) and following paragraph] The expected draw counts in (40) are constructed from the chosen ν=0.4814882 via Eq. (38), and then the final model 're-estimates' ν=0.48149; the paper says this is 'as it should be, since we built the results assuming the value in (39).' This is a circular consistency check, not a validation of the draw-propensity estimate or of the resulting ratings r=(238.0, 0, 140.4). The text should clearly state that the historical ratings are conditional on an assumed ν borrowed from Table 5 and that the agreement only verifies internal algebraic consistency. As written, the presentation gives the impression of independent confirmation.
  3. [Section 5.2, Eq. (29)] The paper explicitly says: 'If this iteration is not guaranteed to converge, one could also use Newton's method, and we do not pursue this question further.' This gap is load-bearing because the numerical demonstration of order-reversal in Section 6.3 and Test 3 in Section 8.3 relies on iteration (29) producing the value σ1/σ2 = 0.9586. If that fixed-point iteration fails to converge or converges to a non-MLE, the counterexample is unsupported. The authors should either prove a convergence result for (29) (or a suitable descent property) or replace the numerical results with a safeguarded Newton/optimization method and confirm that the reported values remain.
  4. [Section 7, 'Reassessment'] The central recommendation to fix ν=2 when estimating strengths rests on 'It can reasonably be argued that two draws are equivalent to a win and a loss between a pair of players in assessing relative strength.' This is an asserted normative premise, not a theorem, and it is load-bearing: it licenses the two-stage procedure and disqualifies simultaneous estimation. The paper does not derive this equivalence from a decision-theoretic or predictive criterion, nor does it test its consequences when draw propensity reflects style, risk aversion, or collusion. The authors should explicitly label this as a modeling axiom and add a sensitivity analysis (e.g., compare strength estimates under ν=2 vs. ν estimated simultaneously, or a predictive comparison on held-out matches) to support the claim that the constrained procedure is 'most correct.'
minor comments (3)
  1. [Section 4.1, Eq. (23)] Equation (23) contains a typesetting error: the expression '(√πi√πi + √πj)^2' should be '(√πi/(√πi+√πj))^2'; as printed, the denominator is missing from the fraction.
  2. [Section 6.3] In the sentence 'When t = 4, there are t = 4 in independent parameters', 'in' is extraneous; it should read 't = 4 independent parameters.'
  3. [References] Reference [5] is listed as 'arXi:2505.24783v1'; the prefix should be 'arXiv:'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the two-stage prescription rests on an explicit assumption, and the Section 10 re-estimated draw parameter is openly labeled a built-in consistency check.

full rationale

The paper's derivation chain is self-contained. The Davidson and Alternative likelihood equations are derived algebraically from their stated assumptions; the nu=2 special case in Equation (22) follows from the Alternative model's Property 3; the Consistency counterexamples in Section 6 are direct computations; and the two-stage Constrained Alternative procedure in Section 7.1 is explicitly built from the stated premise that two draws carry the same strength information as a win and a loss. That premise is asserted rather than proven, so the recommendation to estimate strengths before draw propensity is conditional on it; this is a limitation of the argument, not a circular reduction. Section 10 contains a built-in loop: the expected draw counts are constructed from an externally estimated nu=0.4814882, and the model then returns nu=0.48149; however, the paper itself says 'as it should be, since we built the results assuming the value in (39)', thereby labeling it a consistency check and not an independent prediction. The final ratings (238, 0, 140) depend on the borrowed nu from Table 5, but that is an extrapolation made with stated assumptions, not a prediction equivalent to its own input. No load-bearing step reduces to a self-citation or to an imported uniqueness theorem. Hence no significant circularity.

Assumptions & free parameters 2 free parameters · 6 assumptions · 0 invented entities

The central method rests on the Davidson tie model, the Bradley-Terry strength representation, and the paper-specific equivalence assumption that two draws equal a win and a loss. The historical example adds the assumptions of uniform random byes and a draw propensity borrowed from other matches; the draw propensity parameter nu is the principal free parameter.

free parameters (2)
  • nu (draw-propensity parameter) = 0.4814882 (Constrained Alternative on Table 5; 0.532327 under Constrained Davidson)
    In the final 1821 application, nu is not estimated from the pool data (which lack draw information) but borrowed from the 1821-1836 match dataset, then used to construct expected draw counts via t_ij = nu * sqrt(s_e_ij * s_e_ji).
  • Per-player strength parameters in the Table 5 model = Not reported in full
    These are fitted as nuisance parameters to estimate nu from the 9-player match dataset; they are not the central outputs but are required by the estimation procedure.
assumptions (6)
  • ad hoc to paper Two draws are equivalent to a win and a loss for assessing relative strength
    Stated in Section 7 as 'It can reasonably be argued that two draws are equivalent to a win and a loss...'. This is the load-bearing modeling assumption behind the Constrained Alternative model.
  • domain assumption Davidson's draw relation d_ij = nu * sqrt(p_ij * p_ji)
    Inherited from Davidson (1970), used as Equation (9) in both the Davidson and Alternative models.
  • standard math Bradley-Terry strength representation p_ij = pi_i / (pi_i + pi_j)
    Equation (7); the standard paired-comparison model underlying the paper.
  • domain assumption The player initially left out in each 1821 pool is chosen uniformly at random or rotated
    Section 1: 'It is also likely that the player to be initially left out was either selected randomly in each round, or rotated through the rounds.' Used to compute expected game counts such as n_CD = (1 - (2/3) q_B) N / (1 - p_0).
  • domain assumption Drawn games are replayed until a decisive result, so the encounter win probability is e_BC = p_BC / (1 - d_BC)
    Equation (1); reflects the historical convention that draws are inconclusive and replayed.
  • domain assumption The Table 5 match results are representative of draw propensity for the 1821 pool players
    Used to fix nu = 0.4815 for the 1821 pools despite differences in era, format, and handicaps.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Tied Pools and Drawn Games." pith.science (2026). https://pith.science/paper/CEZK54WK

@misc{pith2026250703894,
  author       = {Pith},
  title        = {Pith review of: Tied Pools and Drawn Games},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CEZK54WK}},
  note         = {Machine review of arXiv:2507.03894}
}
read the original abstract

We consider the problem of estimating `preference' or `strength' parameters in three-way comparison experiments, each composed of a series of paired comparisons, but where only the single `preferred' or `strongest' candidate is known in each trial. Such experiments arise in psychology and market research, but here we use chess competitions as the prototypical context, in particular a series of `pools' between three players that occurred in 1821. The possibilities of tied pools, redundant and therefore unplayed games, and drawn games must all be considered. This leads us to reconsider previous models for estimating strength parameters when drawn games are a possible result. In particular, Davidson's method for ties has been questioned, and we propose an alternative. We argue that the most correct use of this method is to estimate strength parameters first, and then fix these to estimate a draw-propensity parameter, rather than estimating all parameters simultaneously, as Davidson does. This results in a model that is consistent with, and provides more context for, a simple method for handling draws proposed by Glickman. Finally, in pools with incomplete information, the number of drawn games can be estimated by adopting a draw-propensity parameter from related data with more complete information.

Figures

Figures reproduced from arXiv: 2507.03894 by the authors.

Figure 1
Figure 1. Win probabilities in the Davidson and Alternative model with various values of the draw [PITH_FULL_IMAGE:figures/full_fig_p011_1.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

16 extracted references · 15 canonical work pages

  1. [1]

    The rank analysis of incomplete block designs, 1. The method of paired comparisons,

    Bradley, R.A. and Terry, M.E. “The rank analysis of incomplete block designs, 1. The method of paired comparisons,” Biometrika, 39:324–345 (1952)

  2. [2]

    David, H. A. The Method of Paired Comparisons, 2nd edition (London, Chapman and Hall, 1988)

  3. [3]

    On extending the Bradley-Terry model to accommodate ties in paired comparison ex- periments

    Davidson, R.R. “On extending the Bradley-Terry model to accommodate ties in paired comparison ex- periments.” Journal of the American Statistical Association, 65:317–328 (1970)

  4. [4]

    Parameter estimation in large dynamic paired comparison experiments

    Glickman, M.E. “Parameter estimation in large dynamic paired comparison experiments.” Applied Statis- tics, 48: 377–394 (1999)

  5. [5]

    “Paired comparison models with strength-dependent ties and order effects

    Glickman, M.E. “Paired comparison models with strength-dependent ties and order effects. arXi:2505.24783v1 (2025)

  6. [6]

    A generalization of the Bradley-Terry model for draws in chess with an application to collusion

    Hankin, R.K.S., “A generalization of the Bradley-Terry model for draws in chess with an application to collusion.” Journal of Economic Behavior and Organization, 180: 325–333 (2020)

  7. [7]

    The Oxford Companion to Chess (Oxford, Oxford University Press, 1984)

    Hooper, D., and Whyld, K. The Oxford Companion to Chess (Oxford, Oxford University Press, 1984)

  8. [8]

    George Walker

    Murray, H.J.R. “George Walker.” British Chess Magazine, vol.26 (1906), pp.189–194

Show all 16 references
  1. [9]

    A History of Chess (Oxford, Oxford University Press, 1913)

    Murray, H.J.R. A History of Chess (Oxford, Oxford University Press, 1913)

  2. [10]

    Ranking in triple comparisons

    Pendergrass, R.N. and Bradley, R.A. “Ranking in triple comparisons.” in: Olkin, I, et al. (eds.), Con- tributions to Probability and Statistics (Stanford University Press, Stanford, 1960), pp. 331–351

  3. [11]

    Ties in paired-comparison experiments: a generalization of the Bradley- Terry model,

    Rao, P. V. and Kupper, L. L. “Ties in paired-comparison experiments: a generalization of the Bradley- Terry model,” Journal of the American Statistical Association, 62:194–204 (1967)

  4. [12]

    de Saint-Amant, P.C.F. (ed.). Le Palam` ede, new series, vol. 4 (1844)

  5. [13]

    A Century of British Chess (David McKay, Philadelphia, 1934)

    Sergeant, P.W. A Century of British Chess (David McKay, Philadelphia, 1934)

  6. [14]

    Staunton, H. (ed.). The Chess Player’s Chronicle, vol.1 (1841)

  7. [15]

    De la Bourdonnais versus McDonnell, 1834 (Jefferson, McFarland, 2005)

    Utterberg, C. De la Bourdonnais versus McDonnell, 1834 (Jefferson, McFarland, 2005)

  8. [16]

    The light and lustre of chess

    Walker, G. “The light and lustre of chess.” The Chess Player’s Chronicle 4:215–222, 248–254, 279–286 (1843). 33

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.