REVIEW 3 minor 32 references
Engine-Equal, Human-Unequal: A Reproducible Outcome Skew in Engine-Assessed Equal Chess Positions
T0 review · 0 major / 3 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read Among chess positions a strong engine calls dead even, human results are not balanced: positions carry small, stable side skews that reproduce across disjoint player groups, so the engine's evaluation is not a sufficient statistic for human
desk verdict Engine-equal positions carry small reproducible human outcome skews; sub-family selection is the honest soft spot, but this deserves serious refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the within-family replication slope: each position's outcome skew (the mean deviation of its games' results from a rating-calibrated expected score, in White's point of view) is measured once in each of two disjoint account groups, then the two measurements are regressed on each other after subtracting each opening-family's mean skew. Family demeaning removes opening-level repertoire selection, so a positive slope means positions within the same opening carry their own replicated tilt. The slope is estimated by weighted least squares with effective sample sizes that discount repeated games by the same account, and significance is assessed by permutation tests that sh
What would settle it
A randomized assigned-play trial, as the paper's companion study plans, that assigns players to both sides of engine-equal positions and finds no systematic per-position outcome difference by side would falsify the causal reading of the skew; alternatively, a re-analysis on a different platform's games (e.g., chess.com) that yields a within-family replication slope near zero would falsify the claim that the skew is a stable property of human play from these positions.
Extended reading notes
Core claim
The central discovery is a measurable, reproducible outcome skew at engine-equal positions: for 1,661 opening positions rated within 10 centipawns of zero by Stockfish 18 at high depth and reached at least 1,000 times by Lichess players in October 2025, the gap between actual results and rating-predicted results varies by position and persists when measurements are split across disjoint player accounts. In the primary split, a position's skew measured in one account group predicts the skew measured in a completely disjoint group at slope 0.69 (95% CI [0.65, 0.74], permutation p = 0.001); the disfavoured side also spends longer thinking. The authors emphasize that existence of the skew is the
Load-bearing premise
The result stands only if the replicated within-family skew is not substantially caused by which players choose to play a line—narrower sub-repertoires with different player pools within an ECO code—rather than by the position itself; the paper explicitly scopes its estimand to the naturally-reached position and defers causation to a randomized companion study.
Editorial extensions
If this is right
- Engine evaluation alone is insufficient to predict human results at even positions; practical difficulty is position-specific and reproducible.
- A replicated 0.05 skew corresponds to about a 50-point rating gap, and the largest replicated skews to 150 points or more, giving players a familiar currency for the imbalance.
- Most of the skew variance lies within opening families, so adjusting for opening choice alone cannot remove the imbalance.
- The think-time asymmetry gives a behavioral handle: the disfavoured side reliably spends more clock, which could inform training tools and interface design.
- Out-of-sample replication eight months later implies the effect is not a one-month artifact or a quirk of a single player population.
Reading between the lines
- If the skew is partly caused by preparation or repertoire selection, then other games with machine ground truth—such as Go, poker, or AI-assisted education—may carry similar human-specific imbalances that scalar evaluations miss.
- The monotonic rise of the replication slope with position popularity suggests that rare positions may have unmeasured skews; larger corpora or adaptive sampling could map the full landscape.
- The within-family residual selection confound could be tested by comparing skews of transpositions that reach the same position via different move orders; if the skew tracks the move order, the property is not the position alone.
- Pairing think-time data with move-quality measures (e.g., error rates) could separate 'harder' from 'less studied' and sharpen the causal question for the randomized companion study.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies 1,661 opening positions from Lichess October 2025 games that Stockfish 18 evaluates as equal (|eval| ≤ 10cp at depth 28, depth-stable) and that humans reach at least 1,000 times. For each position it computes a skew δ — the mean deviation of game results from a rating-calibration fractional-logit expectation — and tests whether δ measured in one random half of player accounts predicts δ in the disjoint other half after removing ECO-family means. The within-family weighted replication slope is β̂ = 0.691 (family-clustered CI [0.646, 0.736]; permutation p = 0.001), with similar results on temporal, rating-band, and external-month (June 2026, β = 0.904) axes. Robustness checks include a covariate-matched placebo (0.046), a covariate-means re-fit (0.599), within-account demeaning (0.526), line-cluster collapse (0.681), and time-control stratification. A secondary clock analysis finds the disfavoured side spends more think time. The paper explicitly scopes the estimand to the 'naturally-reached position' and disclaims causation, deferring to a pre-registered randomised study.
Significance. If correct, the paper gives a large-scale, carefully identified demonstration that a scalar engine evaluation near zero is not a sufficient statistic for human outcomes at the level of individual positions, and that the residual is reproducible rather than noise. The paper's defensive design is a major strength: the panel and estimator are fixed before the external month is read; the permutation null is calibrated on synthetic data; the covariate-matched placebo bounds calibration-misspecification; secondary analyses are labelled as such; and the code/data release supports independent verification. The main interpretive caveat — sub-family player selection — is acknowledged in §5.3 and is compatible with the scoped claim, though it limits causal or context-free readings. The paper does not rely on machine-checked proofs, but its reproducible code and archived dataset are strong assets.
minor comments (3)
- [§5, §5.3] The Discussion's sentence 'Most of the skew’s variance lies within opening family, so it attaches to specific decision problems rather than to the openings players choose' overstates what the design can establish. §5.3 correctly acknowledges that sub-repertoire selection (line-specific preparation, repertoire comfort, transposition history) can generate a positive within-ECO replication slope. Every replication axis preserves the natural selection process into positions, so the data cannot separate position-level difficulty from line-selection. Recommend rephrasing to 'finer-grained than the ECO family' or explicitly attaching the claim to the naturally-reached position as defined in §5.3.
- [Figure 1] The caption states the fit sample has 1,630 positions (1,661 minus 31 single-member-family positions), but panel (a) labels the plotted sample as n = 1,624. Please reconcile the count or explain the additional six excluded positions.
- [Table 2 / §4.1 (fast-moving-rating filter)] The 'fail(band)' verdict for the fast-moving-rating filter is interpreted as mechanical attenuation because the ICC drops to 0.41/0.32 on the filtered sample. This is plausible, but the paper should state explicitly that the ICC is recomputed on the filtered data, so the attenuation explanation is not an independent check. A brief sentence noting this would prevent over-reading.
Circularity Check
No significant circularity: the replication slope is a cross-cell regression of position-level residuals, and the calibration that could in principle manufacture skew omits position identity by construction.
full rationale
The paper's central claim is a reproducibility claim about position-level skews, and the derivation chain does not reduce to its inputs. The expected-score calibration is fitted on the anchor cell and applied unchanged to the other cell; position identity is deliberately not a calibration covariate, as the paper states: "Position identity is not a calibration covariate, so the calibration cannot absorb position-level skew" (§3.4). The skew is a residual mean aggregated per position, so the replication slope is a regression of one independently measured position-level quantity on another, not a refit of a fitted parameter. The covariate-matched placebo (slope 0.046 vs. headline 0.691) and the covariate-means re-fit (0.599) directly test whether calibration misspecification could manufacture the co-skew, and they bound that channel rather than assuming it away. The within-family demeaning removes family-level repertoire effects, and the external-month replication—"restricting June to games in which neither account appears among the October analysed games ... gives a slope of 1.01" (§4.1)—provides out-of-sample evidence not generated by the same fitted values. The paper's own §5.3 limitation, that sub-family repertoire selection remains inside the naturally-reached-position estimand, is a scope boundary and an acknowledged non-causal caveat, not a circular step: it does not define the replication slope in terms of the outcome, and the causal question is explicitly deferred to a randomized study. There are no load-bearing self-citations, no imported uniqueness theorems, and no ansatz smuggled in via citation; the discussion explicitly credits the raw skew to Lichess folklore and frames the contribution as an audit. The central result is therefore self-contained against the data and does not reduce to its construction.
Assumptions & free parameters
free parameters (3)
- Rating-calibration GLM coefficients =
fitted on group A; parity edges 0.515/0.510/0.505 (blitz/rapid/classical)
- Engine-equal membership thresholds =
|eval| ≤ 10cp at depth 28; last three ladder depths within 50cp; discovery-half occurrence ≥ 1000
- Line-cluster overlap threshold =
0.50 conditional game overlap
assumptions (5)
- domain assumption Lichess Glicko-2 ratings and rating gaps, with time-control interactions, are a valid expected-score calibration after fractional-logit fitting.
- ad hoc to paper Stockfish 18 depth-28 evaluation with depth-stability is the operational definition of 'engine-equal'.
- domain assumption ECO codes are a meaningful opening-family partition for absorbing family-level repertoire selection.
- standard math The within-family permutation test with 1,000 draws and the add-one p-value is calibrated.
- domain assumption Lichess game records of results, ratings, clocks, account IDs, and ECO tags are accurate enough for the analysis.
Cite this review
Pith. "Pith review of Engine-Equal, Human-Unequal: A Reproducible Outcome Skew in Engine-Assessed Equal Chess Positions." pith.science (2026). https://pith.science/paper/UTXYKUZ5
@misc{pith2026260725655,
author = {Pith},
title = {Pith review of: Engine-Equal, Human-Unequal: A Reproducible Outcome Skew in Engine-Assessed Equal Chess Positions},
year = {2026},
howpublished = {\url{https://pith.science/paper/UTXYKUZ5}},
note = {Machine review of arXiv:2607.25655}
}
abstract
Among chess opening positions that a strong engine judges essentially equal (Stockfish 18 evaluation within 10 centipawns of zero, depth-stable) and that humans actually reach on Lichess (October 2025; 1,661 positions, 16.1M occurrences), human results are not balanced. Positions carry outcome skews, each the gap between its games' actual results and what the players' ratings predict, whose directions are stable properties of the naturally-reached position: some positions favour White, others Black. These skews reproduce across three re-partitions -- disjoint player-account sets (primary), time, and disjoint rating bands -- and on an out-of-sample month eight months later. On the primary split, each position's skew is measured once in each account group, and the replication slope asks how well one measurement predicts the other after removing rating and opening-family effects: one means undiminished carry-over; zero, no linear relation. We find 0.69 (family-clustered 95% CI [0.65, 0.74]), rising to 0.94 on the most-popular, best-measured positions. The slope's value depends on the position mix. Existence is the invariant claim: it survives every tighter evaluation band, search depth, calibration, and popularity cutoff we test, and replicates within blitz and rapid separately. The typical skew is small (median $|\delta| \approx 0.018$, about two percentage points of White score), yet it reproduces, position by position, across disjoint accounts. At these positions the disfavoured side also thinks longer. Even where the evaluation is most confident, it is not a sufficient statistic for human outcomes. The result is observational, and the causal question is left to a pre-registered randomised companion study.
Figures
Reference graph
Works this paper leans on
-
[1]
Assessing human error against a benchmark of perfection
Ashton Anderson, Jon Kleinberg, and Sendhil Mullainathan. Assessing human error against a benchmark of perfection. InProceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), pages 705–714, 2016. doi: 10.1145/2939672. 2939803
doi:10.1145/2939672 2016
-
[2]
Ashton Anderson, Jon Kleinberg, and Sendhil Mullainathan. Assessing human error against a benchmark of perfection.ACM Transactions on Knowledge Discovery from Data (TKDD), 11 (4):45:1–45:25, 2017. doi: 10.1145/3046947
-
[3]
Not all Chess960 positions are equally complex.arXiv preprint, 2025
Marc Barthelemy. Not all Chess960 positions are equally complex.arXiv preprint, 2025. arXiv:2512.14319v3
arXiv 2025
-
[4]
Fragility of chess positions: Measure, universality, and tipping points
Marc Barthelemy. Fragility of chess positions: Measure, universality, and tipping points. Physical Review E, 111:014314, 2025. doi: 10.1103/PhysRevE.111.014314
-
[5]
Chess variation entropy and engine relevance for humans.arXiv preprint,
Marc Barthelemy. Chess variation entropy and engine relevance for humans.arXiv preprint,
-
[6]
Yoav Benjamini and Yosef Hochberg. Controlling the false discovery rate: A practical and powerful approach to multiple testing.Journal of the Royal Statistical Society: Series B (Methodological), 57(1):289–300, 1995. doi: 10.1111/j.2517-6161.1995.tb02031.x
arXiv 1995
-
[7]
Zipf’s law in the popularity distribution of chess openings
Bernd Blasius and Ralf Tönjes. Zipf’s law in the popularity distribution of chess openings. Physical Review Letters, 103:218701, 2009. doi: 10.1103/PhysRevLett.103.218701
-
[8]
A. Colin Cameron and Douglas L. Miller. A practitioner’s guide to cluster-robust inference. Journal of Human Resources, 50(2):317–372, 2015. doi: 10.3368/jhr.50.2.317. 25
Show all 32 references
-
[9]
Joseph Hoane, and Feng-hsiung Hsu
Murray Campbell, A. Joseph Hoane, and Feng-hsiung Hsu. Deep Blue.Artificial Intelligence, 134(1–2):57–83, 2002
2002
-
[10]
Quantifying human performance in chess.Scientific Reports, 13:2113, 2023
Sandeep Chowdhary, Iacopo Iacopini, and Federico Battiston. Quantifying human performance in chess.Scientific Reports, 13:2113, 2023. doi: 10.1038/s41598-023-27735-9
2023 doi
-
[11]
Why opening statistics are hard
D2D4C2C4. Why opening statistics are hard. Lichess community blog,https://lichess. org/@/D2D4C2C4/blog/why-opening-statistics-are-hard/9g61F9Uc, 2024. accessed 2026- 07-14
2024
-
[12]
Elo.The Rating of Chessplayers, Past and Present
Arpad E. Elo.The Rating of Chessplayers, Past and Present. Arco Publishing, New York, 1978
1978
-
[13]
Thompson
Chris Frost and Simon G. Thompson. Correcting for regression dilution bias: Comparison of methods for a single predictor variable.Journal of the Royal Statistical Society: Series A (Statistics in Society), 163(2):173–189, 2000. doi: 10.1111/1467-985X.00164
-
[14]
Glickman
Mark E. Glickman. Example of the Glicko-2 system. Technical note, Boston University, http://www.glicko.net/glicko/glicko2.pdf, 2022. revision of 22 March 2022; accessed 2026-07-16
2022
-
[15]
Cognitive performance in competitive environments: Evidence from a natural experiment.Journal of Public Economics, 139:40–52,
Julio González-Díaz and Ignacio Palacios-Huerta. Cognitive performance in competitive environments: Evidence from a natural experiment.Journal of Public Economics, 139:40–52,
-
[16]
Assessing the difficulty of chess tactical problems.International Journal on Advances in Intelligent Systems, 7(3&4):728–738, 2014
Dayana Hristova, Matej Guid, and Ivan Bratko. Assessing the difficulty of chess tactical problems.International Journal on Advances in Intelligent Systems, 7(3&4):728–738, 2014
2014
-
[17]
John Wiley & Sons, New York, 1965
Leslie Kish.Survey Sampling. John Wiley & Sons, New York, 1965
1965
-
[18]
Indoor air quality and strategic decision making
Steffen Künn, Juan Palacios, and Nico Pestel. Indoor air quality and strategic decision making. Management Science, 69(9):5354–5377, 2023. doi: 10.1287/mnsc.2022.4643
2023
-
[19]
Accuracy / win% model
Lichess. Accuracy / win% model. https://lichess.org/page/accuracy, 2023. accessed 2026-07-14
2023
-
[20]
Aligning super- human AI with human behavior: Chess as a model system
Reid McIlroy-Young, Siddhartha Sen, Jon Kleinberg, and Ashton Anderson. Aligning super- human AI with human behavior: Chess as a model system. InProceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), pages 1677–1687, 2020. doi...
2020
-
[21]
Chessformer: A unified architecture for chess modeling
Daniel Monroe, George Eilender, Philip Chalmers, Zhenwei Tang, and Ashton Anderson. Chessformer: A unified architecture for chess modeling. InInternational Conference on Learning Representations (ICLR), 2026
2026
-
[22]
Papke and Jeffrey M
Leslie E. Papke and Jeffrey M. Wooldridge. Econometric methods for fractional response variables with an application to 401(k) plan participation rates.Journal of Applied Econometrics, 11(6):619–632, 1996. doi: 10.1002/(SICI)1099-1255(199611)11:6<619::AID-JAE418>3.0.CO; 2-1
1996 doi
-
[23]
Regan and Guy McC
Kenneth W. Regan and Guy McC. Haworth. Intrinsic chess ratings. InProceedings of the 25th AAAI Conference on Artificial Intelligence, pages 834–839, 2011. 26
2011
-
[24]
Mariano Sigman, Pablo Etchemendy, Diego Fernández Slezak, and Guillermo A. Cecchi. Response time distributions in rapid chess: A large-scale decision making experiment.Frontiers in Neuroscience, 4:60, 2010. doi: 10.3389/fnins.2010.00060
2010 arXiv
-
[25]
The Sonas rating formula — better than Elo? ChessBase News,https://en
Jeff Sonas. The Sonas rating formula — better than Elo? ChessBase News,https://en. chessbase.com/post/the-sonas-rating-formula-better-than-elo , 2002. published 22 October 2002; regression on 266,000 games, 1994–2001; accessed 2026-07-16
2002
-
[26]
Spearman
C. Spearman. The proof and measurement of association between two things.The American Journal of Psychology, 15(1):72–101, 1904. doi: 10.2307/1412159
1904 doi
-
[27]
Life cycle patterns of cognitive performance over the long run.Proceedings of the National Academy of Sciences, 117(44): 27255–27261, 2020
Anthony Strittmatter, Uwe Sunde, and Dainis Zegners. Life cycle patterns of cognitive performance over the long run.Proceedings of the National Academy of Sciences, 117(44): 27255–27261, 2020. doi: 10.1073/pnas.2006653117
2020 doi
-
[28]
Maia-2: A unified model for human-AI alignment in chess
Zhenwei Tang, Difan Jiao, Reid McIlroy-Young, Jon Kleinberg, Siddhartha Sen, and Ashton Anderson. Maia-2: A unified model for human-AI alignment in chess. InAdvances in Neural Information Processing Systems (NeurIPS), 2024. arXiv:2409.20553v2
2024 arXiv
-
[29]
WDL_model: Fitting a Win–Draw–Loss model from played games
The Stockfish Team. WDL_model: Fitting a Win–Draw–Loss model from played games. https://github.com/official-stockfish/WDL_model, 2023. accessed 2026-07-14
2023
-
[30]
Removing skill bias from gaming statistics.arXiv preprint, 2018
I-Sheng Yang. Removing skill bias from gaming statistics.arXiv preprint, 2018. arXiv:1803.05484
2018 arXiv
-
[31]
Human- aligned chess with a bit of search
Yiming Zhang, Athul Paul Jacob, Vivian Lai, Daniel Fried, and Daphne Ippolito. Human- aligned chess with a bit of search. InInternational Conference on Learning Representations (ICLR), 2025. arXiv:2410.03893. 27
2025 arXiv
-
[2016]
doi: 10.1016/j.jpubeco.2016.05.001
2016 doi
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.