{"id":"af323da5-2d6c-4436-a37a-98d6427b8099","arxiv_id":"2607.25655","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Engine-equal chess positions carry small, reproducible human outcome skews that replicate across disjoint player groups, calendar halves, rating bands, and an out-of-sample month.","lead":"Chess positions that a strong computer engine rates as perfectly even are not even for humans: games from those positions show small, stable side biases that repeat even when players are split into two separate groups. The paper shows a machine's 'equal' does not capture everything about how people actually play, with implications for engine-based ratings and difficulty estimates.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Residual sub-family selection (the §5.3 confound) is the load-bearing soft spot: all replications preserve who chooses the line, so a history-based test on unprepared players is needed to distinguish position-level skew from repertoire selection.","rationale":"The reader's weakest assumption and this pass converge on the same point: within-family sub-repertoire selection is the largest unresolved threat to the most interesting reading of the result. I agree with the reader, however, that the paper's explicit estimand — the naturally-reached position — already contains this selection, and the abstract and §5.3 repeatedly state that the result is observational and that causation is deferred to a pre-registered randomized companion study. The statistical support for the narrower claim is strong: the primary slope of 0.691, the external-month 0.904, the account-absent June restriction (1.01), the calibrated permutation null, the line-cluster collapse, and the covariate-matched placebo all point to a real, stable, rating-adjusted position-level outcome skew in naturally-reached play. The sub-family confound does not invalidate that claim as stated; it bounds how the result should be interpreted. The concrete test I propose would sharpen this boundary: by isolating players with no prior exposure to the line, it directly targets the leading mechanism inside the residual confound. If the slope survives in unprepared players, the 'position as decision problem' reading is materially strengthened. If it does not, the paper's own scoping already provides the correct fallback. For these reasons the reader's ACCEPT verdict remains appropriate, with confidence appropriately moderate rather than high.","tokens_in":24239,"tokens_out":12980,"duration_ms":165942,"concrete_test":"Re-run the primary account-disjoint replication on a pre-specified subset of occurrences in which neither player had previously reached the same position — or any position in the same ECO sub-line (ECO code plus the last four plies) — in the Lichess history preceding October 2025. If the within-family replication slope on this 'unprepared' subset remains materially positive (e.g., lower 95% CI > 0.3, permutation p < 0.01), the skew is not primarily driven by line-specific preparation or repertoire comfort, and the concern is substantially discharged. If the slope collapses toward zero, the replicated skew is better described as a property of who chooses the line, and the 'position-level' interpretation in §5 should be weakened to 'player-position system'. Report the fraction of occurrences retained and, if survival is low, a sensitivity analysis using 'no prior occurrence of the exact FE","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's own §5.3 states the central unresolved confound: within an ECO code there can be narrower sub-repertoires with systematically different player pools, so a positive within-ECO replication slope can still reflect selection on who plays a line rather than difficulty of the position itself. Every axis tested — account-disjoint, temporal, rating-band, and external-month — preserves the same natural selection process into positions. Line-specific preparation, transposition history, and repertoire comfort are each compatible with the full pattern of results, including the out-of-sample month and the absent-account June restriction (those players may still have prepared the line elsewhere). The paper explicitly scopes its estimand to the 'naturally-reached position', and the causal question is deferred to a randomized study. Within that scope the claim is internally consistent and the replications support it. The risk is interpretive rather than statistical: if the within-family slope is largely produced by sub-family player pools, the discussion's stronger framing — that the skew attaches to 'specific decision problems rather than the openings players choose' (§5) — is not supported. The within-account demeaning and covariate-means re-fit bound only the accounted-for channels; the residual signal could still be selection. This is precisely the load-bearing assumption the reader identified, and it remains open in the paper.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies 1,661 opening positions from Lichess October 2025 games that Stockfish 18 evaluates as equal (|eval| ≤ 10cp at depth 28, depth-stable) and that humans reach at least 1,000 times. For each position it computes a skew δ — the mean deviation of game results from a rating-calibration fractional-logit expectation — and tests whether δ measured in one random half of player accounts predicts δ in the disjoint other half after removing ECO-family means. The within-family weighted replication slope is β̂ = 0.691 (family-clustered CI [0.646, 0.736]; permutation p = 0.001), with similar results on temporal, rating-band, and external-month (June 2026, β = 0.904) axes. Robustness checks include a covariate-matched placebo (0.046), a covariate-means re-fit (0.599), within-account demeaning (0.526), line-cluster collapse (0.681), and time-control stratification. A secondary clock analysis finds the disfavoured side spends more think time. The paper explicitly scopes the estimand to the 'naturally-reached position' and disclaims causation, deferring to a pre-registered randomised study.","tokens_in":24486,"tokens_out":9880,"duration_ms":107146,"significance":"If correct, the paper gives a large-scale, carefully identified demonstration that a scalar engine evaluation near zero is not a sufficient statistic for human outcomes at the level of individual positions, and that the residual is reproducible rather than noise. The paper's defensive design is a major strength: the panel and estimator are fixed before the external month is read; the permutation null is calibrated on synthetic data; the covariate-matched placebo bounds calibration-misspecification; secondary analyses are labelled as such; and the code/data release supports independent verification. The main interpretive caveat — sub-family player selection — is acknowledged in §5.3 and is compatible with the scoped claim, though it limits causal or context-free readings. The paper does not rely on machine-checked proofs, but its reproducible code and archived dataset are strong assets.","major_comments":[],"minor_comments":[{"comment":"The Discussion's sentence 'Most of the skew’s variance lies within opening family, so it attaches to specific decision problems rather than to the openings players choose' overstates what the design can establish. §5.3 correctly acknowledges that sub-repertoire selection (line-specific preparation, repertoire comfort, transposition history) can generate a positive within-ECO replication slope. Every replication axis preserves the natural selection process into positions, so the data cannot separate position-level difficulty from line-selection. Recommend rephrasing to 'finer-grained than the ECO family' or explicitly attaching the claim to the naturally-reached position as defined in §5.3.","section":"§5, §5.3"},{"comment":"The caption states the fit sample has 1,630 positions (1,661 minus 31 single-member-family positions), but panel (a) labels the plotted sample as n = 1,624. Please reconcile the count or explain the additional six excluded positions.","section":"Figure 1"},{"comment":"The 'fail(band)' verdict for the fast-moving-rating filter is interpreted as mechanical attenuation because the ICC drops to 0.41/0.32 on the filtered sample. This is plausible, but the paper should state explicitly that the ICC is recomputed on the filtered data, so the attenuation explanation is not an independent check. A brief sentence noting this would prevent over-reading.","section":"Table 2 / §4.1 (fast-moving-rating filter)"}],"recommendation":"minor_revision","confidential_remarks":"The paper is unusually careful and the statistical case for the scoped existence claim is strong. The only substantive issue is the Discussion's slightly stronger framing, which the authors themselves qualify in §5.3; this is a local wording fix. I recommend minor revision rather than accept only because of that framing and the Figure 1 count discrepancy."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Chris — this one deserves a real referee. The headline result is that among opening positions Stockfish 18 calls equal (|eval|≤10cp, depth-stable), human games on Lichess show per-position outcome skews that replicate across disjoint player accounts at slope 0.69 (permutation p=0.001), across time, across rating bands, and on a June 2026 month the panel was fixed before reading. The skews are small — median about two percentage points of White score — but they are position-specific, some favour White, some Black, and they reproduce. I came in skeptical and the design convinced me. They calibrate expected score on one split cell and apply it to the other, so position identity can't be a calibration covariate; the permutation null is calibrated on synthetic data; the placebo test (covariate-matched pseudo-positions) gives 0.046 against 0.691; the covariate-means re-fit gives an upper bound of ~13% on the composition channel; the line-cluster collapse leaves the slope at 0.681. They even report the checks that fail the bandwidth — fast-moving-rating filter, one-position-per-game, within-time-control — and interpret them as reliability effects rather than disappearing acts. That is honest, reproducible work.\n\nThe soft spot is the one they name: sub-family selection. Within an ECO code, narrower sub-repertoires can have systematically different player pools, so a positive within-family replication slope can still reflect who picks the line rather than something about the position as a decision problem. Every replication axis preserves the natural selection process into positions. The paper scopes the estimand to the 'naturally-reached position', which is defensible, but the Discussion's stronger framing — that the skew attaches to 'specific decision problems rather than the openings players choose' — goes beyond what the design can support. The think-time result has the same issue: the favoured side moves faster, which smells like home preparation. The paper acknowledges this too. So the existence claim holds for the observed ecology, not for the position in isolation.\n\nThis is for anyone working on human-AI divergence, chess analytics, or rating systems. It deserves a serious referee: the archived code and per-position estimates mean the result is checkable, and the claim is important enough that referee time is warranted. The pre-registered randomised companion study is the right next step. My recommendation: send to peer review, and push the authors to make the scope line — naturally-reached position — stick in the final version.","headline":"Engine-equal positions carry small reproducible human outcome skews; sub-family selection is the honest soft spot, but this deserves serious refereeing.","tokens_in":25001,"tokens_out":2978,"would_cite":true,"duration_ms":32491,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Among chess positions a strong engine calls dead even, human results are not balanced: positions carry small, stable side skews that reproduce across disjoint player groups, so the engine's evaluation is not a sufficient statistic for human","keywords":["chess openings","engine evaluation","human performance","outcome skew","replication","Lichess","Stockfish","decision-making"],"falsifier":"A randomized assigned-play trial, as the paper's companion study plans, that assigns players to both sides of engine-equal positions and finds no systematic per-position outcome difference by side would falsify the causal reading of the skew; alternatively, a re-analysis on a different platform's games (e.g., chess.com) that yields a within-family replication slope near zero would falsify the claim that the skew is a stable property of human play from these positions.","tokens_in":24035,"feed_emoji":"♟️","tokens_out":3792,"duration_ms":39599,"temperature":0.7,"pith_summary":"This paper claims that when Stockfish 18 evaluates an opening position as essentially equal, human games from that position are still systematically unbalanced in a way the engine's number cannot see. The imbalance is small per position—about two percentage points of White score—but it is stable: a position that tilts toward White in one random half of player accounts tilts the same way in the other half, with a replication slope of 0.69, rising to 0.94 on the most-played positions. The same per-position pattern reproduces across time splits, rating bands, and an out-of-sample month eight months later. The authors conclude that an engine evaluation, however deep, is not a sufficient statistic for human outcomes, and that the skew lives mostly inside opening families rather than being a property of whole openings. They frame the result as observational, leaving causation to a randomized companion study.","feed_headline":"Engine-equal chess positions are not equal for humans","feed_subtitle":"Per-position skews from 16M games predict a disjoint player set's results, and repeat eight months later.","key_machinery":"The central mechanism is the within-family replication slope: each position's outcome skew (the mean deviation of its games' results from a rating-calibrated expected score, in White's point of view) is measured once in each of two disjoint account groups, then the two measurements are regressed on each other after subtracting each opening-family's mean skew. Family demeaning removes opening-level repertoire selection, so a positive slope means positions within the same opening carry their own replicated tilt. The slope is estimated by weighted least squares with effective sample sizes that discount repeated games by the same account, and significance is assessed by permutation tests that sh","core_discovery":"The central discovery is a measurable, reproducible outcome skew at engine-equal positions: for 1,661 opening positions rated within 10 centipawns of zero by Stockfish 18 at high depth and reached at least 1,000 times by Lichess players in October 2025, the gap between actual results and rating-predicted results varies by position and persists when measurements are split across disjoint player accounts. In the primary split, a position's skew measured in one account group predicts the skew measured in a completely disjoint group at slope 0.69 (95% CI [0.65, 0.74], permutation p = 0.001); the disfavoured side also spends longer thinking. The authors emphasize that existence of the skew is the","pith_inferences":["If the skew is partly caused by preparation or repertoire selection, then other games with machine ground truth—such as Go, poker, or AI-assisted education—may carry similar human-specific imbalances that scalar evaluations miss.","The monotonic rise of the replication slope with position popularity suggests that rare positions may have unmeasured skews; larger corpora or adaptive sampling could map the full landscape.","The within-family residual selection confound could be tested by comparing skews of transpositions that reach the same position via different move orders; if the skew tracks the move order, the property is not the position alone.","Pairing think-time data with move-quality measures (e.g., error rates) could separate 'harder' from 'less studied' and sharpen the causal question for the randomized companion study."],"forward_implications":["Engine evaluation alone is insufficient to predict human results at even positions; practical difficulty is position-specific and reproducible.","A replicated 0.05 skew corresponds to about a 50-point rating gap, and the largest replicated skews to 150 points or more, giving players a familiar currency for the imbalance.","Most of the skew variance lies within opening families, so adjusting for opening choice alone cannot remove the imbalance.","The think-time asymmetry gives a behavioral handle: the disfavoured side reliably spends more clock, which could inform training tools and interface design.","Out-of-sample replication eight months later implies the effect is not a one-month artifact or a quirk of a single player population."],"fun_headline_variants":["Engine-equal chess spots still favor one side for humans","Chess: engine says equal, but human results skew by position","Reproducible skew in human outcomes at engine-tied openings","Stockfish says equal, humans don't: per-position skew persists","Engine-neutral chess positions aren't neutral for human players"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The result stands only if the replicated within-family skew is not substantially caused by which players choose to play a line—narrower sub-repertoires with different player pools within an ECO code—rather than by the position itself; the paper explicitly scopes its estimand to the naturally-reached position and defers causation to a randomized companion study.","fun_headline_variants_meta":{"raw":{"variants":["Engine-equal chess spots still favor one side for humans","Chess: engine says equal, but human results skew by position","Reproducible skew in human outcomes at engine-tied openings","Stockfish says equal, humans don't: per-position skew persists","Engine-neutral chess positions aren't neutral for human players"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1233,"prompt_tokens":896,"completion_tokens":337,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":640,"completion_tokens_details":{"reasoning_tokens":251}},"tokens_in":640,"tokens_out":337,"duration_ms":4745,"temperature":1.0,"reasoning_tokens":251,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T01:49:06.729908+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A randomized assigned-play trial, as the paper's companion study plans, that assigns players to both sides of engine-equal positions and finds no systematic per-position outcome difference by side would falsify the causal reading of the skew; alternatively, a re-analysis on a different platform's games (e.g., chess.com) that yields a within-family replication slope near zero would falsify the claim that the skew is a stable property of human play from these positions.","supporting_citations":[],"review_version":1}