{"id":"2c4e4508-2c10-4d58-9613-ee71f051dd35","arxiv_id":"2607.20952","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In a chess latent-reasoning model, replacing or removing the silent thought vectors barely changes moves, so the RL improvement appears to be encoded in the weights, not in a consulted scratchpad.","lead":"A chess-playing language model trained with silent 'latent thoughts' plus reinforcement learning makes more legal moves, but causal tests show the thoughts are almost board-invariant and can be replaced or removed without changing play. This suggests the RL gain lives in the weights, not in an inference-time scratchpad, challenging a common assumption about latent reasoning.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim rests on the exact-zero condition, which the paper itself labels an out-of-distribution input; the only significant before/after gap (1% vs 9%) may reflect differential OOD sensitivity rather than reduced thought reliance.","rationale":"The paper is genuinely careful: it reports both the retention-ratio and raw-flip framings in A.7, fixes the evaluation harness, includes six structurally different interventions, and calibrates J-lens against known content before trusting its null result. The content-invariance finding — Substitute and Noise cost essentially nothing on both checkpoints — is a solid negative result about thought-content reliance. My concern is narrower but load-bearing: the between-checkpoint robustness claim, which is what supports 'RL reshapes the weights rather than improving inference-time computation,' has only one statistically significant condition, and that condition is explicitly OOD. The raw-flip analysis in the paper's own appendix shows the opposite direction, so the evidence for the headline mechanistic claim depends on a normalization choice applied to a near-collapse endpoint. The n=1000 replication does not fix this because it only tests Rung-3 within one checkpoint, not the Stage-2 vs Rung-3 comparison. This is the same weakest assumption the reader identified. Since the reader already assigned CONDITIONAL, and the proposed graded-α test would resolve whether the effect is real or an OOD artifact, the appropriate verdict is unchanged: conditional acceptance pending that check.","tokens_in":27632,"tokens_out":4574,"duration_ms":49622,"concrete_test":"Run the full six-condition battery on both Stage-2 and Rung-3 at n=1000 with the zero condition expanded into a graded family α·T* for α ∈ {1, 0.5, 0.1, 0.01, 0.001, 0}, and report both retention ratios and raw legal-to-illegal flips. If the Stage-2 vs Rung-3 gap appears only at α=0, or disappears when raw flips are used, the exact-zero result is an OOD endpoint artifact and the robustness claim is not supported. If the gap grows monotonically as α→0 and persists under raw-flip analysis at small nonzero α, the concern is resolved in the paper's favor.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing evidence for 'RL adds robustness, not content reliance' is the between-checkpoint retention gap under Zero (§4.4, Table 7) — the only condition that clears significance at n=100. Yet §3.3 explicitly defines Zero as testing 'tolerance for an input shape the model has never seen, not content.' The paper therefore uses an acknowledged OOD perturbation as the sole significant support for the claim that RL changes how much the model needs its thoughts. Appendix A.7 discloses that the raw-count framing points the other way: under Zero, Stage-2 loses 47 legal boards and Rung-3 loses 49, and a paired Wilcoxon on raw flips is not significant for any severe condition. The reported +0.134 gap is a ratio of small numbers (1/48 vs 9/58) after normalizing by each checkpoint's own baseline; Rung-3's higher baseline (58% vs 48%) makes proportional retention more favorable to it. At n=100, 9 vs 1 retained boards could easily be OOD noise. The n=1000 replication (A.8) only reruns Rung-3's side within one checkpoint, so it cannot validate the before/after comparison on which the central claim rests. This does not undermine the content-invariance results (Substitute/Noise cost nothing) or the behavioral gains, but those results alone do not establish that RL's benefit lives in the weights rather than in some other training-induced change.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper trains a Qwen3-14B LoRA chess model through an SFT baseline, an explicit-reasoning RL checkpoint (Rung 1), a staged latent-thought curriculum (Stage-2), and latent GRPO (Rung-3). It reports a legality gain from 48% to 61%, a complete elimination of checkmate confabulation, and a six-condition causal intervention battery applied to the latent thought positions both before and after RL. Content-preserving substitution and noise leave performance essentially unchanged; ablation costs a little; exact-zero corruption causes collapse, with a larger proportional retention for Rung-3 than for Stage-2 (1% vs 9% retained legality). The paper interprets this as evidence that RL adds robustness to disruption rather than reliance on thought content, and further claims the RL gain is encoded in a low-rank weight change concentrated in MLP/gate projections. A J-lens analysis reports near-constant cross-board thought-vector similarity (~0.99), and a Gumbel-reparameterized RL control does not outperform the deterministic recipe.","tokens_in":1480,"tokens_out":1449,"duration_ms":52828,"significance":"If the central claim were fully established, the paper would be a valuable direct before/after test of the latent-thought-as-scratchpad assumption under RL, in a domain where RLVR over latent reasoning has previously been reported to fail. The paper has notable strengths: a fixed 100-position harness, a six-condition intervention design, explicit disclosure of the raw-count analysis that runs against the paper's preferred framing, a ten-times-larger replication for the post-RL side, a calibrated J-lens null on thought content, and an exact SVD of the LoRA delta. However, the headline mechanistic claim currently rests on one condition that the paper itself identifies as an out-of-distribution input, and the raw board-level counts move in the opposite direction for that condition. The result is therefore interesting but not yet load-bearing in its current form.","major_comments":[{"comment":"The central claim that RL adds robustness to disruption rather than reliance on thought content rests on the Zero condition, but Section 3.3 explicitly defines Zero as testing 'tolerance for an input shape the model has never seen, not content.' Appendix A.7 further discloses that under Zero, Stage-2 loses 47 legal boards and Rung-3 loses 49, a difference in the opposite direction, and a paired Wilcoxon on raw flips is not significant for any severe condition. The only significant before/after gap (+0.134 retention ratio) is therefore an artifact-sensitive comparison between two checkpoints with different baseline legal rates (48% vs 58%), using proportional normalization. Because exact-zero vectors are out-of-distribution for both checkpoints, the differential collapse could reflect different OOD sensitivity rather than reduced functional reliance on thoughts. The paper should either pr","section":"4.4; Appendix A.7; 3.3"},{"comment":"The n=1000 replication is presented as confirming the pattern, but it only reruns Rung-3's own side of the battery. It cannot validate the paired Stage-2-versus-Rung-3 retention-gap comparison on which the central claim rests. Moreover, Lenmatch-Ablate at n=1000 drops Format compliance to 83.3%, so its legal-rate gap against Baseline is partly driven by output-format failures, and the appendix itself notes that forcing attention onto pad tokens introduces its own out-of-distribution cost. Thus the large-sample replication supports the post-RL checkpoint's internal ordering but does not provide the missing before/after evidence.","section":"Appendix A.8"},{"comment":"The content-invariance results (Substitute and Noise) are convincingly null in both checkpoints and show that neither checkpoint relies on specific thought content. But this does not by itself establish that RL's benefit lives in the weights rather than in some other training-induced change, because the only significant before/after robustness difference is the OOD Zero condition. The Ablate and Lenmatch-Ablate comparisons trend in the predicted direction but are not individually significant at n=100, and the n=1000 replication does not include Stage-2. A causal patch/ablate of the dominant singular directions identified in Section 4.7 — suggested in Future work — would provide a much more direct test of the weight-localization claim and should be reported before the mechanistic conclusion is drawn at this strength.","section":"4.4; 4.7"}],"minor_comments":[{"comment":"The baseline discrepancy (61% in Table 1 vs 58% in Table 2) is explained by NF4 dequantization-order variation, but it is unusual to rely on approximate dequantization for a headline number while using a different internal baseline for the causal battery. Clarify why the same evaluation script gives two different baseline values for the same adapter.","section":"Table 2 note"},{"comment":"The Lenmatch-Ablate Format drop to 83.3% means the McNemar legal-status comparison for that condition is partially a formatting-failure comparison. It would be helpful to report the legal rate conditional on format compliance, or to use a format-robust legality measure.","section":"Appendix A.8"},{"comment":"The cross-recipe cosine similarity of +0.034 is small; the claim that it is 'a real, if modest, signal rather than noise' would be more convincing with a null distribution from randomized or permuted deltas, rather than only the argument that unrelated high-dimensional vectors concentrate near zero.","section":"4.7"}],"recommendation":"major_revision","confidential_remarks":"The paper is transparent and mechanically careful, but the central claim is currently supported by only one significant condition, and that condition is OOD by the paper's own definition. The raw-count reversal in A.7 is a serious concern for the headline interpretation. I would be willing to look at a revision that either supplies a non-OOD significant before/after comparison or reframes the contribution around the content-invariance null and the behavioral gains."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know this paper for the before/after RL causal battery. It is the first to run the same six interventions on the same model before and after RL, in chess rather than math. The behavioral findings are clean: legality climbs from 48% to 61%, confabulation goes to zero, and the paper is upfront that accuracy stays flat. The content-invariance result is solid: substituting or noising the thought vectors costs nothing at either checkpoint, and the J-lens work is a decent attempt at a calibrated null.\n\nThe soft spot is the central mechanistic claim. The 'RL adds robustness, not content-reliance' headline rests on one condition: exact-zero vectors, where Rung-3 retains 9% legality vs 1% for Stage-2. The paper itself admits Zero is an OOD perturbation, not content. The raw-count framing (board flips) points the other way and is non-significant. The n=1000 replication only reruns Rung-3's side, so it can't validate the before/after gap. At n=100 with one seed, 9 vs 1 is thin. The paper is honest about all of this — Appendix A.7 discloses the tension directly — but that means the load-bearing evidence is weaker than the abstract implies.\n\nWhat the paper does well: it reports the two framings, it gives the Bonferroni-corrected CIs, it flags the single-seed limitation, and it engages directly with Switch. The design is a good template for future work even if the conclusions are unsettled.\n\nMy take: the behavioral findings are probably right but narrow. The content-invariance claim is well supported. The robustness-gap claim is suggestive, not established. I'd send it to peer review — a good referee will make the authors strengthen the before/after comparison, ideally with more seeds and a non-OOD severity condition, or soften the central claim. I'd cite it for the battery design and the honest reporting, not for the conclusion.\n\nRecommendation: serious referee, expect heavy revision.","headline":"A genuinely new before/after-RL causal battery on latent thoughts, honest about its own soft spots, but the central robustness claim rests on a single OOD condition whose raw counts point the other way.","tokens_in":28438,"tokens_out":1867,"would_cite":true,"duration_ms":20041,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Reinforcement learning on a latent-reasoning chess model improves resilience to disrupted thoughts, not reliance on their content; the gain lives in the weights.","keywords":["latent reasoning","continuous thoughts","reinforcement learning","causal intervention","chess","weight localization","confabulation","robustness"],"falsifier":"Find a single board from the frozen 100-position harness where substituting the real thought vectors with the fixed average vector T* changes the legal move from legal to illegal (or vice versa) in the post-RL checkpoint; the paper reports zero such flips, so one clean counterexample would falsify content-invariance.","tokens_in":27499,"feed_emoji":"♟️","tokens_out":4034,"duration_ms":41814,"temperature":0.7,"pith_summary":"This paper trains a chess-playing language model through a staged latent-reasoning curriculum followed by reinforcement learning, and asks whether the latent 'thought' vectors are an actively consulted scratchpad. It reports that replacing or noising the thought content leaves legality unchanged, removing the thoughts costs only a little, and only exact-zero corruption causes collapse. The zero-collapse gap between the pre-RL and post-RL checkpoints (1% vs 9% retained legality) is the key evidence: reinforcement learning appears to add robustness to disruption, not reliance on thought content. The paper concludes that in this setting latent reasoning's main effect is to shape the weights during training, and it localizes the change to low-rank, concentrated weight adjustments. A sympathetic reader would care because this challenges the field's default assumption that silent thoughts function as an inference-time scratchpad, and it demonstrates a working RL gain in a domain where similar recipes have been reported to fail.","feed_headline":"RL's chess gain lives in weights, not silent thoughts","feed_subtitle":"A six-condition causal battery shows thought content is replaceable; only exact-zero disruption separates pre- and post-RL.","key_machinery":"The load-bearing instrument is a six-condition causal intervention suite applied identically to the same checkpoint before and after reinforcement learning, comparing each intervention's effect on legal-move rate. The latent thought vector — a hidden state fed back as the next input embedding instead of being decoded into a word — is the object intervened upon. The suite separates three questions: whether thought content matters (substitution, noise), whether the positions need to exist (ablation, length-matched ablation), and whether the model tolerates an unseen input (exact zero). A Jacobian-lens calibration and cross-board cosine similarity (~0.99) further support the claim that thoughts","core_discovery":"The paper's central claim is that in a chess-playing model trained through a staged latent-thought curriculum followed by reinforcement learning, neither the pre-RL nor the post-RL checkpoint relies on the specific content of its latent thought vectors. A six-condition causal battery — replacing thoughts with a fixed vector, random noise, removing them, length-matched removal, and exact-zeroing — shows content-preserving substitutions leave legality unchanged, removal costs little, and only exact-zero vectors cause collapse. The zero-collapse gap between checkpoints (1% pre-RL vs 9% post-RL retained legality) is the key evidence: RL does not teach the model to consult its thoughts more effec","pith_inferences":["Editorial inference: exact-zero vectors are an out-of-distribution input the model never saw in training; the 1%-vs-9% gap may reflect differential sensitivity to OOD inputs rather than a clean measure of content reliance, a concern the paper itself acknowledges in Section 3.3.","Editorial inference: the result is demonstrated on one 14B-parameter LoRA-tuned model in chess; the field-default scratchpad assumption could still hold in math/logic domains or in full fine-tuning at larger scale, and the paper's scope explicitly excludes those settings.","Editorial inference: a testable extension would run the same six-condition battery on a model trained with a reward that does not gate legality, to see whether the robustness gap is an artifact of the gated reward design rather than of RL per se.","Editorial inference: if the weight-localization result generalizes, one could attempt direct weight edits (patching singular directions) as a cheaper alternative to RL for adding robustness, provided the behavior transfers across checkpoints."],"forward_implications":["If the central claim holds, latent-reasoning pipelines that feed hidden states back as inputs do not need to preserve thought content at inference time; the same behavior should be obtainable with arbitrary filler vectors.","The RL robustness gain is carried by the weights, so downstream analyses or steering tools that read thought-vector content would be reading a scaffold, not the computation.","The gated legality reward alone eliminates checkmate confabulation, suggesting that a single gate design can remove a hallucination failure mode without a dedicated reward term.","Weight-change localization predicts that patching dominant singular directions from the post-RL delta into the pre-RL checkpoint should transfer part of the legality gain; the paper identifies layer 19's down-projection as the prime candidate.","The monotone severity ordering across conditions (substitute/noise < ablate < lenmatch-ablate < zero) provides a template for testing content-invariance in other latent-reasoning models."],"fun_headline_variants":["Chess RL gain survives thought erasure, lives in weights","Silent thoughts replaceable: RL chess model's real change in weights","Exact-zero thoughts only trigger collapse in chess RL model","Weight shaping, not scratchpad: chess RL's latent reasoning","RL makes chess thoughts disposable, not consulted"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central interpretation treats the exact-zero intervention as 'total signal loss' whose differential collapse pre/post RL proves reduced reliance on thought content; but exact-zero vectors are a never-seen input shape, so the gap could reflect different sensitivity to out-of-distribution inputs rather than different reliance.","fun_headline_variants_meta":{"raw":{"variants":["Chess RL gain survives thought erasure, lives in weights","Silent thoughts replaceable: RL chess model's real change in weights","Exact-zero thoughts only trigger collapse in chess RL model","Weight shaping, not scratchpad: chess RL's latent reasoning","RL makes chess thoughts disposable, not consulted"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000155,"raw_usage":{"total_tokens":1106,"prompt_tokens":853,"completion_tokens":253,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":597,"completion_tokens_details":{"reasoning_tokens":170}},"tokens_in":597,"tokens_out":253,"duration_ms":3696,"temperature":1.0,"reasoning_tokens":170,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T08:53:41.908607+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Find a single board from the frozen 100-position harness where substituting the real thought vectors with the fixed average vector T* changes the legal move from legal to illegal (or vice versa) in the post-RL checkpoint; the paper reports zero such flips, so one clean counterexample would falsify content-invariance.","supporting_citations":[],"review_version":1}