{"id":"8783f051-a907-41e3-a88c-7af47d62f2af","arxiv_id":"2608.03219","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"LLM benchmark gains often come from producing already-reachable answers more reliably, not from making new answers reachable under a matched probe.","lead":"This paper distinguishes 'realized' answers (what a model produces by default) from 'reachable' answers (what a fixed sampling probe can find), and audits LLM benchmark gains along both axes. Across routing experiments, causal MLP interventions, and RLVR training comparisons, the gains mostly change realization rather than expanding reachability, so benchmark numbers alone do not establish capability growth.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fixed-probe reachability may misread RLVR output-format or diversity collapse as a ceiling contraction, undercutting the DAPO example.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing concern: reachability is operationalized by one fixed probe, and a flat or falling O_K may be a probe artifact. This is not a manufactured worry. The paper's own Setup admits that finite-probe non-finding does not mean impossibility, and the routing section excludes three format-collapse cells, showing the authors know probe sensitivity can corrupt comparisons. The RLVR section applies the same fixed probe without a corresponding format-collapse control. Table 6's K-doubling is a partial control, but if RLVR collapses candidate diversity or format, sampling more candidates under the same probe will not recover lost reachability. The central normative claim—'claims of capability expansion should report both realized performance and reachability'—is reasonable as a reporting standard, but the empirical evidence for the starkest case (DAPO: +14.7 deployed, -13.3 reachable) is not yet secure. My read does not change the reader's CONDITIONAL verdict; it reinforces the condition. I therefore keep the verdict unchanged and agree that this is the weakest assumption.","tokens_in":10313,"tokens_out":3863,"duration_ms":41578,"concrete_test":"On the same 300 DART questions used for Table 5, evaluate base Qwen2.5-32B and DAPO-32B with three independent reachability probes: (1) current temp 0.8 top-p default; (2) temp 1.2/top-p 0.95; (3) answer-extraction-normalized grading with multiple prompt variants, including prompts without explicit format instructions. Each probe uses K=200 and records O_K. Also compute candidate-output diversity (e.g., distinct n-gram or format-type counts) for both checkpoints. If DAPO's O_K is below base under all three probes, the ceiling contraction is robust; if any probe reverses or removes the gap, the 'reachable ceiling falls' conclusion is a probe artifact and the central claim needs qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that RLVR improves realization without expanding reachability depends on O_K being a faithful ceiling. The paper itself restricts reachability: 'an answer not found by a finite probe is unobserved under that protocol, not impossible for the model' (Setup). Yet Tables 5-6 and Figure 6b treat a fall in O_K as a fall in the reachable set. RLVR is known to narrow output diversity and alter answer format; Table 8 shows format accounts for 41-50% of the realization gain, and Figure 2a had to exclude three format-collapse cells elsewhere. If DAPO's distribution collapsed under the fixed temperature-0.8 probe, K=200/400 candidates may mostly share one wrong format, so the -13.3 point delta reflects probe sensitivity, not loss of solvability. The matched-format control is not enough if format collapse is induced by training rather than controlled by the prompt. Thus the flagship 'deployed +14.7, ceiling -13.3' may be an artifact, and the universal reporting prescription is not yet empirically grounded.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a question-level audit framework that distinguishes 'realized' answers (produced by the default deployment procedure) from 'reachable' answers (found by a fixed probe within a budget K). It then applies this framework to three settings: inference-time layer routing, MLP-block silencing, and RLVR training. The routing experiments report that random layer paths match or exceed structured oracle searches in 43 'clean' model–task cells, while answer-blind selectors recover little of the oracle gain. The MLP experiments localize a predefined recognition-correct/generation-wrong failure to a single MLP block in six cases, with component, specificity, and reciprocal controls. The RLVR experiments report that deployed accuracy rises while oracle reachability stays flat or falls, with DAPO as the sharpest case (+14.7 deployed points, -13.3 reachability points), and that newly realized questions were already base-reachable. The paper concludes that benchmark gains do not by themselves establish capability expansion, and that both realized performance and reachability should be reported.","tokens_in":10616,"tokens_out":6146,"duration_ms":73610,"significance":"If the empirical claims hold, the distinction between realization and reachability is a useful evaluation discipline and could change how benchmark gains are interpreted. The paper has real strengths: matched budgets and temperatures, random baselines for routing, predefined failure sets for MLP localization, reciprocal perturbation controls, and public code. The central caveat is that reachability is defined by a fixed finite probe, and several headline conclusions depend on treating that probe's success count as a faithful 'reachable ceiling' across checkpoints whose output distributions differ. This is an empirical audit rather than a derivation, and the main claims need to be robust to probe sensitivity before the reporting prescription is fully grounded.","major_comments":[{"comment":"The Setup explicitly says 'an answer not found by a finite probe is unobserved under that protocol, not impossible for the model,' yet the RLVR conclusion treats a drop in O_K as a drop in the reachable ceiling. For DAPO, Table 5 reports -13.3 points (40 fewer reachable questions), and Table 6 shows this persists at K=400. This is only valid if the fixed probe (temperature 0.8, base prompt, K<=400) has comparable sensitivity for base and trained checkpoints. RLVR is known to alter output format and diversity; Table 8 attributes 41-50% of the realization gain to prompt format, and Figure 2a elsewhere excludes 'format-collapse' cells. If DAPO's trained distribution collapses under the base-format probe, the -13.3 point drop may be a probe-sensitivity artifact rather than a contraction of solvable questions. Please provide candidate-output diagnostics (format, length, diversity) for the DAP","section":"Setup; RLVR and Reachability (Fig. 6b, Tables 5-6)"},{"comment":"The Abstract claims random routes 'match or exceed structured search in all 43 model and task settings,' but Figure 2a reports '43 clean model–task cells; three format-collapse cells are excluded.' This is a post hoc exclusion, not 'all settings.' The paper never states the exclusion criterion or reports what happens in the excluded cells. Because the routing result is one of the paper's four contributions, the universal phrasing is misleading unless the exclusion rule is defined before the comparison and the excluded cells are shown. Please report the three excluded cells individually and restate the claim as 'all non-collapsed cells' or provide a pre-specified definition of 'clean.'","section":"Figure 2a and Abstract"},{"comment":"Most headline tables and figures report point estimates without confidence intervals, and some of the supporting differences are small relative to the sample size. For example, Table 5's Qwen3.5 row has n=150 and a reachable-count change of -1; Table 6's SimpleRL row has n=150 and +1; these are within sampling noise. Yet the aggregate claim 'deployed performance rises while the reachable ceiling remains flat or falls' in 'five of six' matched evaluations depends on distinguishing zero from small nonzero changes. Please report paired bootstrap or McNemar confidence intervals for D, O_K, and their difference, and include the per-question transition matrices behind Figure 6a. This would also clarify whether the DAPO result is the only statistically secure divergence or whether the other cells support it.","section":"Tables 2, 5, 6, 7; Figure 6"}],"minor_comments":[{"comment":"Table 1 mixes different budgets across rows, and the caption says the matched comparison is 'given in Figures 1 and 2.' For readability, add a budget column or use the same K across rows in the table; otherwise the row-to-row comparison is hard to interpret.","section":"Table 1"},{"comment":"The 'recognition' auxiliary readout is defined as 'ranks the correct option under multiple-choice likelihood,' but the paper also uses tasks such as GSM8K and DART-Math that are not multiple-choice. State how recognition is operationalized for free-form generation tasks, or clarify that the MLP-localization census is restricted to multiple-choice tasks.","section":"Setup (Recognition)"},{"comment":"The phrase 'clean model–task cells' appears without definition in the main text. Specify in the caption or Setup which cells are clean and why the three format-collapse cells are excluded.","section":"Figure 2a caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is a serious empirical audit with a potentially useful distinction, but the strongest illustrative result (DAPO) is vulnerable to the fixed-probe sensitivity concern. The routing 'all 43 settings' claim also needs to be qualified because of the post hoc exclusions. These are fixable with additional controls and reanalysis, so I would not reject, but the current version is not yet accept-ready. The authors may also want to soften the abstract's universal phrasing until the probe-robustness analysis is in place."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The realized/reachable split is worth taking seriously. The paper defines it cleanly under fixed budgets, temperatures, and formats, then actually follows question-level fates across routing, MLP silencing, and RLVR checkpoints. That protocol is the contribution, and it should become a reporting habit for anyone claiming a score gain is a capability gain.\n\nWhat it does well: the routing audit uses matched random paths as a baseline, which is the right control, and the finding that random routes match structured search across most clean cells is a solid negative result. The MLP localization is careful for what it is: predefined failure set, same-layer attention control, healthy-panel breakage, and reciprocal perturbation. Small samples, but the controls are real. The RLVR comparison is also honest about matched checkpoints and includes a propensity analysis showing gains concentrate on already-sampled questions. Code is out. Credit where earned.\n\nWhere it gets soft. The flagship DAPO numbers (deployed +14.7, reachable ceiling −13.3) depend on the fixed probe staying faithful after training. RLVR is known to collapse output diversity and shift formats, and the probe draws K candidates at temperature 0.8 under a fixed prompt. If the trained model's samples pile up in one wrong format, the oracle looks lower even when the answer is still possible. The paper acknowledges that reachability is protocol-relative, and the budget control at K=400 helps, but it doesn't vary temperature or probe diversity. Table 8 shows format explains 41–50% of the realization gain, so the rest is genuine, yet the ceiling read is still one probe away from solid. Second, \"all 43 settings\" is accurate only after three format-collapse cells are excluded; that's disclosed, but it weakens the universal phrasing in the abstract. Third, most tables lack error bars, and several headline cells rest on n=150 or 300, which matters when the deltas are a few points.\n\nNone of this kills the framework. The distinction is real, the matched-audit idea is portable, and the routing and localization parts hold up better than the RLVR-ceiling part. The paper is for evaluation researchers and people working on post-training interpretability. It deserves a serious referee, though the referee should push for probe-robustness checks before the DAPO contraction is reported as fact. I'd bring it to a reading group precisely because the distinction is worth arguing about.","headline":"A real and useful evaluation distinction, with a fixed-probe ceiling that probably overstates how little RLVR expands reachability.","tokens_in":11046,"tokens_out":2531,"would_cite":true,"duration_ms":31642,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A benchmark gain is not evidence of expanded capability unless the reachable set also grows; in DAPO/DART, deployed accuracy rose 14.7 points while oracle reachability fell 13.3 points.","keywords":["LLM evaluation","reachability","realization","benchmark gains","RLVR","layer routing","MLP localization","question-level audit"],"falsifier":"Run the DAPO/DART matched comparison (base vs trained, n=300, K=200, temperature 0.8) again with K=800 and a second temperature (e.g. 1.0) and an alternative prompt format. If the trained model's $O_K$ recovers to at least the base level under any of those probes, the 13.3-point reachability drop fails to generalize and the 'reachable ceiling' is probe-dependent. Also check whether the newly realized questions are exactly the base model's high-hit-rate bin (above 35% at K=200); a different propensity profile would falsify the concentration claim.","tokens_in":10258,"feed_emoji":"📊","tokens_out":9272,"duration_ms":84430,"temperature":0.7,"pith_summary":"This paper argues that a rising benchmark score cannot be read as a growing capability unless we know whether the model is solving new questions or merely producing answers it could already reach. To separate these, it defines a question as realized when the default deployment produces the correct answer, and reachable when a fixed sampling probe finds the correct answer within a specified budget. Auditing layer routing, MLP-block interventions, and six matched RLVR checkpoints, the paper finds that random layer routes match or beat structured search, that silencing one identified MLP block repairs most failures where the model recognises but cannot produce the correct answer, and that RLVR training raises deployed performance while the reachable ceiling stays flat or falls—in the DAPO case, +14.7 points deployed and −13.3 points reachable. The conclusion is that realization and reachability move apart, so capability claims should report both under matched conditions.","feed_headline":"RLVR gains often come from answers the base model could already reach","feed_subtitle":"Matched audits show deployed accuracy rising while the reachable ceiling stays flat or falls—capability claims should report both.","key_machinery":"The load-bearing machinery is the two-number question-level audit $(d_i, o_i(K))$. $d_i$ flags whether the deployment output is correct; $o_i(K)$ flags whether any of $K$ probe candidates is correct. Their aggregates $D$ and $O_K$ are called deployed performance and oracle reachability (or the reachable ceiling). The same pair is recorded for every question on both sides of a matched comparison, so a training run or inference intervention can be classified as changing realization, reachability, both, or neither. The audit is only defined relative to a stated protocol—question set, answer format, grading rule, temperature, probe, and budget—and the oracle label means correctness is used after","core_discovery":"The paper's central claim is that a benchmark gain is underdetermined: the same aggregate score change can come from expanding the set of questions a model can reach, or from making already-reachable answers appear more reliably. It makes the distinction operational by scoring every question twice: $d_i$ for whether the default deployment produces the correct answer, and $o_i(K)$ for whether any of $K$ candidates from a fixed probe is correct. Aggregated, these are deployed performance $D$ and oracle reachability $O_K$, and the comparison is only meaningful when base and trained checkpoints share question set, answer format, temperature, budget, and grader. The empirical payload is threefold","pith_inferences":["Editorial extension: the finite-probe caveat opens a direct artifact check. Re-probe DAPO/DART with higher $K$, a different temperature, and a different prompt format; if the trained checkpoint's $O_K$ stops falling, the measured reachability contraction is an artifact of the probe's sampling distribution rather than a loss of solvable questions.","Editorial extension: the routing result suggests the limiting factor for test-time scaling is candidate selection, not candidate generation; ordinary sampling already reproduces the gap without layer routing, so any new answer-blind selector should be compared against majority voting at equal budget.","Editorial extension: the predefined failure set (recognizes correctly, generates wrong) can serve as a public benchmark for mechanistic interventions—if a proposed edit repairs a large fraction of that panel without breaking healthy questions, it has a concrete measure to beat.","Editorial extension: benchmark designers could publish $D$ and $O_K$ as a standard two-number scorecard, making 'capability expansion' claims falsifiable and comparable across labs."],"forward_implications":["A rising deployed score with a flat or falling $O_K$ is a realization gain, not a reachability expansion; capability claims should report both $D$ and $O_K$ under matched conditions.","RLVR tends to realize questions the base model already sampled frequently, so training gains can be concentration effects rather than additions to the solvable set.","Structured layer routing should not be credited with expanding what a model can answer unless it beats budget-matched random routes under the same scoring; answer-blind selection recovers almost none of the oracle headroom.","Recognition-correct/generation-wrong failures can be causally localized to a single MLP block in some models, giving an intervention-level handle on where realized answers are lost.","Whether RLVR raises or lowers the reachable ceiling depends on the recipe—math-only RLVR lowered oracle reachability in one OLMo lineage while SFT-containing variants raised it—so training-recipe controls are part of any reachability claim."],"supporting_citations":[{"why":"Defines pass@k, the candidate-set success notion that o_i(K) extends to question-level reachability.","marker":"(Chen et al. 2021)"},{"why":"Supplies the random-search control that budget-matched random routes use to test whether structured routing structure matters.","marker":"(Li and Talwalkar 2020)"},{"why":"Provides the SimpleRL-Zoo base and trained checkpoints used in the RLVR matched evaluations.","marker":"(Zeng et al. 2025)"},{"why":"Provides the DAPO checkpoints and training system for the +14.7/−13.3 matched comparison.","marker":"(Yu et al. 2026)"},{"why":"Prior evidence that low-budget gains can coexist with high-budget contraction; the paper's matched question-level audit extends this.","marker":"(Yue et al. 2025)"},{"why":"Establishes that transformer MLP blocks act as key-value memories, motivating the MLP-block silencing experiments.","marker":"(Geva et al. 2021)"},{"why":"Supplies the causal localization approach the paper adapts for component, reciprocal, and specificity controls.","marker":"(Meng et al. 2022)"},{"why":"Claims RLVR implicitly incentivizes correct reasoning in base LLMs; the paper's reachability ceiling results qualify that claim.","marker":"(Wen et al. 2025)"},{"why":"Debates whether RLVR expands or shrinks reasoning capacity; the paper's D/O distinction offers a sharper audit.","marker":"(Yao et al. 2025)"},{"why":"Diagnoses sampling blind spots in math reasoning difficulty estimation, supporting the paper's finite-probe caveat.","marker":"(Zhou et al. 2026)"}],"fun_headline_variants":["Benchmark gains may just surface already-reachable answers","Reachability vs realization: a stricter lens on LLM gains","LLM gains often mean better access, not more capability","Deployed scores rise while reachable ceiling stays flat or falls","New audit separates reachable answers from realized ones"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that $O_K$—sampling $K$ candidates at temperature 0.8 under a fixed format and grader—is the right operational definition of reachability; the paper itself concedes that an answer not found by a finite probe is unobserved, not impossible, so a flat or falling ceiling after RLVR could be a probe artifact if training shifts the sampling distribution (format collapse, output length changes, or reduced diversity).","fun_headline_variants_meta":{"raw":{"variants":["Benchmark gains may just surface already-reachable answers","Reachability vs realization: a stricter lens on LLM gains","LLM gains often mean better access, not more capability","Deployed scores rise while reachable ceiling stays flat or falls","New audit separates reachable answers from realized ones"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000175,"raw_usage":{"total_tokens":1162,"prompt_tokens":820,"completion_tokens":342,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":261}},"tokens_in":564,"tokens_out":342,"duration_ms":4551,"temperature":1.0,"reasoning_tokens":261,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:14:20.987981+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the DAPO/DART matched comparison (base vs trained, n=300, K=200, temperature 0.8) again with K=800 and a second temperature (e.g. 1.0) and an alternative prompt format. If the trained model's $O_K$ recovers to at least the base level under any of those probes, the 13.3-point reachability drop fails to generalize and the 'reachable ceiling' is probe-dependent. Also check whether the newly realized questions are exactly the base model's high-hit-rate bin (above 35% at K=200); a different propensity profile would falsify the concentration claim.","supporting_citations":[],"review_version":1}