{"id":"c911d906-a273-4c12-806c-ead6d0700a1b","arxiv_id":"2607.28576","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":5.5,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Self-inspection methods (Self-Refine, Reflexion, Best-of-N self-verify) lose to equal-token repeated sampling on math from 1.5B to 7B; no tested method reliably wins.","lead":"On math problems with 1.5B–7B models, self-refine, forced Reflexion, and model-chosen Best-of-N never beat plain repeated sampling once token budgets match; all self-inspection comparisons land negative. The result forces test-time method papers to report a cost-matched sampling baseline instead of only beating one chain-of-thought.","discovery_kind":"replication","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The paper’s strongest claim is a carefully scoped negative empirical result, not a universal claim that self-inspection never helps. Within that scope the evidence is unusually tight: cost is measured rather than capped, comparisons are paired, multiplicity is corrected, the Best-of-N confound is removed by holding samples fixed, and the authors themselves surface the Reflexion non-firing failure mode and a scoring bug that would have biased results the other way. The reader correctly flags the main external limitation (domain and cost measure) without treating it as an internal refutation. I find no more load-bearing internal concern; the honest stress-test outcome is that the argument holds and the ACCEPT verdict should stand.","tokens_in":24642,"tokens_out":429,"duration_ms":8730,"concrete_test":"Re-run the independent re-derivation script (Appendix A.6) on the released 19,200 generations and confirm that all 96 checks still match Table 2 and the judges-vs-counting gaps in §5.3; any mismatch >0.5 pp on a significant cell would reopen the claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that none of the seven methods is reliably better than cost-matched repeated sampling, and that all 18 self-inspection comparisons are negative, with Self-Refine and forced Reflexion still 3.6–10.1 pp below at 7B—is internally well supported. The design (paired questions, measured per-method cost, bootstrap CIs, Holm correction, identical-sample judges-vs-counting, documented bug fix and independent re-derivation) matches the claim as stated. The reader’s scope caveat (exact-match math, generated tokens, fixed prompts/hyperparameters) is already disclosed in §6.4–6.5 and does not undermine the claim inside its stated bounds; it only limits extrapolation. No hidden inconsistency or load-bearing statistical flaw rises to the level of overturning the result.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper re-runs budget-matched evaluation of seven test-time reasoning methods (CoT, Plan-and-Solve, Self-Refine, Reflexion and a forced variant, Best-of-N with self-verify, multi-agent debate) against self-consistency on Qwen2.5 models (1.5B/3B, with selected methods at 7B) and two math benchmarks (GSM8K, MATH-500; 150 questions each). Cost is measured as all generated tokens; each method is compared to the self-consistency curve interpolated at its own measured cost, with paired bootstrap CIs and Holm correction. No method is reliably better than equal-cost repeated sampling; all 18 self-inspection comparisons are negative. Holding the same eight samples fixed, model selection loses to majority vote below 7B and is indistinguishable at 7B, while Self-Refine and forced Reflexion remain below baseline at 7B. Reflexion as published never retries on 1.5B. Code, prompts, generations, and verification scripts are released.","tokens_in":24855,"tokens_out":1411,"duration_ms":47931,"significance":"If the result holds inside its stated scope, it is a high-value corrective for the test-time compute literature: gains over single CoT are not evidence that planning, self-critique, or self-selection mechanisms help once token budget is controlled. Strengths that raise the contribution above a routine bake-off include (i) the identical-sample judges-vs-counting design that isolates selection from sampling, (ii) proper paired inference with multiplicity control and an explicit detectability statement, (iii) per-method realized-cost matching rather than a shared cap, (iv) reporting of Reflexion’s engagement rate (exposing silent collapse to CoT), and (v) full artifact release plus an independent re-derivation script and documented scoring/context bugs. The work is a careful statistical re-examination and extension of Wang et al., not a new phenomenon, but that is the right contribution for the claim.","major_comments":[{"comment":"Table 2 and §5: the full seven-method, cost-matched design is reported for 1.5B and 3B only. At 7B the table covers CoT, Self-Refine, forced Reflexion, and Best-of-N; Plan-and-Solve, published Reflexion, and debate are absent. The abstract and title frame the result as holding “from 1.5B to 7B” for the comparison as a whole. That framing is accurate for rewriting and for judges-vs-counting, but overstates coverage for the full method suite. Please state explicitly in the abstract and §5 which claims are supported at 7B and which stop at 3B, so the scale claim cannot be read as a complete six-setting × seven-method grid.","section":"Table 2, §5, Abstract"},{"comment":"§5.1 and §6.1: two groupings of “self-inspection” / “repeated passes over own work” are used to organize the negative pattern. The manuscript correctly labels one grouping as formed after the Best-of-N scoring fix and reports both a naive sign test and a clustered setting-level exact test (p=0.0156 / p=0.06). Because the central claim is already carried by the per-comparison table and the sample-fixed Best-of-N contrast, the post-hoc grouping should not be presented as confirmatory structure in the abstract (“all 18 self-inspection comparisons”). Keep the descriptive grouping in the body; lead the abstract with the pre-specified cost-matched tests and the fixed-sample selection result.","section":"§5.1, §6.1, Abstract"}],"minor_comments":[{"comment":"§4.1 vs §5: the setup section introduces 1.5B and 3B, while results add 7B for a subset of methods. A single sentence in §4.1 stating the 7B extension and which methods were affordable would remove the discontinuity.","section":"§4.1"},{"comment":"Figure 1 is dense (many rows). Consider splitting by model size or ordering methods consistently within each setting so significant (coloured) intervals are easier to scan.","section":"Figure 1"},{"comment":"§4.7 / input-token reconstruction: the total-token re-analysis for Best-of-N is valuable and strengthens the paper. A short pointer in the main results (not only threats) that the Best-of-N gap widens under total tokens would help readers who only skim §5.","section":"§4.7, §5.3"},{"comment":"Appendix A.4–A.5: the context-truncation bias and Best-of-N scoring bug are exemplary scientific hygiene. Consider a one-sentence forward reference in §4 or §5 so readers know the main numbers are post-correction without waiting for the appendix.","section":"Appendix A.4–A.5"},{"comment":"Typos / polish: “ap-value” spacing in the introduction; “V alidating” and “T able” broken initial capitals in §4; “MA TH-500” spacing artifacts appear in several places—likely PDF ligature/encoding issues worth a pass.","section":"§1, §4"},{"comment":"Related work: the discussion of trained verifiers (Zhang et al.) and imperfect-verifier limits (Stroebl et al.) is well placed. A brief explicit contrast with process-reward / trained self-verifiers in the conclusion would further prevent readers from over-generalizing to trained selectors.","section":"§3, §7"}],"recommendation":"minor_revision","confidential_remarks":"This is stronger open experimental work than much of the test-time-methods literature: the design matches the claim, limitations are disclosed in the right places, and the artifact trail (including bug postmortems and a second-path re-derivation) is unusually good. I would not block on the incomplete 7B grid if the abstract is tightened. Fit is appropriate for a methods/evaluation venue in CL/ML; novelty is primarily rigorous re-examination plus the fixed-sample selection result and small-model regime, which the authors already frame honestly."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple and earned. Once you match generated-token cost and do proper paired inference, none of the seven methods reliably beats repeated sampling on these math benchmarks, and every self-inspection comparison lands negative. Self-Refine and forced Reflexion are still several points under the baseline at 7B; only the Best-of-N judge-vs-count gap shrinks toward zero.\n\nWhat is actually new is not the qualitative claim—Wang et al. already said the baseline often wins—but the design that makes it testable: per-method measured cost with interpolation on a K=16 pool, paired bootstrap and Holm (within-setting and all-36), the 1.5B/3B regime, and especially the fixed-sample judges-vs-counting contrast that isolates selection from sampling. The Reflexion non-firing result on 1.5B is a useful caution about adaptive methods that can silently become CoT. They also document a real Best-of-N scoring bug they caught by cross-check, re-ran the truncated cells, and ship an independent re-derivation script plus all generations. That is above the usual bar for this genre.\n\nSoft spots are mostly scope, and they already say so. Exact-match math is where majority vote is strongest and critique has the least room to help; generated tokens (not input or latency) is the accounting that flatters sequential methods; configs and prompts are fixed, not tuned. The self-inspection grouping is partly post-hoc and labeled as descriptive. None of that overturns the claim inside the stated bounds. Temperature is checked via SC@1; Monte-Carlo stability is tiny; grader validation is thorough.\n\nThis is for people who care about test-time compute evaluation and small/mid open models. Cite it when you need a cost-matched baseline or a warning about self-triggered control flow. I would send it to referees; the evidence matches the claim and the release is serious. Engage with it.","headline":"Solid statistical re-run of Wang et al.: no method beats cost-matched sampling, and self-inspection stays negative through 7B, with unusually clean artifacts.","tokens_in":25541,"tokens_out":507,"would_cite":true,"duration_ms":10766,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"No method that makes a model inspect its own answers beats simply sampling more times once every generated token is counted the same way.","keywords":["test-time compute","self-consistency","self-refine","reflexion","budget-matched evaluation","self-verification","mathematical reasoning","language models"],"falsifier":"Rerun the identical paired, cost-matched protocol on a larger model or with a trained verifier and show any self-inspection method landing reliably above the interpolated majority-vote curve on the same questions.","tokens_in":25451,"feed_emoji":"🔁","tokens_out":1007,"duration_ms":23816,"temperature":0.7,"pith_summary":"Many popular test-time tricks—self-critique, reflection, best-of-N selection, multi-agent debate—also make the model write far more text than one chain of thought. Extra text alone raises accuracy, so a win over a single sample does not prove the trick’s idea helped. This paper re-runs that comparison as a controlled experiment on open 1.5B–7B models and two math benchmarks: every method is measured by total generated tokens and scored against repeated sampling at that same cost, with paired bootstrap intervals and multiplicity correction. None of the thirty-six comparisons finds a reliable win for the fancy method; all eighteen that involve the model judging or rewriting its own output come out negative. As models grow, letting the model pick among its samples stops hurting relative to a majority vote, but forcing rewrite-and-retry loops still wastes tokens compared with drawing another independent attempt.","feed_headline":"Sample more: self-critique loses at equal token cost","feed_subtitle":"From 1.5B to 7B, rewriting and self-picks trail plain majority vote when every token is counted","key_machinery":"Cost-matched self-consistency curve: generate a pool of independent chains once, read majority-vote accuracy at every sample count from the same pool, then place each competing method against the curve at that method’s own measured completion-token cost.","core_discovery":"At equal generated-token cost, none of seven common test-time reasoning methods is reliably better than repeated sampling with majority vote on GSM8K and MATH-500 for Qwen2.5 models from 1.5B to 7B. Every comparison in which the model assesses or rewrites its own output is negative; Self-Refine and forced Reflexion remain several points below the matched baseline even at 7B. Holding the same eight samples fixed, counting the most common answer beats asking the model to choose, by large margins below 7B and by amounts no longer distinguishable from zero at 7B.","pith_inferences":["On open-ended tasks where answers cannot be majority-voted, the practical rival to self-critique shrinks to “pick one sample,” so the same mechanisms could look better there without contradicting these math results.","If override accuracy of an untrained self-verifier never crosses 50% on disagreements, scaling alone will not make judging beat counting; training or external verifiers would be required.","Evaluations that charge for input tokens or wall-clock latency would likely widen the gap against sequential self-refinement and debate, because those methods re-read long contexts and cannot parallelize rounds.","Silent control-flow collapse (a method that stops acting while keeping its name) is a general evaluation hazard for any agent loop gated on the model’s own correctness judgment."],"forward_implications":["A reported gain over one chain of thought is not evidence that planning, critique, reflection, or debate caused the gain unless an equal-token sampling baseline is beaten.","With a fixed generation budget on checkable math, spending tokens on another independent attempt is a better default than spending them on self-critique or self-rewrite.","Adaptive methods must report how often their self-triggered loop actually fires; otherwise a method can score well by silently collapsing into a cheap single sample.","Untrained same-model selection among samples reaches parity with majority vote near 7B by agreeing more often with the tally, not by becoming a better override judge.","Future method papers can settle the budget question cheaply by releasing one repeated-sampling accuracy-versus-cost curve on the same model and questions."],"fun_headline_variants":["Repeated sampling beats self-critique at equal token cost","Self-refine trails majority vote from 1.5B to 7B","No self-inspection method beats plain resampling","Rewriting loses to sampling more at matched budgets","Model self-picks lag majority vote even at 7B"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That counting generated tokens on exact-match math problems, with fixed untuned prompts and hyperparameters, is a fair test of whether self-inspection helps in the settings people actually care about.","fun_headline_variants_meta":{"raw":{"variants":["Repeated sampling beats self-critique at equal token cost","Self-refine trails majority vote from 1.5B to 7B","No self-inspection method beats plain resampling","Rewriting loses to sampling more at matched budgets","Model self-picks lag majority vote even at 7B"]},"model":"grok-4.5","effort":"low","cost_usd":0.004892,"raw_usage":{"total_tokens":1521,"prompt_tokens":1004,"num_sources_used":0,"completion_tokens":66,"cost_in_usd_ticks":48924000,"prompt_tokens_details":{"text_tokens":1004,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":451,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":1004,"tokens_out":66,"duration_ms":7810,"temperature":1.0,"reasoning_tokens":451,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T03:17:06.737464+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Rerun the identical paired, cost-matched protocol on a larger model or with a trained verifier and show any self-inspection method landing reliably above the interpolated majority-vote curve on the same questions.","supporting_citations":[],"review_version":1}