{"id":"b144db1e-c69b-4fe1-938e-1fa75a3e31af","arxiv_id":"2608.05148","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new collection of 50 procedural generators designed for completion-supervised fine-tuning beats three existing procedural collections and a no-procedural baseline on reasoning benchmarks at 3B scale in mean scores.","lead":"Researchers built Reasoning Core, a set of 50 computer programs that generate millions of solvable reasoning questions for training AI language models. In matched fine-tuning runs on small models, this data scored higher on reasoning benchmarks than three existing procedural collections and no procedural data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Design-principle claims rest on small-model rankings with poor cross-family and duration stability (Tables 3, 6); the causal link from compact targets/calibrated difficulty to the 3B advantage is unsupported.","rationale":"The reader's weakest assumption is the load-bearing one: task-level utility rankings from 135M/360M models at 300 steps are used to guide generator design, yet the headline empirical claim is measured at 3B/2,400 steps. I agree with this assessment. The literal headline claim—highest mean scores on DROP, LogiQA, and ARC-Challenge in Table 2—is directly supported by the reported point estimates, so the numerical result itself is not the soft spot. The vulnerability is the causal narrative: the conclusion states that compact targets and calibrated difficulty matter, generalizing a negative result on step-by-step rationales (Section 5.2.3) that appears to have been validated only at small scale and short durations. Table 6's collection-order agreement for OLMo (0.00 from 300 to 1,200 updates) and Table 3's low cross-family task-level agreement (0.34–0.38 for Core) show that the development signal is fragile. The paper repeatedly hedges that the diagnostic is a low-cost heuristic, but it still draws design conclusions from it without scale caveats. A 3B task-level panel for a subset of generators would settle whether the 360M rankings are predictive; if they are not, the design-principle contribution is unvalidated, though the resource and the direct 3B comparison remain valuable. The reader's CONDITIONAL verdict appropriately reflects this: the paper can be repaired by adding such a transfer check or by softening the causal claims. No change to the verdict is needed.","tokens_in":18142,"tokens_out":10252,"duration_ms":125028,"concrete_test":"Run isolated 300-update interventions for a stratified sample of 15 Reasoning Core generators (5 math/proof, 5 logic/state, 5 graph/language) on SmolLM3-3B-Base, with the same main stream, 20% token budget, and paired seeds as Table 2; compute Kendall's τb between the resulting BBH-development NLL utility ranks and the SmolLM2-360M 300-update ranks from the released diagnostics. If the 95% confidence interval for τb includes 0 or even slightly positive (e.g., upper bound below 0.5), the 360M-based task rankings do not predict relative utility at 3B, so the design decisions made from them are unsupported and the claim that compact targets/calibrated difficulty drive the headline advantage loses its evidence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution is design guidance: compact canonical answers and calibrated difficulty are said to drive training utility. That guidance is based on the development loop in §4.3: 300-update isolated-task runs on SmolLM2-360M and OLMo-1B, scored by BBH-development NLL and FineWeb NLL. The reliability of this signal is weak exactly where it matters. Table 3 reports Kendall's τb of only +0.34 (135M–OLMo) and +0.38 (360M–OLMo) for Reasoning Core task ranks; the paper uses these ranks to make generator-level decisions. Table 6 shows OLMo's collection-order agreement is 0.00 from 300 to 1,200 updates (3/6 pairwise orders), meaning the ordering of the four collections at the development duration is uncorrelated with the ordering at 1,200 steps on the very model used for development. No task-level agreement is reported between the 360M development model and the 3B model used for the headline result. The compact-vs-rationale negative result in §5.2.3 (parsing and graph pathfinding) is presented without stating that it was validated at small scale only; the conclusion generalizes it: 'compact canonical answers outperformed faithful step-by-step traces of the correct algorithm.' If small-model 300-step utility does not transfer to 3B/2,400 steps, the 3B advantage of Reasoning Core may reflect collection-specific task semantics rather than the design factors the paper emphasizes, and the design-principle contribution would be unvalidated. The paper is honest about using the small-model diagnostic as a heuristic, but it does not establish that the heuristic is predictive in the setting where the headline comparison is made.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Reasoning Core, a collection of 50 procedural generators for verifiable reasoning problems, and compares it with Procedural Warmup, Reasoning Gym, and SynLogic under a matched completion-supervised fine-tuning protocol. The headline result is that at 3B parameters and 2,400 updates, Reasoning Core achieves the highest mean scores on DROP, LogiQA, and ARC-Challenge, exceeding the main-only baseline and the three alternative procedural collections. The paper also reports development diagnostics that lead to design guidance (compact targets and calibrated difficulty), a negative result on step-by-step rationale targets, a zero-shot solvability analysis, and a single-seed verifier-backed RL demonstration. Extensive reproducibility artifacts, semantic audits, and public releases of code and data are described.","tokens_in":18427,"tokens_out":5042,"duration_ms":60368,"significance":"If the headline result holds, the paper provides a useful public resource and one of the more controlled comparisons of procedural data design for completion-supervised reasoning training. The matched paired-seed protocol, shared optimization configuration, and repository-scale audit are genuine strengths, and the explicit discussion of limitations is welcome. The main broader claims, however, are design principles derived from small-model, short-duration development panels whose cross-family and cross-duration stability is weak; those claims currently outrun the evidence. The resource and the audit methodology are likely to be valuable to the community regardless of the outcome of the design-principle claims.","major_comments":[{"comment":"The design guidance in Sections 5.2.2 and 7 rests on 300-step isolated-task panels using SmolLM2-135M, SmolLM2-360M, and OLMo-1B. Table 3 reports Kendall tau values of only 0.34 (135M to OLMo) and 0.38 (360M to OLMo) for Reasoning Core task ranks, and Table 6 shows that OLMo's collection ordering at 300 updates is uncorrelated with its ordering at 1,200 updates (tau = 0.00, preserved pairwise orderings 3/6). No task-level agreement is reported between the 360M development model and the 3B model used for the headline result. The paper should either provide evidence that these developmental signals transfer to the 3B/2,400-step setting or explicitly scope the design-principle claims to the measured regimes.","section":"Section 5.2.1; Tables 3 and 6"},{"comment":"The negative result on step-by-step rationale targets is stated for parsing and graph pathfinding tasks, but the section does not report the model size, training duration, or number of seeds for those experiments; given the development setup in Section 4.3, it appears to be based on small-model, 300-step runs. The conclusion in Section 7 nevertheless states globally that 'compact canonical answers outperformed faithful step-by-step traces of the correct algorithm.' This overgeneralizes the small-scale finding, and the claim should be either validated at the 3B scale or explicitly restricted to the settings in which it was measured.","section":"Section 5.2.3; Section 7"},{"comment":"The primary 3B comparison is reported as means with sample standard deviations, but no paired-difference significance tests or confidence intervals are provided. Several differences are within one standard deviation of each other (for example, LogiQA: Reasoning Core 47.8±0.7 vs SynLogic 47.1±0.5; ARC-Challenge: Reasoning Core 51.3±0.5 vs Reasoning Gym 51.1±0.2), and with five seeds these differences may not be reliable. Because the abstract and Section 5.1 claim that Reasoning Core 'exceeds' the alternatives, the paper should report paired differences with confidence intervals or a test across seeds, and should address multiple comparisons across the four proxy benchmarks.","section":"Section 5.1; Table 2"},{"comment":"Development decisions were made using BBH-development NLL and FineWeb NLL, while the primary held-out compound includes BBH-test from the same benchmark family. The non-algorithmic/algorithmic partition reduces direct overlap, but the paper does not justify why BBH-development utility is a valid proxy for BBH-test and the other held-out benchmarks. This matters because Section 5.1 reports that BBH-development utility does not transfer uniformly across evaluations; the reader needs a clearer statement of which development decisions are assumed to generalize and which are only in-sample diagnostics.","section":"Section 4.3; Section 5.1"}],"minor_comments":[{"comment":"The color legend ('darker teal') is not accessible or reproducible in grayscale; consider adding numeric deltas or textual annotations for improvements over the main-only baseline.","section":"Table 2"},{"comment":"The caption notes that panel-specific y-axis ranges prevent cross-panel comparison, but the axes are not labeled on each panel; adding per-panel labels would make the figure easier to read.","section":"Figure 1"},{"comment":"The Core/Gym equal mix achieves a larger BBH-test NLL reduction (36.25) than Reasoning Core alone (29.17); the main text does not discuss this, and it should be addressed as a possible interaction or mixture effect.","section":"Appendix D.1, Table 5"},{"comment":"The phrase 'exceeding both the baseline without procedural data and all three alternative procedural collections' should be qualified as a mean-level claim, given the statistical concerns in the main comparison.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The resource and audit work are solid and the paper is unusually transparent about limitations. The main risk is that the design-principle conclusions are presented more broadly than the small-model, short-duration evidence supports. I would be willing to accept after the authors add paired-difference inference and explicitly scope the cross-family, cross-duration, and cross-scale claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this paper ships a real resource—50 permissively licensed generators with semantic scorers, difficulty controls, audit scripts, and a 10B-token released dataset—and the audit work is the most interesting part. They found material defects in 13 of 105 Reasoning Gym tasks and 9 SynLogic generators, including scorers that give credit to any nonempty answer and logic puzzles whose displayed constraints contradict the stored solution. That is a concrete service to the community, and the paper is appropriately careful not to call it a quality ranking of the two libraries.\n\nThe SFT comparison is also better designed than most of this literature: paired seeds, shared data order, fixed optimization config across collections, and matched main-only baselines. The headline 3B result is honestly stated as highest mean scores on DROP, LogiQA, and ARC-C, which Table 2 supports, though the error bars overlap and there are no significance tests. The paper also discloses that the RL comparison is single-seed and illustrative. That is the right level of restraint.\n\nThe soft spot is the design-guidance story. The claims that compact canonical targets and calibrated difficulty drive training utility are inferred from 300-step isolated-task runs on 135M/360M models, and the paper's own numbers show this diagnostic is shaky exactly where it matters. Table 3 reports cross-family Kendall tau of only 0.34–0.38 between the SmolLM2 models and OLMo-1B for Reasoning Core task ranks. Table 6 shows OLMo's collection-order agreement is 0.00 from 300 to 1,200 updates—the ordering at the development duration is uncorrelated with the ordering later in training on the very model used for development. No task-level agreement is reported between the 360M development model and the 3B model behind the headline result. The compact-vs-rationale negative result is also presented as a general conclusion even though it was validated only at small scale. I don't think this invalidates the resource—the 3B comparison is still a real measurement, and the collection may well be better for the stated reasons—but the paper should either provide direct evidence that the small-model signal transfers to the 3B setting or soften the causal language.\n\nMissing significance testing is a fixable problem. The authors can report paired-difference intervals or at least paired permutation tests across the five seeds, and they should reframe the design principles as hypotheses suggested by the development loop rather than conclusions established by it.\n\nWho gets value from this: anyone doing post-training data engineering, especially people building verifiable reasoning datasets. The audit methodology alone is worth reading. It deserves a serious referee; I would recommend conditional acceptance with the statistical and framing revisions above.","headline":"A genuinely useful procedural-data resource with an honest matched comparison, but the design-principle claims rest on small-model diagnostics whose transfer to the 3B headline setting is not established.","tokens_in":18984,"tokens_out":1387,"would_cite":true,"duration_ms":20383,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Procedural data from 50 generators beats three rival reasoning collections on a matched 3B training run.","keywords":["procedural generation","completion-supervised fine-tuning","reasoning training","verifiable rewards","synthetic data","task utility","semantic auditing"],"falsifier":"Train the 3B model on the same 50 generators with step-by-step trace targets instead of compact answers, holding the token budget and seeds fixed; if Reasoning Core no longer beats the other collections, the paper's compact-answer design claim fails. A second check is to rank the 50 generators by 135M-model utility and train a 3B model on the bottom-ranked half; the transfer assumption predicts it should transfer worse.","tokens_in":17917,"feed_emoji":"🧩","tokens_out":4966,"duration_ms":52425,"temperature":0.7,"pith_summary":"The paper tries to establish that broad procedural generation can be turned into effective completion-supervised fine-tuning data, not just a source for reinforcement learning warmups or narrow abstract sequences. It introduces Reasoning Core, a collection of 50 generators spanning mathematics, logic, planning, state tracking, formal languages, structured data, games, causality, and code, each with semantic scoring, difficulty controls, and compact canonical targets. In the primary 3B comparison after 2,400 updates, Reasoning Core achieves the highest mean scores on DROP, LogiQA, and ARC-Challenge, beating both the matched main-only baseline and the three alternative procedural collections. The paper also argues that semantic validity alone does not ensure training utility: target format and difficulty calibration matter, and compact answers outperform faithful step-by-step traces of the correct algorithm.","feed_headline":"50 procedural generators beat three rival reasoning datasets","feed_subtitle":"In a matched 3B training run, broad verifiable tasks beat narrow warmups and RL-style environments.","key_machinery":"The central object is Reasoning Core, a library of 50 procedural generators covering mathematics, logic, planning, state tracking, formal languages, structured data, games, causality, and code. Each generator maps a seed and difficulty level to a prompt, a canonical compact reference answer, and a semantic scorer that accepts any valid answer even when the training target is one deterministic serialization. The argument is carried by a matched experimental protocol: the same seed, optimizer, and 80/20 main-to-auxiliary token split are used for every collection, paired with a main-only baseline, so reported effects isolate the marginal contribution of auxiliary procedural data.","core_discovery":"On the paper's own terms, the central discovery is a design recipe for procedural training data plus evidence that it works. Given a fixed token budget and a matched completion-supervised protocol, a heterogeneous collection of 50 verifiable generators with semantic scorers, difficulty levels, and deterministic compact targets transfers better to held-out reasoning benchmarks than the main-only baseline, Procedural Warmup, Reasoning Gym, or SynLogic. The recipe's load-bearing design choices are answer representation and difficulty: compact canonical completions beat longer, semantically correct algorithm traces, and tasks that saturate native reward do not transfer better than tasks with intermediate progress. The paper further reports that semantic audits, combining model-assisted review, human adjudication, and regression testing, expose material defects in the external collections, so procedural generation alone is not a guarantee of correctness.","pith_inferences":["If compact targets are the active ingredient, answer serialization becomes a transferable design knob that other synthetic-data pipelines could tune without changing generators; the paper does not test this across other collections.","The paper's fixed 20% auxiliary token share is validated on only one collection at one duration, so the optimal auxiliary ratio for Reasoning Core remains an open testable question.","Because cross-family task-rank agreement is much lower than within-family agreement, the authors' own developmental diagnostic suggests that generator selections may need to be re-measured per model family; the paper stops short of claiming a universal task ranking."],"forward_implications":["Reasoning Core can serve as a supervised fine-tuning layer before reinforcement learning, since its scorers double as outcome rewards and a matched RL run shows viability.","Broad, heterogeneous procedural mixtures are a viable alternative to narrow procedural warmups for improving downstream reasoning benchmarks.","Answer serialization is a training-data design axis: compact canonical targets are better than faithful step-by-step traces under this protocol.","Semantic auditing of procedural repositories is necessary; defects in scoring or generation can silently corrupt comparisons.","The released 10-billion-token procedural pile gives other training pipelines a ready-made source of verifiable reasoning examples."],"supporting_citations":[{"why":"Supplies the Procedural Warmup baseline and the precedent for completion-supervised procedural warmup.","marker":"(Jiang et al., 2026)"},{"why":"Supplies Reasoning Gym, the main alternative collection and the RL training recipe matched in Section 5.4.","marker":"(Stojanovski et al., 2025)"},{"why":"Supplies SynLogic as the second alternative procedural collection of verifiable logical games.","marker":"(Liu et al., 2025)"},{"why":"Supplies DOLCI, the curated instruction and conversation data in the main stream and main-only baseline.","marker":"(Olmo et al., 2025)"},{"why":"Supplies FineWeb-Edu, the text stream blended with DOLCI as the main data.","marker":"(Penedo et al., 2024)"},{"why":"Supplies BBH, split into development and held-out parts for evaluation and task selection.","marker":"(Suzgun et al., 2023)"}],"fun_headline_variants":["Broad tasks, compact answers: recipe for procedural data success","Why 50 procedural generators beat three rival reasoning sets","Reasoning Core tops 3B benchmarks with 50 verifiable generators","Audits expose flaws in procedural data, but Reasoning Core still wins","Design prescription: compact targets, calibrated difficulty transfer best"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The design choices behind Reasoning Core are validated on 135M- and 360M-parameter models after 300 updates, so the headline 3B result assumes that small-model, short-run task rankings predict which generators help a 3B model after 2,400 updates.","fun_headline_variants_meta":{"raw":{"variants":["Broad tasks, compact answers: recipe for procedural data success","Why 50 procedural generators beat three rival reasoning sets","Reasoning Core tops 3B benchmarks with 50 verifiable generators","Audits expose flaws in procedural data, but Reasoning Core still wins","Design prescription: compact targets, calibrated difficulty transfer best"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00167,"raw_usage":{"total_tokens":6610,"prompt_tokens":915,"completion_tokens":5695,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":531,"completion_tokens_details":{"reasoning_tokens":5611}},"tokens_in":531,"tokens_out":5695,"duration_ms":42804,"temperature":1.0,"reasoning_tokens":5611,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T04:13:21.012627+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the 3B model on the same 50 generators with step-by-step trace targets instead of compact answers, holding the token budget and seeds fixed; if Reasoning Core no longer beats the other collections, the paper's compact-answer design claim fails. A second check is to rank the 50 generators by 135M-model utility and train a 3B model on the bottom-ranked half; the transfer assumption predicts it should transfer worse.","supporting_citations":[],"review_version":1}