{"id":"f0f0685a-ce2b-402a-ac1b-cd7dcee518c7","arxiv_id":"2508.16514","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A controlled study of synthetic math data generation yields a new dataset blend (FLAMES) that improves fine-tuned 7B model scores on MATH and OlympiadBench, with caveats on checkpoint selection.","lead":"This paper introduces FLAMES, a standardized framework for comparing synthetic math problem generation strategies, and tests 12 data agents, 6 quality-control methods, and multiple teacher models under one pipeline. The authors use the findings to build an open-source-only training dataset that lifts small math models on benchmarks such as OlympiadBench and MATH.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"OOD/competition gains may be inflated by training/eval overlap: decontamination is only run against GSM8K/MATH test sets, not CollegeMath, GSMPlus, or OlympiadBench; an overlap audit is needed.","rationale":"The reader's weakest assumption—mixture proportions tuned on the same evaluation benchmarks—is a valid concern, but I see an even more direct threat to the headline numbers: the absence of decontamination against the OOD, robustness, and competition benchmarks that dominate the claimed gains. The paper explicitly reports decontamination only against GSM8K and MATH test sets, while the largest deltas are on OlympiadBench and GSMPlus. Because the synthesis models are pretrained on the open web and the Distraction Insertion agent mimics the GSMPlus perturbation style, near-duplicate leakage is plausible and would inflate the central comparison. This is a concrete, checkable threat rather than a stylistic or interpretational objection. The benchmark-guided mixture selection remains relevant, but it is secondary: the mixture search in Appendix D.2 considered only a handful of candidates, while a small number of overlapping OlympiadBench/GSMPlus examples could shift the reported margins by several points. I therefore recommend keeping the reader's conditional verdict: the paper is a useful controlled study, but the dataset-superiority claim should be accepted only after a decontamination audit of all evaluation benchmarks, not just GSM8K/MATH. My agreement_with_reader is 'partial' because the reader identified mixture selection as the weakest assumption, whereas I see the decontamination gap as at least as load-bearing and more directly falsifiable.","tokens_in":21811,"tokens_out":11904,"duration_ms":142841,"concrete_test":"Run the paper's own decontamination criterion (95% of the 8-grams in a test item contained in a synthetic problem, as described in Appendix A) between FLAMES Large and FLAMES XL and each of CollegeMath, GSMPlus (including the Distraction subset), and OlympiadBench. Count how many test-set problems have a near-duplicate in the training FLAMES data. If any are flagged, recompute the Table 3 and abstract deltas after removing those overlapping synthetic problems and check whether FLAMES still beats refreshed ScaleQuest by the claimed margins. If the overlap count is zero or negligible, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that FLAMES Large/XL outperforms public datasets on MATH, CollegeMath, GSMPlus, and OlympiadBench (Tables 3, 4, and abstract). The paper's decontamination step (Appendix A, Section 2) removes synthetic problems with high 8-gram overlap only against the GSM8K and MATH test sets. No decontamination is reported for CollegeMath, GSMPlus, or OlympiadBench. This is the most load-bearing gap because those are exactly the benchmarks where FLAMES reports its largest relative gains (e.g., +18.3 over refreshed ScaleQuest on OlympiadBench in Table 3, +4.9 on GSMPlus). The Taxonomy-Based Key Concepts agent generates problems from a broad taxonomy with no seed, so Qwen2.5-32B-Instruct—pretrained on the open web—may reproduce or near-duplicate existing competition or college problems. The Distraction Insertion agent modifies GSM8K problems by inserting misleading details, the same perturbation family as GSMPlus, making near-duplicates plausible. If even a few dozen of the 675 OlympiadBench items or a few hundred GSMPlus items appear in FLAMES Large/XL, the 'easy-to-hard generalization' and 'outperforms public datasets' claims are materially inflated. Additionally, the mixture proportions were chosen using the same five evaluation benchmarks (Appendix D.2, Table 11), and checkpoints were selected by highest average on GSM8K/MATH test sets (Appendix A), so the reported margins come from a pipeline optimized on the evaluation sets. The limitations section does not mention overlap or selection as risks. The concern is not intentional leakage; it is that the claimed margins have not been shown to survive an overlap check.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces FLAMES, a standardized framework for controlled comparison of synthetic math data pipelines. It studies 10 existing plus 2 novel data synthesis agents, 6 quality-control strategies, 2 problem-generation and 2 solution-generation models, and mixture proportions. The empirical findings are that complexity-enhancing agents help most, problem coverage matters more than solution precision, easy-to-hard generalization occurs, and a tuned blend (FLAMES Small/Large/XL) outperforms public datasets on MATH, CollegeMath, GSMPlus, and OlympiadBench. The internal comparisons are carefully controlled, but the dataset-level claims depend on decontamination and selection protocols that are currently incomplete.","tokens_in":22309,"tokens_out":4798,"duration_ms":49093,"significance":"If the dataset gains are robust, this is a valuable contribution: the framework enables apples-to-apples comparison of data synthesis strategies, the public baselines are refreshed with the same solution generator as a fairness control, and the final data pipeline uses only open-source models. The two novel agents are plausible and directly tested. The main risk is that the headline superiority over public datasets may be inflated by evaluation-set contamination and by tuning the mixture and checkpoints on the same benchmarks used for certification.","major_comments":[{"comment":"Decontamination is described only against the GSM8K and MATH test sets (Appendix A: 'decontaminating synthetic problems against the GSM8K and MATH test sets'). The largest reported gains are on OlympiadBench, CollegeMath, and GSMPlus, and these are exactly the benchmarks not audited. The Distraction Insertion agent (Figure 9) modifies GSM8K problems by inserting misleading details, the same perturbation family as GSMPlus; the Taxonomy-Based Key Concepts agent generates unseeded problems from a web-derived taxonomy and may reproduce or near-duplicate existing competition problems. Please run an overlap audit of FLAMES Large/XL against CollegeMath, GSMPlus, and OlympiadBench, report near-duplicate counts, and re-evaluate after removing them. The Limitations section (Section 7) does not mention this gap.","section":"§2, Appendix A; Tables 3, 9"},{"comment":"The FLAMES mixture proportions were selected by evaluating candidate mixtures on the same five evaluation benchmarks that later certify the dataset (Table 11, Appendix D.2), and checkpoints were selected by highest average on GSM8K/MATH (Appendix A). The public datasets were not tuned in this way. This makes the claimed superiority over public datasets partly a selection artifact. Please provide a holdout evaluation or otherwise quantify the effect of benchmark-guided selection on the reported margins.","section":"§5.3, Appendix D.2, Table 11"},{"comment":"The headline numbers are computed against different baselines. The abstract's +15.7/+4.5/+6.5/+3.1 match the original (non-refreshed) ScaleQuest in Table 9, while Section 5.3's +12.8/+5.8/+4.9/+1.3 appear to be per-column best refreshed baselines from Table 3. This makes the claimed improvements difficult to interpret. Please report all deltas against a single consistent baseline, preferably the best refreshed public dataset at comparable scale.","section":"Abstract vs §5.3, Tables 3 and 9"}],"minor_comments":[{"comment":"In the text describing mixture-D, 'IDC agent' should be 'IQC agent'.","section":"Appendix D.2"},{"comment":"The label 'DeepSeek-7B' for the solution-generation model should be expanded to 'DeepSeek-Math-7B-RL' to avoid confusion with the student model.","section":"Table 5"},{"comment":"The 'GSM8K/MATH 15K' row should clarify whether this is the combined train split size or a subsample.","section":"Table 3"},{"comment":"The statement 'data agents, while using only GSM8K and MATH as seed sets, lead to improvement over competition level benchmarks' is inaccurate for the QFT and Taxonomy agents, which do not use those seed sets. Please qualify.","section":"§5.2"},{"comment":"The limitations section should also acknowledge the decontamination scope and the benchmark-guided mixture selection as potential threats to the dataset-level claims.","section":"Section 7"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope and the controlled-framework contribution is solid. The main risk is contamination/selection artifact in the headline dataset comparisons; this is fixable with an overlap audit and a holdout-based evaluation. No concerns about attribution or novelty."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"FLAMES is worth reading. It provides a controlled comparison of 12 data agents and 6 quality-control strategies under a single pipeline, with the sensible touch of refreshing public datasets with the same solution teacher. The two new agents—Taxonomy-Based Key Concepts and Distraction Insertion—are modest twists on existing ideas, but they are clearly described and their targeted effects (OOD and robustness) show up in the right places. The dataset blend also lifts small models substantially on competition benchmarks, which is a real finding if it holds.\n\nThe central internal comparisons are carefully run: same student model, same 150K scale, same solution generator. The finding that coverage beats precision in quality control is plausible and independently useful, as is the result that solution-teacher quality matters more than problem-teacher quality.\n\nThe soft spots are mostly about the external claims. First, decontamination is only reported against GSM8K/MATH test sets. CollegeMath, GSMPlus, and OlympiadBench are not checked, and the Taxonomy agent generates unseeded problems while the Distraction Insertion agent perturbs GSM8K in ways that could resemble GSMPlus. The largest reported gains are precisely on those benchmarks. So the easy-to-hard generalization and 'outperforms public datasets' claims need an overlap audit. Second, the final mixture proportions were chosen by looking at the same five evaluation sets (Table 11), and checkpoints were selected by test-set performance. That doesn't invalidate the relative agent ranking, but it means the margins over public datasets are partly a selection artifact. Third, there are no error bars or seeds, and the dataset is not released, so external validation is impossible right now.\n\nThe limitations section acknowledges teacher dependence and bias, but does not mention benchmark overlap or selection on eval sets, which are the more pressing risks.\n\nThis paper is aimed at anyone building synthetic math data or comparing data-synthesis methods. The framework itself is a contribution worth citing. It deserves a serious referee—the right outcome is peer review, not desk rejection—but the referee should ask for a contamination check on all five benchmarks, error bars or multiple seeds, and ideally a release of the FLAMES dataset.","headline":"FLAMES is a well-controlled testbed for synthetic math data, but its headline gains over public datasets are likely inflated by eval-set tuning and missing decontamination.","tokens_in":22794,"tokens_out":1839,"would_cite":true,"duration_ms":20610,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 50-20-20-10 blend of four data-synthesis agents makes a 7B math model reach 81.4% on MATH, beating much larger models.","keywords":["math reasoning","synthetic data pipeline","data synthesis agents","fine-tuning","out-of-domain generalization","distraction robustness","easy-to-hard generalization","FLAMES dataset"],"falsifier":"Retrain a FLAMES mixture with the proportions chosen on only four of the five benchmarks, then evaluate on the held-out benchmark; if the margin over refreshed ScaleQuest shrinks or reverses on the held-out set, the reported superiority is partly benchmark tuning. A second check: replace the Qwen2.5-Math-7B-Instruct solution teacher with a weaker model and see whether the 'coverage beats precision' ordering flips on MATH.","tokens_in":21770,"feed_emoji":"🧮","tokens_out":6694,"duration_ms":61743,"temperature":0.7,"pith_summary":"FLAMES is a controlled framework for studying the math-data synthesis pipeline one factor at a time: same seed problems, same problem generator, same solution generator, same student model, same evaluation. Using it, the paper claims that the biggest gains come from agents that increase problem complexity, that keeping more problems with imperfect solutions beats strict precision filtering (the tested solvability filter even rejects 30% of real MATH problems), and that GSM8K/MATH-seeded synthetic data transfers to competition-level benchmarks. From those findings it builds FLAMES datasets—a 50/20/20/10 blend of Suggester-Editor, IQC, Taxonomy-Based Key Concepts, and Distraction Insertion—and reports that FLAMES Large outperforms all tested public math datasets on OlympiadBench, CollegeMath, GSMPlus, and MATH, with Qwen2.5-Math-7B reaching 81.4% on MATH. The reason to care: prior synthetic-math works used incomparable setups, so practitioners could not tell which pipeline choices actually drive reasoning gains.","feed_headline":"A 50-20-20-10 data mix lifts a 7B model to 81.4% on MATH","feed_subtitle":"Standardized pipeline shows complexity-boosting agents and broad coverage beat strict filtering—and beat larger models.","key_machinery":"The load-bearing object is the FLAMES framework: a standardized synthetic-data pipeline with fixed seed datasets (GSM8K and MATH), fixed problem-generation model (Qwen2.5-32B-Instruct), fixed solution-generation model (Qwen2.5-Math-7B-Instruct), fixed student model (DeepSeek-Math-7B), and fixed first-solution quality control. Inside it, the paper compares 12 data agents—prompt-based problem-synthesis strategies—including two new ones: Taxonomy-Based Key Concepts, which generates problems from a curated math taxonomy with no seed problems, and Distraction Insertion, which adds irrelevant details without changing the answer. The framework's one-factor-at-a-time design is what converts a jumble","core_discovery":"On its own terms, the paper's discovery is that the math data synthesis pipeline can be decomposed into controllable factors and that a few simple design rules dominate: complexity-enhancing agents (Suggester-Editor and IQC) give the best overall improvements; coverage of generated problems matters more than the reliability of their solutions, so 'keep the first solution' and 'majority + first' outperform solvability and reward-model filtering; and the choice of solution-generation teacher matters more than the choice of problem-generation model. It then claims that combining 50% Suggester-Editor, 20% IQC, 20% Taxonomy-Based Key Concepts, and 10% Distraction Insertion produces a dataset that","pith_inferences":["The exact 50/20/20/10 proportions were picked on the same five benchmarks used for the final comparison, so part of the margin over public datasets could be selection artifact; a cleaner test would tune on a subset and hold out one benchmark.","The coverage-beats-precision result was established with Qwen2.5-Math-7B-Instruct as solution teacher; with a weaker teacher, noisy solutions could plausibly reverse the ordering.","The two novel agents suggest a general recipe: explicitly target the failure mode (OOD breadth via taxonomy, distraction robustness via insertion) and add a small portion of that data to a strong base mixture.","The easy-to-hard transfer result is tested only on English benchmarks; applying the same pipeline to other languages or formal math domains would show whether the finding generalizes."],"forward_implications":["With a fixed generation budget, complexity-enhancing agents should be preferred: they improved in-domain, robustness, and competition metrics simultaneously.","Strict solvability and reward-model filtering can be skipped; keeping first solutions and broader problem coverage performs better and costs less compute.","Synthetic data seeded only by GSM8K and MATH improved OlympiadBench, so easy problems can transfer to hard competition problems.","Investing in a stronger solution-generation teacher yields more student-model improvement than a stronger problem generator.","A small share (10%) of distraction-insertion data improves robustness and overall average without hurting in-domain scores."],"supporting_citations":[{"why":"Supplies the Suggester-Editor and Ask-Me-Anything agents and the OrcaMath baseline dataset, on which the FLAMES mixture improves.","marker":"(Mitra et al., 2024)"},{"why":"Supplies the Paraphrasing, Self-Verification, and FOBAR agents, plus MetaMathQA as a baseline dataset.","marker":"(Yu et al., 2023a)"},{"why":"Supplies the QFT agent, the solvability filter, and the ScaleQuest baseline that FLAMES is compared against; also the reward-model filtering method studied in Section 4.","marker":"(Ding et al., 2024)"},{"why":"Supplies the Iterative Question Composing (IQC) agent whose data is 20% of the FLAMES mixture, and the MMIQC baseline dataset.","marker":"(Liu et al., 2024b)"},{"why":"Supplies the Key Concepts and Seeded Key Concepts agents and the key-concept generation approach that the novel Taxonomy-Based Key Concepts agent extends.","marker":"(Huang et al., 2024a)"},{"why":"Supplies the Few-Shot agent and OpenMathInstruct-2, a large open-source baseline dataset.","marker":"(Toshniwal et al., 2024a)"},{"why":"Supplies the GSM8K seed problems and in-domain evaluation set used throughout the framework.","marker":"(Cobbe et al., 2021)"},{"why":"Supplies the MATH seed problems and evaluation set, and the MATH500 subset used to show the solvability filter rejects real problems.","marker":"(Hendrycks et al., 2021)"},{"why":"Supplies GSMPlus, the robustness evaluation benchmark whose distraction subset is used to validate the Distraction Insertion agent.","marker":"(Li et al., 2024)"},{"why":"Supplies DeepSeek-Math-7B, the student model used for the main controlled comparisons.","marker":"(Shao et al., 2024)"}],"fun_headline_variants":["Data mix 50-20-20-10 pushes 7B model past 405B","Smarter math data synthesis: 7B model tops 405B","FLAMES reveals data mix recipe for math reasoning","Data mix beats model size: 7B wins on MATH","Synthetic data recipe: complexity and coverage win"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing assumption is that the FLAMES mixture proportions, selected by running candidate mixtures on the same five evaluation benchmarks later used to certify the dataset, do not create an unfair advantage over public datasets whose compositions were not tuned on those benchmarks.","fun_headline_variants_meta":{"raw":{"variants":["Data mix 50-20-20-10 pushes 7B model past 405B","Smarter math data synthesis: 7B model tops 405B","FLAMES reveals data mix recipe for math reasoning","Data mix beats model size: 7B wins on MATH","Synthetic data recipe: complexity and coverage win"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000617,"raw_usage":{"total_tokens":2746,"prompt_tokens":834,"completion_tokens":1912,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":578,"completion_tokens_details":{"reasoning_tokens":1823}},"tokens_in":578,"tokens_out":1912,"duration_ms":14879,"temperature":1.0,"reasoning_tokens":1823,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:15:02.470442+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain a FLAMES mixture with the proportions chosen on only four of the five benchmarks, then evaluate on the held-out benchmark; if the margin over refreshed ScaleQuest shrinks or reverses on the held-out set, the reported superiority is partly benchmark tuning. A second check: replace the Qwen2.5-Math-7B-Instruct solution teacher with a weaker model and see whether the 'coverage beats precision' ordering flips on MATH.","supporting_citations":[{"cited_title":"small- size","cited_arxiv_id":null,"evidence_quote":"Supplies the MATH seed problems and evaluation set, and the MATH500 subset used to show the solvability filter rejects real problems."}],"review_version":1}