REVIEW 4 major objections 6 minor 1 cited by
Improving Mathematical Reasoning Capabilities of Small Language Models via Feedback-Driven Distillation
T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A feedback loop in which the student's right and wrong answers decide what new questions the teacher writes lifts sub-billion-parameter models to state-of-the-art math reasoning, with FlanT5-Large reaching 49.43% on GSM8K and 67.55%…
desk verdict A plausible feedback-driven data-generation idea undermined by missing controls—worth sending to review but only after the authors address the data-size confound. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the easy/hard feedback split that turns the student's own performance into curriculum signal. After each fine-tuning round, the student SLM is asked to solve every question in the distillation pool; a question that the student answers correctly is classified easy and one it answers incorrectly is classified hard. The teacher LLM is then prompted with two different templates — one that asks for a more complex version of the seed question, one that asks for a new question of similar difficulty — so that the newly generated data is targeted at exactly the skills the student does and does not have. The accepted questions, filtered by majority voting over several PoTs, are added to the pool and the student is retrained from scratch, and the loop repeats with a shrinking seed pool that drops the already-mastered easy questions. This mechanism is what converts the student's mistakes into larger, more complex, and more diverse training data.
What would settle it
Take the newly generated questions from a finished FDD run, have independent annotators or a strong verifier solve them without seeing the recorded answers, and measure how often the majority-vote answer is wrong; a substantial error rate would show the gains could come from training on flawed patterns rather than from the feedback mechanism. A second decisive check is to replace the easy/hard targeting with random selection of the same number of new questions: if the accuracy gap disappears, the feedback split is not what carries the claim.
Extended reading notes
Core claim
The central discovery is a three-stage distillation loop the paper calls Feedback-Driven Distillation (FDD). In the initialization stage, a large teacher model writes Program-of-Thought (PoT) solutions — short Python programs that an interpreter executes — for the GSM8K training problems, and the solutions whose executed answers match the gold answers form the first distillation dataset. After fine-tuning the small model on this dataset, the student's performance splits the questions into easy (solved) and hard (failed) buckets; the teacher then generates more complex variants of easy questions and similar-difficulty variants of hard ones, accepts a new question when at least one of several PoTs agrees with the majority-vote answer, and merges everything into an enlarged dataset used to fine-tune the student from scratch. Over three rounds this yields 49.43% GSM8K accuracy and a 67.55% average across GSM8K, ASDiv, SVAMP, and MultiArith for FlanT5-Large, and the ablation studies attribute the gains to combining both generation strategies, to more rounds, and to more reasoning paths per question.
Load-bearing premise
The load-bearing premise is that the teacher LLM's generated questions and the majority-vote answers chosen as their gold labels are correct, because Section 3.2 accepts a new question whenever at least one Program-of-Thought solution matches the voting result, with no human validation, no external ground truth, and no measurement of how often flawed questions slip through.
Editorial extensions
If this is right
- Models far below a billion parameters can reach accuracy levels previously seen only in much larger models, so mathematically capable assistants become deployable on low-resource devices.
- The gains transfer outside the training distribution: models fine-tuned only on GSM8K-derived questions improve on ASDiv, SVAMP, and MultiArith, evidence that the loop teaches generalizable procedures rather than memorized answers.
- Distillation data can be manufactured rather than collected: the loop keeps producing novel questions from existing ones, so the bottleneck becomes teacher-LLM compute instead of human annotation.
- Each extra round and each extra reasoning path per question adds measurable accuracy, implying that data scale and diversity are active ingredients of the distillation gain, not incidental side effects.
Reading between the lines
- I would expect the easy/hard split to give the largest advantage early in the loop, when the student's errors are most informative; at very large data budgets the gap versus uniform question generation should shrink, a prediction the paper's ablations do not test.
- The majority-vote filter screens rationales, not questions: a plausible but mathematically wrong new question whose generated programs consistently agree would still enter the training set, so teacher reliability may be the real ceiling on the method.
- The same loop should transfer to other domains where correctness is cheaply checkable, such as code generation and formal or executable reasoning; in domains with no executable verifier, the voting filter would be the bottleneck.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Feedback-Driven Distillation (FDD), a framework to improve mathematical reasoning of small language models (SLMs, ≤1B parameters) by iteratively expanding a distillation dataset. An initial dataset is built by prompting ChatGPT (gpt-3.5-turbo) to produce Program-of-Thought rationales for GSM8K training questions, which is used to fine-tune FlanT5 models. The fine-tuned SLM then classifies questions as easy or hard; for easy questions the LLM generates more complex variants, and for hard questions it generates similar-difficulty variants. These generated questions (with majority-vote answers and matching PoTs) are added to the distillation dataset and the SLM is fine-tuned from scratch; the process is repeated over multiple rounds. Experiments on GSM8K, ASDiv, SVAMP, and MultiArith report substantial gains over the initialization-only baseline, with FlanT5-Large reaching 49.43% on GSM8K and 67.55% average accuracy, which the paper describes as state-of-the-art for SLMs.
Significance. If the causal claim is correct, FDD would be a useful and practical contribution: using the student's own errors to guide data generation could boost sub-billion-parameter models to performance levels normally associated with much larger models, with broad implications for low-resource deployment. The paper has genuine strengths: it evaluates on four external benchmarks, provides internal ablations for generation strategy, number of rounds, and number of reasoning paths, and it attempts a data-leakage analysis. The ablation results in Table 2 and Figures 2 and 3 are internally consistent and suggest that iterative data expansion helps. However, the central contribution—that the easy/hard feedback split is what drives the gains—is not cleanly identified because the ablation arms differ in dataset size and generation strategy simultaneously. The quality of the self-generated training data is also not validated. These gaps currently leave the main claim underdetermined, though they are addressable within the scope of the manuscript.
major comments (4)
- [4.4, Table 2] The ablation for question generation strategies compares adding 1,666 complex questions, 1,928 diverse questions, and all 3,594 questions against the initialization-only baseline, but these arms vary both the strategy and the number of added questions at the same time. There is no control that adds the same number of generated questions while using a uniform or reversed assignment of seeds to the complex/diverse instruction types. Since prior work (MetaMath, MuggleMath, WizardLM) shows that unguided question augmentation alone can improve math reasoning, the observed +8.0 GSM8K improvement from 'All Questions' could be a simple data-quantity or data-diversity effect rather than evidence that the easy/hard feedback mechanism is causally responsible. Please add a matched-size control (e.g., random assignment of seeds to the two instructions, or a uniform generation strategy with the same total data budget) and report the comparison.
- [3.2] New questions are accepted whenever at least one Program-of-Thought matches the majority vote among the generated PoTs, and the majority-voted answer is used as the gold answer. The paper provides no human validation, no estimate of label noise, and no analysis of how many generated questions are flawed, ambiguous, or unanswerable. Because the method's premise is that feedback-driven generation produces a high-quality training signal, a sample-based human evaluation or an independent agreement filter is needed to rule out the possibility that the gains come from memorizing incorrect patterns rather than from learning better reasoning.
- [4.3, Table 1] The state-of-the-art claim is based on cross-paper comparisons with different base models, training budgets, and decoding settings; for example, Zhu et al. [5] and [6] use different SLM backbones and dataset sizes. No controlled comparison is reported under matched base model and data budget. The internal FDD gains are convincing as an ablation, but the headline 'SOTA' claim should either be rephrased as 'strong performance relative to published numbers' or supported by a controlled run against the closest prior distillation method on the same FlanT5 backbone and a matched training budget.
- [4.7, Eq. (4), Figure 4] The ROUGE-L similarity of roughly 0.11–0.12 is not a sufficient test for data leakage. Average lexical overlap between generated questions and test questions would not detect semantically equivalent rephrasings, swapped numbers, or multi-hop compositions, which are exactly the kind of variations the method generates. The conclusion that the results 'effectively eliminat[e] the impact of data leakage' is therefore too strong. I recommend a stronger test (e.g., n-gram overlap, model-based paraphrase detection, or measuring whether test-question variants appear among generated questions) or, at minimum, a caveat that this test only addresses surface-level lexical leakage.
minor comments (6)
- [3.4] Section 3.4 contains the typo 'Initiation stage' where 'Initialization stage' is meant.
- [Figure 4] The legend labels 'Complex' and 'Diverse' are ambiguous with respect to the multi-round setup; please clarify which bars correspond to generations from easy seeds versus hard seeds in each round.
- [Table 1] The column heading appears as 'A VG' in the table body but 'AVG' in the text; please normalize the formatting.
- [4.1] The text writes 'SV AMP' with a space; the correct benchmark name is 'SVAMP'.
- [Appendix A] The instruction templates are labeled Table A.3 and A.4 but are not referenced by those numbers in the main text; please align the cross-references.
- [4.3] The 'AVG' column in Table 1 averages over different sets of benchmarks for different models (e.g., MultiArith is missing for several proprietary and open-source LLMs); please state this explicitly so the averages are not read as directly comparable.
Circularity Check
No significant circularity; benchmark evaluations are external and the central claim is not defined by its inputs.
full rationale
The paper's central claim is that feedback-driven question generation (complex variants of easy questions, diverse variants of hard questions) followed by re-fine-tuning improves SLM mathematical reasoning. The reported improvements are measured on held-out external test sets (GSM8K test, ASDiv, SVAMP, MultiArith) that are never used to construct the distillation data; the paper even includes a ROUGE-L leakage analysis to show generated questions are not near test questions. The only self-referential element is that gold answers for newly generated questions are produced by the same ChatGPT model that generated the questions, with answers selected by PoT majority voting rather than by an external oracle. This is a weak-supervision or self-training design choice and a data-quality caveat, not a circular derivation: the evaluation target (accuracy on external benchmarks) is not defined in terms of, nor forced by, the ChatGPT-generated labels. The self-citations to Zhu et al. [5,6] supply the PoT fine-tuning recipe, but that baseline is itself evaluated externally and is not used as a uniqueness theorem to forbid alternatives; hence it is not load-bearing circularity. Ablation comparisons in Table 2 do not control for dataset size, but an uncontrolled confound is an experimental weakness, not an equivalence-by-construction. No equation in the paper defines the predicted accuracy in terms of the fitted or training quantities. Therefore no significant circularity is present.
Assumptions & free parameters
free parameters (4)
- Number of distillation rounds =
3
- Number of PoTs generated per question =
4
- Decoding temperatures =
1.0 for question generation, 0.7 for PoT generation
- Demonstration set size k =
not reported
assumptions (4)
- domain assumption ChatGPT (gpt-3.5-turbo) generates correct and useful math questions and PoT solutions at the quality assumed by the method.
- domain assumption Accuracy on GSM8K, ASDiv, SVAMP, and MultiArith measures mathematical reasoning and transfers to real-world use.
- domain assumption Fine-tuning the SLM from scratch on the enlarged dataset is the right protocol and avoids catastrophic forgetting.
- ad hoc to paper A ROUGE-L similarity around 0.11-0.12 implies no data leakage from the test sets.
Cite this review
Pith. "Pith review of Improving Mathematical Reasoning Capabilities of Small Language Models via Feedback-Driven Distillation." pith.science (2026). https://pith.science/paper/FRM66FPS
@misc{pith2026241114698,
author = {Pith},
title = {Pith review of: Improving Mathematical Reasoning Capabilities of Small Language Models via Feedback-Driven Distillation},
year = {2026},
howpublished = {\url{https://pith.science/paper/FRM66FPS}},
note = {Machine review of arXiv:2411.14698}
}
abstract
Large Language Models (LLMs) demonstrate exceptional reasoning capabilities, often achieving state-of-the-art performance in various tasks. However, their substantial computational and memory demands, due to billions of parameters, hinder deployment in resource-constrained environments. A promising solution is knowledge distillation, where LLMs transfer reasoning capabilities to Small Language Models (SLMs, $\le$ 1B parameters), enabling wider deployment on low-resource devices. Existing methods primarily focus on generating high-quality reasoning rationales for distillation datasets but often neglect the critical role of data quantity and quality. To address these challenges, we propose a Feedback-Driven Distillation (FDD) framework to enhance SLMs' mathematical reasoning capabilities. In the initialization stage, a distillation dataset is constructed by prompting LLMs to pair mathematical problems with corresponding reasoning rationales. We classify problems into easy and hard categories based on SLM performance. For easy problems, LLMs generate more complex variations, while for hard problems, new questions of similar complexity are synthesized. In addition, we propose a multi-round distillation paradigm to iteratively enrich the distillation datasets, thereby progressively improving the mathematical reasoning abilities of SLMs. Experimental results demonstrate that our method can make SLMs achieve SOTA mathematical reasoning performance.
Figures
Figures from the paper (1 more)
Forward citations
Cited by 1 Pith paper
-
VeriThinker: Learning to Verify Makes Reasoning Model Efficient
VeriThinker shows that fine-tuning a reasoning model only on a solution-verification task reduces chain-of-thought length on MATH500 and AIME by 20-45% while preserving or slightly improving accuracy.
Reference graph
Works this paper leans on
-
[5]
X. Zhu, B. Qi, K. Zhang, X. Long, Z. Lin, B. Zhou, Pad: Program-aided distillation can teach small models reasoning better than chain-of-thought fine-tuning (2023). arXiv:2305.13888. URL https://arxiv.org/abs/2305.13888v1
arXiv 2023
-
[6]
X. Zhu, J. Li, Y. Liu, C. Ma, W. Wang, Distilling mathematical reasoning capabilities into small language models, Neural Networks 179 (2024) 19 106594. doi:https://doi.org/10.1016/j.neunet.2024.106594. URL https://www.sciencedirect.com/science/article/pii/ S0893608024005185
arXiv 2024
-
[1]
C.-Y. Hsieh, C.-L. Li, C.-k. Yeh, H. Nakhost, Y. Fujii, A. Ratner, R. Krishna, C.-Y. Lee, T. Pfister, Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes, in: A. Rogers, J. Boyd-Graber, N. Okazaki (Eds.), Findings of the Association for Computational Linguistics: ACL 2023, Association for Computati...
-
[2]
N. Ho, L. Schmid, S.-Y. Yun, Large language models are reasoning teachers, in: A. Rogers, J. Boyd-Graber, N. Okazaki (Eds.), Pro- ceedings of the 61st Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), Association for Com- putational Linguistics, Toronto, Canada, 2023, pp. 14852–14882. doi: 10.18653/v1/2023.acl-long...
-
[3]
Y. Fu, H. Peng, L. Ou, A. Sabharwal, T. Khot, Specializing smaller lan- guage models towards multi-step reasoning, in: A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, J. Scarlett (Eds.), Proceedings of the 40th International Conference on Machine Learning, Vol. 202 of Proceed- ings of Machine Learning Research, PMLR, 2023, pp. 10421–10430. URL ...
work page 2023
-
[4]
K. Shridhar, A. Stolfo, M. Sachan, Distilling reasoning capabilities into smaller language models, in: A. Rogers, J. Boyd-Graber, N. Okazaki (Eds.), Findings of the Association for Computational Linguistics: ACL 2023, Association for Computational Linguistics, Toronto, Canada, 2023, pp. 7059–7073. doi:10.18653/v1/2023.findings-acl.441. URL https://aclanth...
-
[7]
Y. Wang, Y. Kordi, S. Mishra, A. Liu, N. A. Smith, D. Khashabi, H. Hajishirzi, Self-instruct: Aligning language models with self-generated instructions, in: A. Rogers, J. Boyd-Graber, N. Okazaki (Eds.), Pro- ceedings of the 61st Annual Meeting of the Association for Computa- tional Linguistics (Volume 1: Long Papers), Association for Compu- tational Lingu...
-
[8]
L. Yu, W. Jiang, H. Shi, J. YU, Z. Liu, Y. Zhang, J. Kwok, Z. Li, A. Weller, W. Liu, Metamath: Bootstrap your own mathematical ques- tions for large language models, in: The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=N8N0hgNDRt
work page 2024
Show all 25 references
-
[9]
C. Li, Z. Yuan, H. Yuan, G. Dong, K. Lu, J. Wu, C. Tan, X. Wang, C. Zhou, MuggleMath: Assessing the impact of query and response augmentation on math reasoning, in: L.-W. Ku, A. Martins, V. Srikumar (Eds.), Proceedings of the 62nd Annual Meeting of the Association for Computat...
2024 doi
-
[10]
C. Xu, Q. Sun, K. Zheng, X. Geng, P. Zhao, J. Feng, C. Tao, Q. Lin, D. Jiang, WizardLM: Empowering large pre-trained language models to follow complex instructions, in: The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=...
2024
-
[11]
Jiang, C
Y. Jiang, C. Chan, M. Chen, W. Wang, Lion: Adversarial distillation of proprietary large language models, in: H. Bouamor, J. Pino, K. Bali (Eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Association for Computational Linguistics,...
2023 doi
-
[12]
N. Lee, T. Wattanawong, S. Kim, K. Mangalam, S. Shen, G. Anu- manchipalli, M. Mahoney, K. Keutzer, A. Gholami, LLM2LLM: Boosting LLMs with novel iterative data enhancement, in: L.-W. Ku, A. Martins, V. Srikumar (Eds.), Findings of the Association for Computational Lin- guistic...
2024 doi
-
[13]
J. Ying, M. Lin, Y. Cao, W. Tang, B. Wang, Q. Sun, X. Huang, S. Yan, LLMs-as-instructors: Learning from errors toward automating model improvement, in: Y. Al-Onaizan, M. Bansal, Y.-N. Chen (Eds.), Find- ings of the Association for Computational Linguistics: EMNLP 2024, Associa...
2024
-
[14]
Cobbe, V
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, J. Schulman, Training verifiers to solve math word problems, CoRR abs/2110.14168 (2021). arXiv:2110.14168. URL https://arxiv.org/abs/2110.14168
2021 arXiv
-
[15]
Miao, C.-C
S.-y. Miao, C.-C. Liang, K.-Y. Su, A diverse corpus for evaluating and developing English math word problem solvers, in: D. Jurafsky, J. Chai, N. Schluter, J. Tetreault (Eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Association...
2020
-
[16]
Patel, S
A. Patel, S. Bhattamishra, N. Goyal, Are NLP models really able to solve simple math word problems?, in: K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty, Y. Zhou (Eds.), Proceedings of the 2021 Conference of the ...
2021 doi
-
[17]
S. Roy, D. Roth, Solving general arithmetic word problems, in: L. M` arquez, C. Callison-Burch, J. Su (Eds.), Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, As- sociation for Computational Linguistics, Lisbon, Portugal, 2015, pp. 1743–1...
2015 doi
-
[18]
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma, A. Webson, S. S. Gu, Z. Dai, M. Suz- gun, X. Chen, A. Chowdhery, A. Castro-Ros, M. Pellat, K. Robinson, D. Valter, S. Narang, G. Mishra, A. Yu, V. Zhao, Y. Huang, A. Dai, H. Yu,...
2024
-
[19]
Achiam, S
OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, R. Avila, I. Babuschkin, S. Balaji, V. Balcom, P. Baltescu, H. Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-Shapiro, C. Berner, L....
2024 arXiv
-
[20]
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, E. Chu, J. H. Clark, L. E. Shafey, Y. Huang, K. Meier-Hellstern, G. Mishra, E. Moreira, M. Omer- nick, K. Robinson, S. Ruder, Y. Tay, K. Xiao, Y. Xu, Y. Zhang, G. H. Ab...
2023 arXiv
-
[21]
Touvron, L
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hoss...
2023 arXiv
-
[22]
Rozi` ere, J
B. Rozi` ere, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan, Y. Adi, J. Liu, R. Sauvestre, T. Remez, J. Rapin, A. Kozhevnikov, I. Evtimov, J. Bitton, M. Bhatt, C. C. Ferrer, A. Grattafiori, W. Xiong, A. D´ efossez, J. Copet, F. Azhar, H. Touvron, L. Martin, N. Usunier,...
2024 arXiv
-
[23]
A. N. Lee, C. J. Hunter, N. Ruiz, Platypus: Quick, cheap, and powerful refinement of llms (2024). arXiv:2308.07317. URL https://arxiv.org/abs/2308.07317 24
2024 arXiv
-
[24]
H. Luo, Q. Sun, C. Xu, P. Zhao, J. Lou, C. Tao, X. Geng, Q. Lin, S. Chen, D. Zhang, Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct (2023). arXiv:2308.09583. URL https://arxiv.org/abs/2308.09583
2023 arXiv
-
[25]
Z. Gou, Z. Shao, Y. Gong, yelong shen, Y. Yang, M. Huang, N. Duan, W. Chen, ToRA: A tool-integrated reasoning agent for mathematical problem solving, in: The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=Ep0TtjVoap 25 A...
2024
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.