REVIEW 4 major objections 4 minor 167 references
Evaluation of LLMs for mathematical problem solving
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper compares GPT-4o, DeepSeek-V3, and Gemini-2.0 on three mathematics datasets and claims GPT-4o is the most stable model, even though its own accuracy tables place Gemini-2.0 first on every dataset.
desk verdict The multi-dimensional evaluation idea and MIT dataset are useful, but the paper's central claim that GPT-4o is most stable is contradicted by its own tables; the work needs major corrections before peer review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Structured Chain-of-Thought (SCoT) framework is the central instrument: it decomposes each model's solution into steps and scores five dimensions—final-answer accuracy, step completeness, step validity, intermediate calculation accuracy, and problem comprehension. The framework also includes an automated grading model built on DeepSeek-V3 for the university-level dataset, a 1-to-5 scoring rubric for non-perfect solutions, and a ten-category manual error taxonomy. This machinery converts raw LLM outputs into comparable quality scores and produces the per-subject and per-difficulty comparisons that the paper's conclusions rest on.
What would settle it
Re-evaluate the three datasets across repeated runs with fixed prompts, record accuracy and score distributions per model, and have human graders independently score a random subset of the university problems; if GPT-4o's accuracy remains the lowest on all three datasets and its run-to-run variance is no smaller than Gemini-2.0's, the 'most stable and consistent' claim is falsified.
Extended reading notes
Core claim
The paper's central assertion, stated in the abstract and in Section 4.5, is that GPT-4o is the most stable and consistent performer across GSM8K, MATH500, and the University dataset, with especially strong results on high-level open-courseware problems. It also claims DeepSeek-V3 is competitive in well-structured computational and optimisation domains but loses accuracy in statistical inference, and that Gemini-2.0 has strong linguistic understanding but weak multi-step reasoning and symbolic logic. The intended discovery is thus a differentiated capability profile per model, produced by the SCoT evaluation rather than by final-answer accuracy alone. The reported accuracy figures in the same paper tell a different rank order on the accuracy dimension: GPT-4o scores lower than the other two on every dataset, and Gemini-2.0 has the highest accuracy on all three. In the paper's own terms, the evidence supports a model-specific strengths-and-weaknesses profile, but the 'most stable' label is not what the accuracy tables show.
Load-bearing premise
The load-bearing premise is that the SCoT scores, including those produced by an automated DeepSeek-V3 grader for the university dataset, are fair and consistent enough to rank the models, because the reported accuracy tables alone do not support calling GPT-4o the most stable performer.
Editorial extensions
If this is right
- If the stability claim were taken at face value, educators choosing a general-purpose model for mixed mathematics courses would prefer GPT-4o despite its lower reported accuracy.
- If DeepSeek-V3's optimisation strengths are real, its best classroom use is computational and optimisation problem sets rather than statistical inference.
- If Gemini-2.0's documented weakness in multi-step reasoning and symbolic logic is real, it needs structured step-by-step prompts or verification before use on proof-heavy material.
- Prompt scaffolding changes outcomes for all three models, so any single-prompt evaluation understates what the models can do when guided.
- The discrepancy between the abstract and the accuracy tables means that an evaluation summary built on the SCoT dimensions must report per-dimension scores explicitly before a reader can judge a stability claim.
Reading between the lines
- The paper leaves implicit that 'stability' and 'accuracy' are different quantities; if stability means low variance across repeated runs, the claim and the tables could both be true, but the paper never reports the repeated-run variance that would verify it.
- Because the automated grader for the university dataset is itself DeepSeek-V3, one of the models under comparison, the SCoT scores for that dataset could carry an undetected grader bias; a fairer design would use a hold-out grader or human marking.
- A testable extension: rerun the same three datasets with several temperatures and seeds, record score distributions rather than point accuracies, and compare variance across models—this would turn 'stable and consistent' into a measured quantity.
- The roughly 50-percent accuracies on GSM8K echo a broader pattern that LLM arithmetic reasoning degrades sharply on multi-step word problems, so the practical educational lesson may be that these models need verification layers rather than raw generation in tutoring tools.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper compares GPT-4o, DeepSeek-V3, and Gemini-2.0 on three mathematics datasets (GSM8K, MATH500, and a University/MIT OCW dataset) using a Structured Chain-of-Thought (SCoT) framework with five evaluation dimensions. It reports accuracy and qualitative error analyses and concludes that GPT-4o is the most stable and consistent model across all datasets, with Gemini strong in structured problems, DeepSeek strong in computation but weak in inference, and GPT-4o lacking explanation precision. The paper also introduces an automated grading component and reports accuracy tables for each dataset.
Significance. If the evaluation were sound, the cross-model comparison on datasets of varying difficulty, including a university-level dataset, would provide useful evidence for educators selecting LLMs for mathematics tutoring and assessment. The authors make code and data publicly available and apply a multi-dimensional evaluation framework, which are positive features. However, the central claim of the paper is directly contradicted by its own accuracy tables, and the University-dataset grading procedure uses one of the compared models as the automated grader without validation. These issues undermine the reliability of the reported rankings and the paper's headline conclusion.
major comments (4)
- [Abstract and §4.5] The claim that "GPT-4o is the most stable and consistent in performance across all the datasets" is contradicted by the paper's own accuracy tables. Table 5 shows GPT-4o at 50.0% vs Gemini 53.9% on GSM8K; Table 7 shows GPT-4o at 36.8% vs Gemini 54.7% on MATH500; Table 10 shows GPT-4o at 35.6% vs Gemini 54.4% on the University dataset; and Table 17 repeats these values. GPT-4o has the lowest accuracy of the three models on every dataset, and Gemini has the highest. No run-to-run consistency metric or alternative aggregation is reported that would support placing GPT-4o first on stability.
- [§4.5] The summary section contains numerical inconsistencies that further undermine the central claim. It states that "the accuracy of GPT-4o, DeepSeek, and Gemini for the GSM8K dataset is 54.7%, 54.4%, and 53.9%", which matches neither Table 5 (50.0%, 42.6%, 53.9%) nor Table 17. It also states that "The overall accuracy for University problems is 29%", while Table 17 reports 35.6% for GPT-4o, 42.8% for DeepSeek, and 54.4% for Gemini. Later in the same section, GPT-4o is said to have "an average correctness rate of 82%" on University questions, which is inconsistent with the 35.6% in Table 10. These contradictions make the reported results and the headline conclusion unreliable.
- [§4.3] The evaluation of GPT-4o on the University dataset uses "an automated grading model based on Deepseek-V3 that evaluates each solution step." This means one of the models under comparison acts as the grader for another model's outputs. No validation of this grader against human judgments, no inter-rater reliability, and no agreement statistics are reported. Since the University dataset is the only source for the GPT-4o stability claim, the rankings on that dataset rest on an unverified and potentially biased grading process. This is a load-bearing methodological flaw, not a minor concern.
- [§3.4 and §4] The paper lists "Consistency" as a core evaluation dimension and uses it in the headline claim, but it never reports a quantitative consistency metric. No variance across repeated runs, no standard deviation, and no agreement measure is given for any model or dataset. The only support for "stable and consistent" appears to be the qualitative statements in §4.5. Without an operationalized consistency measure, the central claim is not empirically supported.
minor comments (4)
- [Throughout] There are numerous typos and inconsistencies, including "DeekSeek" (§4.3), "vaggue" (Table 18), "Univesrity" (Table 19 title), and inconsistent use of "MATH 500" vs "MATH500." The manuscript would benefit from a careful proofreading pass.
- [Introduction and Table 2] The introduction refers to the "University of New South Sales" dataset, while Table 1 and Table 2 describe "MIT Open Courseware" as the source. This naming inconsistency should be resolved, and the relationship between the UNSW and MIT OCW datasets should be clarified.
- [Table 9] The MATH500 subject-level and difficulty-level accuracy numbers for DeepSeek and GPT-4o in Table 9 are internally inconsistent: for example, DeepSeek's listed subject accuracies (e.g., Prealgebra 52%, Algebra 40%) are difficult to reconcile with its overall accuracy of 51.1% given the reported difficulty distribution. The authors should provide per-cell counts or clarifying details.
- [§4.3] The sentence "29% accuracy is not much surprising due to the rigorous criterion. Only a full mark answer can be considered correct for proof problems" is unclear, as the paper reports accuracy values far above 29% in Table 10 and does not explain how the 29% figure was derived. This should be rewritten or removed.
Circularity Check
No circularity found: the paper reports empirical benchmark measurements, and its conclusions are not derived from fitted parameters or self-referential definitions.
full rationale
This manuscript is an empirical evaluation of three LLMs on three mathematics datasets, not a derivation of a predicted quantity from fitted inputs. The accuracy values in Tables 5, 7, 10, and 17 are external measurements against standard answer keys. No parameter is fitted to a subset of data and then renamed as a prediction; no equation defines the claimed outcome in terms of the measurement itself; and no load-bearing uniqueness theorem or ansatz is imported through self-citation. The paper's own results contradict its abstract claim that GPT-4o is 'the most stable and consistent in performance across all the datasets,' since Gemini-2.0 has the highest accuracy on every dataset (53.9%, 54.7%, 54.4%) while GPT-4o has the lowest (50.0%, 36.8%, 35.6%); however, that is an internal-consistency or correctness defect, not circular reasoning. Similarly, Section 4.3's statement that 'an automated grading model based on Deepseek-V3' evaluates GPT-4o solutions creates a potential grader-bias concern, but it does not make the reported University scores equivalent to the paper's inputs by construction. Because no step in the paper's chain of measurement reduces to its own premises, the circularity score is 0.
Assumptions & free parameters
free parameters (3)
- Difficulty level word-length thresholds =
Level 1: 20-25 words; Level 2: 25-30; Level 3: 30-35; Level 4: 35-40; Level 5: >40
- University score rubric =
1-5 integer scale (Table 11)
- Answer-matching tolerance rules =
Equivalent expressions accepted; floating-point margin not specified
assumptions (3)
- domain assumption API outputs from GPT-4o, DeepSeek-V3, and Gemini-2.0 are representative of those models' typical behavior
- domain assumption The SCoT framework is a valid instrument for measuring mathematical reasoning quality
- domain assumption The manual error categorization is consistent across models and raters
Cite this review
Pith. "Pith review of Evaluation of LLMs for mathematical problem solving." pith.science (2026). https://pith.science/paper/A5ZISYDN
@misc{pith2026250600309,
author = {Pith},
title = {Pith review of: Evaluation of LLMs for mathematical problem solving},
year = {2026},
howpublished = {\url{https://pith.science/paper/A5ZISYDN}},
note = {Machine review of arXiv:2506.00309}
}
read the original abstract
Large Language Models (LLMs) have shown impressive performance on a range of educational tasks, but are still understudied for their potential to solve mathematical problems. In this study, we compare three prominent LLMs, including GPT-4o, DeepSeek-V3, and Gemini-2.0, on three mathematics datasets of varying complexities (GSM8K, MATH500, and MIT Open Courseware datasets). We take a five-dimensional approach based on the Structured Chain-of-Thought (SCoT) framework to assess final answer correctness, step completeness, step validity, intermediate calculation accuracy, and problem comprehension. The results show that GPT-4o is the most stable and consistent in performance across all the datasets, but particularly it performs outstandingly in high-level questions of the MIT Open Courseware dataset. DeepSeek-V3 is competitively strong in well-structured domains such as optimisation, but suffers from fluctuations in accuracy in statistical inference tasks. Gemini-2.0 shows strong linguistic understanding and clarity in well-structured problems but performs poorly in multi-step reasoning and symbolic logic. Our error analysis reveals particular deficits in each model: GPT-4o is at times lacking in sufficient explanation or precision; DeepSeek-V3 leaves out intermediate steps; and Gemini-2.0 is less flexible in mathematical reasoning in higher dimensions.
Figures
Reference graph
Works this paper leans on
-
[1]
Liljedahl and M
P. Liljedahl and M. Santos-Trigo, Mathematical problem solving . Springer, 2019
2019
-
[2]
Students’ difficulties in mathemat- ics problem-solving: What do they say?
T. Tambychik and T. S. M. Meerah, “Students’ difficulties in mathemat- ics problem-solving: What do they say?” Procedia-Social and Behav- ioral Sciences, vol. 8, pp. 142–151, 2010
2010
-
[3]
Abstraction in mathematics and mathe- matics learning
M. Mitchelmore and P. White, “Abstraction in mathematics and mathe- matics learning.” International Group for the Psychology of Mathemat- ics Education, 2004
2004
-
[4]
J. T. Kinard and A. Kozulin, Rigorous mathematical thinking: Concep- tual formation in the mathematics classroom . Cambridge University Press, 2008
2008
-
[5]
Mathematics: con- ceptions, learning and teaching,
B. Yushau, M. Bokhari, A. Mji, and D. Wessels, “Mathematics: con- ceptions, learning and teaching,” King Fahd University of Petroleum & Minerals, Department of Mathematical Sciences: Technical Report Se- ries: TR, vol. 322, 2004
2004
-
[6]
The nature, e ffects, and relief of mathematics anxiety,
R. Hembree, “The nature, e ffects, and relief of mathematics anxiety,” Journal for research in mathematics education , vol. 21, no. 1, pp. 33– 46, 1990
1990
-
[7]
Development of ab- stract mathematical reasoning: the case of algebra,
A. Susac, A. Bubic, A. Vrbanc, and M. Planinic, “Development of ab- stract mathematical reasoning: the case of algebra,” Frontiers in Human Neuroscience, vol. 8, p. 679, 2014
2014
-
[8]
The logical thinking ability: Math- ematical disposition and self-regulated learning,
J. Ab, G. Margono, and W. Rahayu, “The logical thinking ability: Math- ematical disposition and self-regulated learning,” in Journal of Physics: Conference Series, vol. 1155, no. 1. IOP Publishing, 2019, p. 012092
2019
Show all 167 references
-
[9]
Developing critical thinking skills of students in mathematics learning,
F. Firdaus, I. Kailani, M. N. B. Bakar, and B. Bakry, “Developing critical thinking skills of students in mathematics learning,” Journal of Educa- tion and Learning (EduLearn), vol. 9, no. 3, pp. 226–236, 2015
2015
-
[10]
Critical thinking: Essence for teaching mathematics and mathematics problem solving skills,
E. E. Peter, “Critical thinking: Essence for teaching mathematics and mathematics problem solving skills,” African Journal of Mathematics and Computer Science Research, vol. 5, no. 3, pp. 39–43, 2012
2012
-
[11]
Artificial intelligence in mathemat- ics education: A systematic literature review,
M. Z. bin Mohamed, R. Hidayat, N. N. binti Suhaizi, M. K. H. bin Mah- mud, S. N. binti Baharuddin et al., “Artificial intelligence in mathemat- ics education: A systematic literature review,” International Electronic Journal of Mathematics Education, vol. 17, no. 3, p. em0694, 2022
2022
-
[12]
The importance of educational technology in teaching,
S. Lazar, “The importance of educational technology in teaching,” In- ternational Journal of Cognitive Research in Science, Engineering and Education, vol. 3, no. 1, pp. 111–114, 2015
2015
-
[13]
Benefits and limitations of the artificial with respect to the traditional learning of mathematics,
M. G. V oskoglou and A.-B. M. Salem, “Benefits and limitations of the artificial with respect to the traditional learning of mathematics,” Math- ematics, vol. 8, no. 4, p. 611, 2020
2020
-
[14]
The use of interactive learning technologies in math classes,
N. V . Telegina, E. G. Galimova, and S. G. Dobrotvorskaya, “The use of interactive learning technologies in math classes,” Modern journal of language teaching methods, vol. 7, no. 4, pp. 106–113, 2017
2017
-
[15]
Going the distance with on- line education,
J. Larreamendy-Joerns and G. Leinhardt, “Going the distance with on- line education,” Review of educational research, vol. 76, no. 4, pp. 567– 605, 2006
2006
-
[16]
Natural language processing in the era of large language models,
A. Zubiaga, “Natural language processing in the era of large language models,” p. 1350306, 2024
2024
-
[17]
An introduction to key technology in artifi- cial intelligence and big data driven e-learning and e-education,
P. Gao, J. Li, and S. Liu, “An introduction to key technology in artifi- cial intelligence and big data driven e-learning and e-education,”Mobile Networks and Applications, vol. 26, no. 5, pp. 2123–2126, 2021
2021
-
[18]
Learning mathematics with large language models: A comparative study with computer algebra systems and other tools,
N. Matzakos, S. Doukakis, and M. Moundridou, “Learning mathematics with large language models: A comparative study with computer algebra systems and other tools,” International Journal of Emerging Technolo- gies in Learning (iJET), vol. 18, no. 20, pp. 51–71, 2023
2023
-
[19]
Johnsen, Large language models (LLMs)
M. Johnsen, Large language models (LLMs). Maria Johnsen, 2024
2024
-
[20]
Talking about large language models,
M. Shanahan, “Talking about large language models,” Communications of the ACM, vol. 67, no. 2, pp. 68–79, 2024
2024
-
[21]
A survey on large language mod- els: Applications, challenges, limitations, and practical usage,
M. U. Hadi, R. Qureshi, A. Shah, M. Irfan, A. Zafar, M. B. Shaikh, N. Akhtar, J. Wu, S. Mirjalili et al., “A survey on large language mod- els: Applications, challenges, limitations, and practical usage,”Authorea Preprints, vol. 3, 2023
2023
-
[22]
A survey of large language models,
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y . Hou, Y . Min, B. Zhang, J. Zhang, Z. Dong et al., “A survey of large language models,” arXiv preprint arXiv:2303.18223, vol. 1, no. 2, 2023
2023 arXiv
-
[23]
Challenges and applications of large language models,
J. Kaddour, J. Harris, M. Mozes, H. Bradley, R. Raileanu, and R. McHardy, “Challenges and applications of large language models,” arXiv preprint arXiv:2307.10169, 2023
2023 arXiv
-
[24]
The evolution of human-computer interaction: From com- mand lines to conversational interfaces powered by large language mod- els,
V . Vemuri, “The evolution of human-computer interaction: From com- mand lines to conversational interfaces powered by large language mod- els,” J Artif Intell Mach Learn & Data Sci 2024, vol. 2, no. 1, pp. 2257– 2266
2024
-
[25]
Enhancing human-computer in- teraction through ai: A study on chatgpt in educational environments,
D. K. Kothari and O. N. N. Fernando, “Enhancing human-computer in- teraction through ai: A study on chatgpt in educational environments,” in 2024 IEEE Conference on Artificial Intelligence (CAI). IEEE, 2024, pp. 500–503
2024
-
[26]
Large language models: a com- prehensive survey of its applications, challenges, limitations, and future prospects,
M. U. Hadi, R. Qureshi, A. Shah, M. Irfan, A. Zafar, M. B. Shaikh, N. Akhtar, J. Wu, S. Mirjalili et al., “Large language models: a com- prehensive survey of its applications, challenges, limitations, and future prospects,” Authorea Preprints, vol. 1, pp. 1–26, 2023
2023
-
[27]
A survey on enhancing reinforce- ment learning in complex environments: Insights from human and llm feedback,
A. R. Laleh and M. N. Ahmadabadi, “A survey on enhancing reinforce- ment learning in complex environments: Insights from human and llm feedback,” arXiv preprint arXiv:2411.13410, 2024
2024 arXiv
-
[28]
A multifaceted vision of the human-ai collaboration: a comprehensive review,
M. Puerta-Beldarrain, O. Gómez-Carmona, R. Sánchez-Corcuera, D. Casado-Mansilla, D. López-de Ipiña, and L. Chen, “A multifaceted vision of the human-ai collaboration: a comprehensive review,” IEEE Access, 2025
2025
-
[29]
Practical and ethical challenges of large lan- guage models in education: A systematic scoping review,
L. Yan, L. Sha, L. Zhao, Y . Li, R. Martinez-Maldonado, G. Chen, X. Li, Y . Jin, and D. Gaševi´c, “Practical and ethical challenges of large lan- guage models in education: A systematic scoping review,” British Jour- nal of Educational Technology, vol. 55, no. 1, pp. 90–112, 2024
2024
-
[30]
Challenges and future directions of big data and artificial intelligence in education,
H. Luan, P. Geczy, H. Lai, J. Gobert, S. J. Yang, H. Ogata, J. Baltes, R. Guerra, P. Li, and C.-C. Tsai, “Challenges and future directions of big data and artificial intelligence in education,”Frontiers in psychology, vol. 11, p. 580820, 2020
2020
-
[31]
Exploring the possibility and implementation of ai- supported online math coaching,
K. Karlström, “Exploring the possibility and implementation of ai- supported online math coaching,” 2024
2024
-
[32]
Artificial intelligence in mathemat- ics education: The good, the bad, and the ugly,
O. A. Opesemowo and M. Ndlovu, “Artificial intelligence in mathemat- ics education: The good, the bad, and the ugly,” Journal of Pedagogical Research, vol. 8, no. 3, pp. 333–346, 2024
2024
-
[33]
Information analysis of advanced mathematics education- adaptive algorithm based on big data,
J. Tan, “Information analysis of advanced mathematics education- adaptive algorithm based on big data,” Mathematical Problems in En- gineering, vol. 2022, no. 1, p. 7796681, 2022
2022
-
[34]
Deep learning,
Y . LeCun, Y . Bengio, and G. Hinton, “Deep learning,”nature, vol. 521, no. 7553, pp. 436–444, 2015
2015
-
[35]
Development and application of artificial neu- ral network,
Y .-c. Wu and J.-w. Feng, “Development and application of artificial neu- ral network,” Wireless Personal Communications, vol. 102, pp. 1645– 1656, 2018
2018
-
[36]
Large language models (llms): survey, technical frameworks, and future challenges,
P. Kumar, “Large language models (llms): survey, technical frameworks, and future challenges,” Artificial Intelligence Review, vol. 57, no. 10, p. 260, 2024. 19
2024
-
[37]
When large language models meet personal- ization: Perspectives of challenges and opportunities,
J. Chen, Z. Liu, X. Huang, C. Wu, Q. Liu, G. Jiang, Y . Pu, Y . Lei, X. Chen, X. Wang et al., “When large language models meet personal- ization: Perspectives of challenges and opportunities,” World Wide Web, vol. 27, no. 4, p. 42, 2024
2024
-
[38]
A survey on large lan- guage models for code generation,
J. Jiang, F. Wang, J. Shen, S. Kim, and S. Kim, “A survey on large lan- guage models for code generation,” arXiv preprint arXiv:2406.00515 , 2024
2024 arXiv
-
[39]
Datasets for large language models: A comprehensive survey,
Y . Liu, J. Cao, C. Liu, K. Ding, and L. Jin, “Datasets for large language models: A comprehensive survey,” arXiv preprint arXiv:2402.18041 , 2024
2024 arXiv
-
[40]
Harnessing large language mod- els’ zero-shot and few-shot learning capabilities for regulatory research,
H. Meshkin, J. Zirkle, G. Arabidarrehdor, A. Chaturbedi, S. Chakravar- tula, J. Mann, B. Thrasher, and Z. Li, “Harnessing large language mod- els’ zero-shot and few-shot learning capabilities for regulatory research,” Briefings in Bioinformatics, vol. 25, no. 5, p. bbae354, 2024
2024
-
[41]
Alto, Modern Generative AI with ChatGPT and OpenAI Models: Leverage the capabilities of OpenAI’s LLM for productivity and inno- vation with GPT3 and GPT4
V . Alto, Modern Generative AI with ChatGPT and OpenAI Models: Leverage the capabilities of OpenAI’s LLM for productivity and inno- vation with GPT3 and GPT4. Packt Publishing Ltd, 2023
2023
-
[42]
Google gemini as a next generation ai educational tool: a review of emerging educational technology,
M. Imran and N. Almusharraf, “Google gemini as a next generation ai educational tool: a review of emerging educational technology,” Smart Learning Environments, vol. 11, no. 1, p. 22, 2024
2024
-
[43]
A survey of deepseek models,
F. Neha and D. Bhati, “A survey of deepseek models,” Authorea Preprints, 2025
2025
-
[44]
Comparative analysis based on deepseek, chatgpt, and google gem- ini: Features, techniques, performance, future prospects,
A. Rahman, S. H. Mahir, M. T. A. Tashrif, A. A. Aishi, M. A. Karim, D. Kundu, T. Debnath, M. A. A. Moududi, and M. Eidmum, “Comparative analysis based on deepseek, chatgpt, and google gem- ini: Features, techniques, performance, future prospects,”arXiv preprint arXiv:2503.04783, 2025
2025 arXiv
-
[45]
Bias and fairness in large language models: A survey,
I. O. Gallegos, R. A. Rossi, J. Barrow, M. M. Tanjim, S. Kim, F. Der- noncourt, T. Yu, R. Zhang, and N. K. Ahmed, “Bias and fairness in large language models: A survey,” Computational Linguistics, vol. 50, no. 3, pp. 1097–1179, 2024
2024
-
[46]
Large language models (llm) in compu- tational social science: prospects, current state, and challenges,
S. Thapa, S. Shiwakoti, S. B. Shah, S. Adhikari, H. Veeramani, M. Nasim, and U. Naseem, “Large language models (llm) in compu- tational social science: prospects, current state, and challenges,” Social Network Analysis and Mining, vol. 15, no. 1, pp. 1–30, 2025
2025
-
[47]
Foundation and large language models: fundamentals, challenges, op- portunities, and social impacts,
D. Myers, R. Mohawesh, V . I. Chellaboina, A. L. Sathvik, P. Venkatesh, Y .-H. Ho, H. Henshaw, M. Alhawawreh, D. Berdik, and Y . Jararweh, “Foundation and large language models: fundamentals, challenges, op- portunities, and social impacts,” Cluster Computing, vol. 27, no. 1, ...
2024
-
[48]
Llm challenges and solutions,
U. Kamath, K. Keenan, G. Somers, and S. Sorenson, “Llm challenges and solutions,” inLarge Language Models: A Deep Dive: Bridging The- ory and Practice. Springer, 2024, pp. 219–274
2024
-
[49]
A step-by-step solution methodology for mathematical expressions,
S. Hosseinpour, M. M. R. Alavi Milani, and H. Pehlivan, “A step-by-step solution methodology for mathematical expressions,”Symmetry, vol. 10, no. 7, p. 285, 2018
2018
-
[50]
Solving mathematical problems using large language models: A survey,
H. Lai, B. Wang, J. Liu, F. He, C. Zhang, H. Liu, and H. Chen, “Solving mathematical problems using large language models: A survey,” Avail- able at SSRN 5002356
-
[51]
Testing gpt-4 with wolfram alpha and code interpreter plug-ins on math and science problems,
E. Davis and S. Aaronson, “Testing gpt-4 with wolfram alpha and code interpreter plug-ins on math and science problems,”arXiv preprint arXiv:2308.05713, 2023
2023 arXiv
-
[52]
Assessing the performance of ai-generated code: A case study on github copilot,
S. Li, Y . Cheng, J. Chen, J. Xuan, S. He, and W. Shang, “Assessing the performance of ai-generated code: A case study on github copilot,” in 2024 IEEE 35th International Symposium on Software Reliability Engi- neering (ISSRE). IEEE, 2024, pp. 216–227
2024
-
[53]
Exploring large language models for code explanation,
P. Bhattacharya, M. Chakraborty, K. N. Palepu, V . Pandey, I. Dindorkar, R. Rajpurohit, and R. Gupta, “Exploring large language models for code explanation,” arXiv preprint arXiv:2310.16673, 2023
2023 arXiv
-
[54]
Automatic programming: Large language models and beyond,
M. R. Lyu, B. Ray, A. Roychoudhury, S. H. Tan, and P. Thongtanunam, “Automatic programming: Large language models and beyond,” ACM Transactions on Software Engineering and Methodology, 2024
2024
-
[55]
Intelligent computing: the latest advances, challenges, and future,
S. Zhu, T. Yu, T. Xu, H. Chen, S. Dustdar, S. Gigan, D. Gunduz, E. Hos- sain, Y . Jin, F. Lin et al., “Intelligent computing: the latest advances, challenges, and future,” Intelligent Computing, vol. 2, p. 0006, 2023
2023
-
[56]
Bridging machine learning and logical reasoning by abductive learning,
W.-Z. Dai, Q. Xu, Y . Yu, and Z.-H. Zhou, “Bridging machine learning and logical reasoning by abductive learning,”Advances in Neural Infor- mation Processing Systems, vol. 32, 2019
2019
-
[57]
Mathbench: Evaluating the theory and application proficiency of llms with a hierarchical mathematics benchmark,
H. Liu, Z. Zheng, Y . Qiao, H. Duan, Z. Fei, F. Zhou, W. Zhang, S. Zhang, D. Lin, and K. Chen, “Mathbench: Evaluating the theory and application proficiency of llms with a hierarchical mathematics benchmark,” arXiv preprint arXiv:2405.12209, 2024
2024 arXiv
-
[58]
Metacog- nitive capabilities of llms: An exploration in mathematical problem solv- ing,
A. Didolkar, A. Goyal, N. R. Ke, S. Guo, M. Valko, T. Lillicrap, D. Jimenez Rezende, Y . Bengio, M. C. Mozer, and S. Arora, “Metacog- nitive capabilities of llms: An exploration in mathematical problem solv- ing,” Advances in Neural Information Processing Systems , vol. 37, pp...
2024
-
[59]
A systematic study and comprehensive evaluation of chat- gpt on benchmark datasets,
M. T. R. Laskar, M. S. Bari, M. Rahman, M. A. H. Bhuiyan, S. Joty, and J. X. Huang, “A systematic study and comprehensive evaluation of chat- gpt on benchmark datasets,” arXiv preprint arXiv:2305.18486, 2023
2023 arXiv
-
[60]
A survey on evaluation of large language mod- els,
Y . Chang, X. Wang, J. Wang, Y . Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y . Wanget al., “A survey on evaluation of large language mod- els,” ACM transactions on intelligent systems and technology , vol. 15, no. 3, pp. 1–45, 2024
2024
-
[61]
Surampudi, Big Data Meets LLMs: A New Era of Incident Monitor- ing
Y . Surampudi, Big Data Meets LLMs: A New Era of Incident Monitor- ing. Libertatem Media Private Limited, 2024
2024
-
[62]
E ffect of chat gpt on the digitized learning process of university students,
A. G. R. Castillo, G. S. Silva, J. F. Arocutipa, H. Q. Berrios, M. A. M. Rodriguez, G. Y . Reyes, H. R. P. Lopez, R. M. V . Teves, H. V . H. Rivera, and J. L. Arias-Gonzáles, “E ffect of chat gpt on the digitized learning process of university students,” Journal of Namibian St...
2023
-
[63]
The e ffects of over-reliance on ai dialogue systems on students’ cognitive abilities: a systematic review,
C. Zhai, S. Wibowo, and L. D. Li, “The e ffects of over-reliance on ai dialogue systems on students’ cognitive abilities: a systematic review,” Smart Learning Environments, vol. 11, no. 1, p. 28, 2024
2024
-
[64]
Chatgpt: The end of online exam in- tegrity?
T. Susnjak and T. R. McIntosh, “Chatgpt: The end of online exam in- tegrity?” Education Sciences, vol. 14, no. 6, p. 656, 2024
2024
-
[65]
Ethical principles for artificial intelligence in education,
A. Nguyen, H. N. Ngo, Y . Hong, B. Dang, and B.-P. T. Nguyen, “Ethical principles for artificial intelligence in education,” Education and infor- mation technologies, vol. 28, no. 4, pp. 4221–4241, 2023
2023
-
[66]
Formal mathematical reasoning: A new frontier in ai,
K. Yang, G. Poesia, J. He, W. Li, K. Lauter, S. Chaudhuri, and D. Song, “Formal mathematical reasoning: A new frontier in ai,” arXiv preprint arXiv:2412.16075, 2024
2024 arXiv
-
[67]
Generative ai and future education: a review, theoretical validation, and authors’ perspective on challenges and solutions,
W. K. Monib, A. Qazi, R. A. Apong, M. T. Azizan, L. De Silva, and H. Yassin, “Generative ai and future education: a review, theoretical validation, and authors’ perspective on challenges and solutions,” PeerJ Computer Science, vol. 10, p. e2105, 2024
2024
-
[68]
The impact of artificial intelligence on the evolution of digital education: A compara- tive study of openai text generation tools including chatgpt, bing chat, bard, and ernie,
N. Y . Motlagh, M. Khajavi, A. Sharifi, and M. Ahmadi, “The impact of artificial intelligence on the evolution of digital education: A compara- tive study of openai text generation tools including chatgpt, bing chat, bard, and ernie,” arXiv preprint arXiv:2309.02029, 2023
2023 arXiv
-
[69]
Machine learning and deep learning,
C. Janiesch, P. Zschech, and K. Heinrich, “Machine learning and deep learning,” Electronic markets, vol. 31, no. 3, pp. 685–695, 2021
2021
-
[70]
Deep learning,
N. Rusk, “Deep learning,” Nature Methods, vol. 13, no. 1, pp. 35–35, 2016
2016
-
[71]
Scaling laws of data- driven machine learning models: A survey and taxonomy,
C. Wang, Y . Wang, Z. Li, Q. Jia, and W. Liu, “Scaling laws of data- driven machine learning models: A survey and taxonomy,” Authorea Preprints, 2024
2024
-
[72]
Deep learning & its applications,
N. P. R. Thota and A. Lakshmanarao, “Deep learning & its applications,” INDEX, VOLUME I, vol. 43, p. 81
-
[73]
Deep learning,
X. Hao, G. Zhang, and S. Ma, “Deep learning,” International Journal of Semantic Computing, vol. 10, no. 03, pp. 417–439, 2016
2016
-
[74]
Mehrotra, Basics of artificial intelligence & machine learning
D. Mehrotra, Basics of artificial intelligence & machine learning. No- tion Press, 2019
2019
-
[75]
A study of challenges and limitations to ap- plying machine learning to highly unstructured data,
S. Singh and S. Hooda, “A study of challenges and limitations to ap- plying machine learning to highly unstructured data,” in 2023 7th In- ternational Conference On Computing, Communication, Control And Automation (ICCUBEA). IEEE, 2023, pp. 1–6
2023
-
[76]
Supervised learning,
B. Liu, “Supervised learning,” in Web Data Mining: Exploring Hyper- links, Contents, and Usage Data. Springer, 2011, pp. 63–132
2011
-
[77]
Self- attention and transformers: Driving the evolution of large language mod- els,
Q. Luo, W. Zeng, M. Chen, G. Peng, X. Yuan, and Q. Yin, “Self- attention and transformers: Driving the evolution of large language mod- els,” in 2023 IEEE 6th International Conference on Electronic Informa- tion and Communication Technology (ICEICT). IEEE, 2023, pp. 401– 405
2023
-
[78]
A review on large language models: Architectures, applications, taxonomies, open issues and challenges,
M. A. K. Raiaan, M. S. H. Mukta, K. Fatema, N. M. Fahad, S. Sakib, M. M. J. Mim, J. Ahmad, M. E. Ali, and S. Azam, “A review on large language models: Architectures, applications, taxonomies, open issues and challenges,” IEEE access, vol. 12, pp. 26 839–26 874, 2024
2024
-
[79]
Large language models, fine tuning, code generation,
L. Z. Juhász, “Large language models, fine tuning, code generation,” in 2024 IEEE 7th International Conference and Workshop Óbuda on Electrical and Power Engineering (CANDO-EPE) . IEEE, 2024, pp. 000 191–000 200. 20
2024
-
[80]
Machine learning and deep learning for big data analytics: A review of methods and applications,
N. L. Rane, M. Paramesha, S. P. Choudhary, and J. Rane, “Machine learning and deep learning for big data analytics: A review of methods and applications,” Partners Universal International Innovation Journal, vol. 2, no. 3, pp. 172–197, 2024
2024
-
[81]
Chatgpt and open-ai models: A preliminary review,
K. I. Roumeliotis and N. D. Tselikas, “Chatgpt and open-ai models: A preliminary review,”Future Internet, vol. 15, no. 6, p. 192, 2023
2023
-
[82]
Gpt (generative pre-trained transformer)–a comprehensive re- view on enabling technologies, potential applications, emerging chal- lenges, and future directions,
G. Yenduri, M. Ramalingam, G. C. Selvi, Y . Supriya, G. Srivastava, P. K. R. Maddikunta, G. D. Raj, R. H. Jhaveri, B. Prabadevi, W. Wang et al. , “Gpt (generative pre-trained transformer)–a comprehensive re- view on enabling technologies, potential applications, emerging chal-...
2024
-
[83]
Research on text classification model based on self-attention mechanism and multi-neural network
X. Wang, Y . Chen, W. Liu, and W. Tai, “Research on text classification model based on self-attention mechanism and multi-neural network.” in ICBASE, 2022, pp. 244–256
2022
-
[84]
Chatgpt–technical research model, capability analysis, and application prospects,
T. Wang and Q. Zhu, “Chatgpt–technical research model, capability analysis, and application prospects,” in 2024 IEEE 7th Advanced In- formation Technology, Electronic and Automation Control Conference (IAEAC), vol. 7. IEEE, 2024, pp. 787–796
2024
-
[85]
A sur- vey of reinforcement learning from human feedback,
T. Kaufmann, P. Weng, V . Bengs, and E. Hüllermeier, “A sur- vey of reinforcement learning from human feedback,” arXiv preprint arXiv:2312.14925, vol. 10, 2023
2023
-
[86]
Summary of chatgpt-related research and perspective towards the future of large language models,
Y . Liu, T. Han, S. Ma, J. Zhang, Y . Yang, J. Tian, H. He, A. Li, M. He, Z. Liu et al. , “Summary of chatgpt-related research and perspective towards the future of large language models,” Meta-radiology, vol. 1, no. 2, p. 100017, 2023
2023
-
[87]
External reasoning: Towards multi-large-language-models interchangeable assistance with human feedback,
A. Liu, “External reasoning: Towards multi-large-language-models interchangeable assistance with human feedback,” arXiv preprint arXiv:2307.12057, 2023
2023 arXiv
-
[88]
Use of language by generative ai tools in mathematical problem solving: The case of chatgpt,
W. Daher and F. Gierdien, “Use of language by generative ai tools in mathematical problem solving: The case of chatgpt,” African Journal of Research in Mathematics, Science and Technology Education , vol. 28, no. 2, pp. 222–235, 2024
2024
-
[89]
Investigating the e ffectiveness of chatgpt in mathematical reasoning and problem solving: Evidence from the viet- namese national high school graduation examination,
X.-Q. Dao and N.-B. Le, “Investigating the e ffectiveness of chatgpt in mathematical reasoning and problem solving: Evidence from the viet- namese national high school graduation examination,” arXiv preprint arXiv:2306.06331, 2023
2023 arXiv
-
[90]
Chatgpt challenges blended learning methodologies in engineering education: A case study in mathematics,
L. M. Sánchez-Ruiz, S. Moll-López, A. Nuñez-Pérez, J. A. Moraño- Fernández, and E. Vega-Fleitas, “Chatgpt challenges blended learning methodologies in engineering education: A case study in mathematics,” Applied Sciences, vol. 13, no. 10, p. 6039, 2023
2023
-
[91]
Mathematical capabilities of chatgpt,
S. Frieder, L. Pinchetti, R.-R. Gri ffiths, T. Salvatori, T. Lukasiewicz, P. Petersen, and J. Berner, “Mathematical capabilities of chatgpt,” Ad- vances in neural information processing systems , vol. 36, pp. 27 699– 27 744, 2023
2023
-
[92]
Deepseek and china’s ai renaissance: The rise of the four dragons and six tigers,
D. C. Youvan, “Deepseek and china’s ai renaissance: The rise of the four dragons and six tigers,” 2025
2025
-
[93]
Deepseek large-scale model: technical analysis and develop- ment prospect,
H. Liao, “Deepseek large-scale model: technical analysis and develop- ment prospect,” Journal of Computer Science and Electrical Engineer- ing, vol. 7, no. 1, pp. 33–37, 2025
2025
-
[94]
Deepseek models: A comprehensive survey of methods and applications
F. D. Puspitasari, C. Zhang, S. K. Dam, M. Zhang, T.-H. Kim, C. S. Hong, S.-H. Bae, C. Qin, J. Wei, G. Wang et al., “Deepseek models: A comprehensive survey of methods and applications.”
-
[95]
A survey on complex reasoning of large language models through the lens of self-evolution,
T. He, H. Li, J. Chen, R. Liu, Y . Cao, L. Liao, Z. Zheng, Z. Chu, J. Liang, M. Liu et al., “A survey on complex reasoning of large language models through the lens of self-evolution,” 2025
2025
-
[96]
Mixture of experts (moe): A big data perspective,
W. Gan, Z. Ning, Z. Qi, and P. S. Yu, “Mixture of experts (moe): A big data perspective,” arXiv preprint arXiv:2501.16352, 2025
2025 arXiv
-
[97]
A survey on post-training of large language models,
G. Tie, Z. Zhao, D. Song, F. Wei, R. Zhou, Y . Dai, W. Yin, Z. Yang, J. Yan, Y . Suet al., “A survey on post-training of large language models,” arXiv preprint arXiv:2503.06072, 2025
2025 arXiv
-
[98]
Ai for mathematics,
Q. Miao and F.-Y . Wang, “Ai for mathematics,” inArtificial Intelligence for Science (AI4S) Frontiers and Perspectives Based on Parallel Intelli- gence. Springer, 2024, pp. 21–39
2024
-
[99]
Dynamic chain-of-thought: Towards adaptive deep reason- ing,
L. Wang, “Dynamic chain-of-thought: Towards adaptive deep reason- ing,” arXiv preprint arXiv:2502.10428, 2025
2025 arXiv
-
[100]
Deepseekmath: Pushing the limits of mathematical reasoning in open language models,
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y . Li, Y . Wuet al., “Deepseekmath: Pushing the limits of mathematical reasoning in open language models,” arXiv preprint arXiv:2402.03300, 2024
2024 arXiv
-
[101]
Ensembling large language mod- els with process reward-guided tree search for better complex reason- ing,
S. Park, X. Liu, Y . Gong, and E. Choi, “Ensembling large language mod- els with process reward-guided tree search for better complex reason- ing,” arXiv preprint arXiv:2412.15797, 2024
2024 arXiv
-
[102]
Deepseek-v3 technical report,
A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan et al., “Deepseek-v3 technical report,”arXiv preprint arXiv:2412.19437, 2024
2024 arXiv
-
[103]
A review of deepseek models’ key in- novative techniques,
C. Wang and M. Kantarcioglu, “A review of deepseek models’ key in- novative techniques,”arXiv preprint arXiv:2503.11486, 2025
2025 arXiv
-
[104]
Deepseek-coder: When the large language model meets programming–the rise of code intelligence,
D. Guo, Q. Zhu, D. Yang, Z. Xie, K. Dong, W. Zhang, G. Chen, X. Bi, Y . Wu, Y . Liet al., “Deepseek-coder: When the large language model meets programming–the rise of code intelligence,”arXiv preprint arXiv:2401.14196, 2024
2024 arXiv
-
[105]
Welcome to the gemini era: Google deepmind and the information industry,
H. R. Saeidnia, “Welcome to the gemini era: Google deepmind and the information industry,”Library Hi Tech News, no. ahead-of-print, 2023
2023
-
[106]
Gemini versus chatgpt: appli- cations, performance, architecture, capabilities, and implementation,
N. Rane, S. Choudhary, and J. Rane, “Gemini versus chatgpt: appli- cations, performance, architecture, capabilities, and implementation,” Journal of Applied Artificial Intelligence, vol. 5, no. 1, pp. 69–93, 2024
2024
-
[107]
Towards system 2 reasoning in llms: Learning how to think with meta chain-of-though,
V . Xiang, C. Snell, K. Gandhi, A. Albalak, A. Singh, C. Blagden, D. Phung, R. Rafailov, N. Lile, D. Mahan et al. , “Towards system 2 reasoning in llms: Learning how to think with meta chain-of-though,” arXiv preprint arXiv:2501.04682, 2025
2025 arXiv
-
[108]
Campesato, Google Gemini for Python: Coding with Bard
O. Campesato, Google Gemini for Python: Coding with Bard . Stylus Publishing, LLC, 2024
2024
-
[109]
Or-llm-agent: Automating modeling and solving of operations research optimization problem with reasoning large lan- guage model,
B. Zhang and P. Luo, “Or-llm-agent: Automating modeling and solving of operations research optimization problem with reasoning large lan- guage model,” arXiv preprint arXiv:2503.10009, 2025
2025 arXiv
-
[110]
Evaluating large language models: Chatgpt-4, mistral 8x7b, and google gemini benchmarked against mmlu,
K. Ono and A. Morita, “Evaluating large language models: Chatgpt-4, mistral 8x7b, and google gemini benchmarked against mmlu,”Authorea Preprints, 2024
2024
-
[111]
To cot or not to cot? chain-of-thought helps mainly on math and symbolic reasoning,
Z. Sprague, F. Yin, J. D. Rodriguez, D. Jiang, M. Wadhwa, P. Singhal, X. Zhao, X. Ye, K. Mahowald, and G. Durrett, “To cot or not to cot? chain-of-thought helps mainly on math and symbolic reasoning,” arXiv preprint arXiv:2409.12183, 2024
2024 arXiv
-
[112]
Solv- ing quantitative reasoning problems with language models,
A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V . Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Soloet al., “Solv- ing quantitative reasoning problems with language models,”Advances in Neural Information Processing Systems, vol. 35, pp. 3843–3857, 2022
2022
-
[113]
Training verifiers to solve math word problems,
K. Cobbe, V . Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano et al., “Training verifiers to solve math word problems,” arXiv preprint arXiv:2110.14168, 2021
2021 arXiv
-
[114]
Measuring mathematical problem solving with the math dataset,
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt, “Measuring mathematical problem solving with the math dataset,” arXiv preprint arXiv:2103.03874, 2021
2021 arXiv
-
[115]
Applying large lan- guage models and chain-of-thought for automatic scoring,
G.-G. Lee, E. Latif, X. Wu, N. Liu, and X. Zhai, “Applying large lan- guage models and chain-of-thought for automatic scoring,” Computers and Education: Artificial Intelligence, vol. 6, p. 100213, 2024
2024
-
[116]
Llm performance assessment in computer science graduate entrance exams,
A. VarastehNezhad, R. Tavasoli, M. Masumi, and F. Taghiyareh, “Llm performance assessment in computer science graduate entrance exams,” in 2024 11th International Symposium on Telecommunications (IST) . IEEE, 2024, pp. 232–237
2024
-
[117]
Evaluating llms at de- tecting errors in llm responses,
R. Kamoi, S. S. S. Das, R. Lou, J. J. Ahn, Y . Zhao, X. Lu, N. Zhang, Y . Zhang, R. H. Zhang, S. R. Vummanthalaet al., “Evaluating llms at de- tecting errors in llm responses,”arXiv preprint arXiv:2404.03602, 2024
2024 arXiv
-
[118]
An exemplar for combining the collection, analysis, and interpretations of verbal and nonverbal data in qualitative research,
A. J. Onwuegbuzie and V . T. Byers, “An exemplar for combining the collection, analysis, and interpretations of verbal and nonverbal data in qualitative research,” International Journal of Education , vol. 6, no. 1, p. 183, 2014
2014
-
[119]
Be- yond final answers: Evaluating large language models for math tutor- ing,
A. Gupta, J. Reddig, T. Calo, D. Weitekamp, and C. J. MacLellan, “Be- yond final answers: Evaluating large language models for math tutor- ing,” arXiv preprint arXiv:2503.16460, 2025
2025 arXiv
-
[120]
Enhancing mathematical capabilities through chatgpt and similar generative artificial intelligence: Roles and challenges in solv- ing mathematical problems,
N. Rane, “Enhancing mathematical capabilities through chatgpt and similar generative artificial intelligence: Roles and challenges in solv- ing mathematical problems,” Available at SSRN 4603237, 2023
2023
-
[121]
How numerical precision affects mathematical reasoning ca- pabilities of llms,
G. Feng, K. Yang, Y . Gu, X. Ai, S. Luo, J. Sun, D. He, Z. Li, and L. Wang, “How numerical precision affects mathematical reasoning ca- pabilities of llms,” arXiv preprint arXiv:2410.13857, 2024
2024 arXiv
-
[122]
Proof automation with large lan- guage models,
M. Lu, B. Delaware, and T. Zhang, “Proof automation with large lan- guage models,” in Proceedings of the 39th IEEE /ACM International Conference on Automated Software Engineering, 2024, pp. 1509–1520
2024
-
[123]
Large language models as an interface to interact with api tools in natural language,
Y . G. Tesfagiorgis and B. M. Monteiro Silva, “Large language models as an interface to interact with api tools in natural language,” 2023
2023
-
[124]
Llms on the fly: Text-to-json for custom api calling,
M. Escarda-Fernández, I. López-Riobóo-Botana, S. Barro-Tojeiro, L. Padrón-Cousillas, S. Gonzalez-Vázquez, A. Carreiro-Alonso, and 21 P. Gómez-Area, “Llms on the fly: Text-to-json for custom api calling,” Proceedings of the SEPLN-CEDI, 2024
2024
-
[125]
Semantification of identifiers in mathe- matics for better math information retrieval,
M. Schubotz, A. Grigorev, M. Leich, H. S. Cohl, N. Meuschke, B. Gipp, A. S. Youssef, and V . Markl, “Semantification of identifiers in mathe- matics for better math information retrieval,” in Proceedings of the 39th International ACM SIGIR conference on Research and Developmen...
2016
-
[126]
The importance of data validation and parsing when working with external data sources,
A. Gallen, “The importance of data validation and parsing when working with external data sources,” 2024
2024
-
[127]
How can we know when language models know? on the calibration of language models for ques- tion answering,
Z. Jiang, J. Araki, H. Ding, and G. Neubig, “How can we know when language models know? on the calibration of language models for ques- tion answering,” Transactions of the Association for Computational Lin- guistics, vol. 9, pp. 962–977, 2021
2021
-
[128]
Exploring the impact of the output format on the evaluation of large language models for code translation,
M. Macedo, Y . Tian, F. Cogo, and B. Adams, “Exploring the impact of the output format on the evaluation of large language models for code translation,” in Proceedings of the 2024 IEEE /ACM First International Conference on AI Foundation Models and Software Engineering, 2024, ...
2024
-
[129]
Mathematical modelling abilities of artificial intelligence tools: The case of chatgpt,
C. Spreitzer, O. Straser, S. Zehetmeier, and K. Maaß, “Mathematical modelling abilities of artificial intelligence tools: The case of chatgpt,” Education Sciences, vol. 14, no. 7, p. 698, 2024
2024
-
[130]
Lemma: Learning from errors for mathematical advancement in llms,
Z. Pan, Y . Li, H. Lin, Q. Pei, Z. Tang, W. Wu, C. Ming, H. V . Zhao, C. He, and L. Wu, “Lemma: Learning from errors for mathematical advancement in llms,” arXiv preprint arXiv:2503.17439, 2025
2025 arXiv
-
[131]
Are large lan- guage models really good logical reasoners? a comprehensive evaluation and beyond,
F. Xu, Q. Lin, J. Han, T. Zhao, J. Liu, and E. Cambria, “Are large lan- guage models really good logical reasoners? a comprehensive evaluation and beyond,” IEEE Transactions on Knowledge and Data Engineering, 2025
2025
-
[132]
Benchmarking llms for real-world applications: From numerical metrics to contextual and qualitative evaluation,
H. I. Ashqar, “Benchmarking llms for real-world applications: From numerical metrics to contextual and qualitative evaluation,” Authorea Preprints, 2025
2025
-
[133]
Logicbench: Towards systematic evaluation of logical reasoning ability of large language models,
M. Parmar, N. Patel, N. Varshney, M. Nakamura, M. Luo, S. Mashetty, A. Mitra, and C. Baral, “Logicbench: Towards systematic evaluation of logical reasoning ability of large language models,” arXiv preprint arXiv:2404.15522, 2024
2024 arXiv
-
[134]
Llms can find mathemati- cal reasoning mistakes by pedagogical chain-of-thought,
Z. Jiang, H. Peng, S. Feng, F. Li, and D. Li, “Llms can find mathemati- cal reasoning mistakes by pedagogical chain-of-thought,”arXiv preprint arXiv:2405.06705, 2024
2024 arXiv
-
[135]
Pkrd-cot: A unified chain-of-thought prompting for multi-modal large language models in autonomous driving,
X. Luo, F. Ding, Y . Song, X. Zhang, and J. Loo, “Pkrd-cot: A unified chain-of-thought prompting for multi-modal large language models in autonomous driving,” arXiv preprint arXiv:2412.02025, 2024
2024 arXiv
-
[136]
The multi-round process matrix,
T. Ho ffreumon and O. Oreshkov, “The multi-round process matrix,” Quantum, vol. 5, p. 384, 2021
2021
-
[137]
Automatic chain of thought prompting in large language models,
Z. Zhang, A. Zhang, M. Li, and A. Smola, “Automatic chain of thought prompting in large language models,”arXiv preprint arXiv:2210.03493, 2022
2022 arXiv
-
[138]
Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models,
I. Mirzadeh, K. Alizadeh, H. Shahrokhi, O. Tuzel, S. Bengio, and M. Farajtabar, “Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models,” arXiv preprint arXiv:2410.05229, 2024
2024 arXiv
-
[139]
From explicit cot to im- plicit cot: Learning to internalize cot step by step,
Y . Deng, Y . Choi, and S. Shieber, “From explicit cot to im- plicit cot: Learning to internalize cot step by step,” arXiv preprint arXiv:2405.14838, 2024
2024 arXiv
-
[140]
Mathematical language models: A survey,
W. Liu, H. Hu, J. Zhou, Y . Ding, J. Li, J. Zeng, M. He, Q. Chen, B. Jiang, A. Zhou et al., “Mathematical language models: A survey,” arXiv preprint arXiv:2312.07622, 2023
2023 arXiv
-
[141]
Test-driven development for code generation,
N. S. Mathews and M. Nagappan, “Test-driven development for code generation,” arXiv preprint arXiv:2402.13521, 2024
2024 arXiv
-
[142]
Research methods in (mathematics) education,
A. H. Schoenfeld, “Research methods in (mathematics) education,” in Handbook of international research in mathematics education. Rout- ledge, 2010, pp. 481–533
2010
-
[143]
Artificial intelligence, computational thinking, and mathematics education,
G. Gadanidis, “Artificial intelligence, computational thinking, and mathematics education,” The International Journal of Information and Learning Technology, vol. 34, no. 2, pp. 133–139, 2017
2017
-
[144]
Mathematics and machine creativity: A survey on bridging mathematics with ai,
S. Liang, W. Zhang, T. Zhong, and T. Liu, “Mathematics and machine creativity: A survey on bridging mathematics with ai,” arXiv preprint arXiv:2412.16543, 2024
2024
-
[145]
Machine learning and ai learning: Understanding the revolution,
V . R. Boppana, “Machine learning and ai learning: Understanding the revolution,” Journal of Innovative Technologies, vol. 5, no. 1, 2022
2022
-
[146]
Shalev-Shwartz and S
S. Shalev-Shwartz and S. Ben-David, Understanding machine learning: From theory to algorithms. Cambridge university press, 2014
2014
-
[147]
Alpaydin, Introduction to machine learning
E. Alpaydin, Introduction to machine learning. MIT press, 2020
2020
-
[148]
Overview of supervised learning,
T. Hastie, R. Tibshirani, J. Friedman, T. Hastie, R. Tibshirani, and J. Friedman, “Overview of supervised learning,” The elements of statis- tical learning: Data mining, inference, and prediction, pp. 9–41, 2009
2009
-
[149]
M. E. Celebi and K. Aydin, Unsupervised learning algorithms . Springer, 2016, vol. 1
2016
-
[150]
Unsupervised learning based on artificial neural network: A review,
H. U. Dike, Y . Zhou, K. K. Deveerasetty, and Q. Wu, “Unsupervised learning based on artificial neural network: A review,” in 2018 IEEE International Conference on Cyborg and Bionic Systems (CBS). IEEE, 2018, pp. 322–327
2018
-
[151]
Deep reinforcement learning: An overview,
Y . Li, “Deep reinforcement learning: An overview,” arXiv preprint arXiv:1701.07274, 2017
2017 arXiv
-
[152]
Neural networks and deep learning: A paradigm shift in information processing, machine learning, and artificial intelli- gence,
S. Fitz and P. Romero, “Neural networks and deep learning: A paradigm shift in information processing, machine learning, and artificial intelli- gence,” The Palgrave handbook of technological finance, pp. 589–654, 2021
2021
-
[153]
Artificial intelligence: a survey on evolution, models, applica- tions and future trends,
Y . Lu, “Artificial intelligence: a survey on evolution, models, applica- tions and future trends,”Journal of Management Analytics, vol. 6, no. 1, pp. 1–29, 2019
2019
-
[154]
Gem- ini: a family of highly capable multimodal models,
G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican et al. , “Gem- ini: a family of highly capable multimodal models,” arXiv preprint arXiv:2312.11805, 2023
2023 arXiv
-
[155]
The benefits of a concise chain of thought on problem-solving in large language models,
M. Renze and E. Guven, “The benefits of a concise chain of thought on problem-solving in large language models,” in 2024 2nd Interna- tional Conference on Foundation and Large Language Models (FLLM). IEEE, 2024, pp. 476–483
2024
-
[156]
Question-analysis prompting improves llm performance in reasoning tasks,
D. Yugeswardeenoo, K. Zhu, and S. O’Brien, “Question-analysis prompting improves llm performance in reasoning tasks,”arXiv preprint arXiv:2407.03624, 2024
2024 arXiv
-
[157]
Towards reasoning era: A survey of long chain-of-thought for reasoning large language models,
Q. Chen, L. Qin, J. Liu, D. Peng, J. Guan, P. Wang, M. Hu, Y . Zhou, T. Gao, and W. Che, “Towards reasoning era: A survey of long chain-of-thought for reasoning large language models,” arXiv preprint arXiv:2503.09567, 2025
2025 arXiv
-
[158]
A review of current trends, techniques, and challenges in large language models (llms),
R. Patil and V . Gudivada, “A review of current trends, techniques, and challenges in large language models (llms),” Applied Sciences, vol. 14, no. 5, p. 2074, 2024
2024
-
[159]
Quantifying the capabilities of llms across scale and precision,
S. Badshah and H. Sajjad, “Quantifying the capabilities of llms across scale and precision,” arXiv preprint arXiv:2405.03146, 2024
2024 arXiv
-
[160]
A sys- tematic survey and critical review on evaluating large language models: Challenges, limitations, and recommendations,
M. T. R. Laskar, S. Alqahtani, M. S. Bari, M. Rahman, M. A. M. Khan, H. Khan, I. Jahan, A. Bhuiyan, C. W. Tan, M. R. Parvez et al., “A sys- tematic survey and critical review on evaluating large language models: Challenges, limitations, and recommendations,” in Proceedings of ...
2024
-
[161]
Matheval: A compre- hensive benchmark for evaluating large language models on mathemati- cal reasoning capabilities
Z. Liu, T. Liu, Z. Chen, M. Tian, W. Luo et al., “Matheval: A compre- hensive benchmark for evaluating large language models on mathemati- cal reasoning capabilities.”
-
[162]
An empirical study of data ability boundary in llms’ math reasoning,
Z. Chen, Y . Chen, J. Han, Z. Huang, J. Qi, and Y . Zhou, “An empirical study of data ability boundary in llms’ math reasoning,” arXiv preprint arXiv:2403.00799, 2024
2024 arXiv
-
[163]
Benchmarking automated theorem proving with large language models,
V . Lama, C. Ma, and T. Ghosal, “Benchmarking automated theorem proving with large language models,” in Proceedings of the 1st Work- shop on NLP for Science (NLP4Science), 2024, pp. 208–218
2024
-
[164]
Harnessing llms for api interactions: A framework for classification and synthetic data generation,
C. Tao, X. Fan, and Y . Yang, “Harnessing llms for api interactions: A framework for classification and synthetic data generation,” arXiv preprint arXiv:2409.11703, 2024
2024 arXiv
-
[165]
Evaluating the problem solving abilities of chatgpt,
F. Zeng, “Evaluating the problem solving abilities of chatgpt,” 2023
2023
-
[166]
Understanding of convolutional neural net- work (cnn): A review,
P. Purwono, A. Ma’arif, W. Rahmaniar, H. I. K. Fathurrahman, A. Z. K. Frisky, and Q. M. ul Haq, “Understanding of convolutional neural net- work (cnn): A review,” International Journal of Robotics and Control Systems, vol. 2, no. 4, pp. 739–748, 2022
2022
-
[167]
Recurrent neural networks (rnns): A gentle introduc- tion and overview,
R. M. Schmidt, “Recurrent neural networks (rnns): A gentle introduc- tion and overview,”arXiv preprint arXiv:1912.05911, 2019. Appendix 6.1. Examples from mathematical problems ’ 22 Figure 4: Selected questions and LLM answers from the GSM8K dataset (a) MATH500 example 1 (b) M...
1912 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.