REVIEW 3 major objections 5 minor 43 references
AI Reasoning Models for Problem Solving in Physics
T0 review · 3 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Five independent runs of OpenAI's o3-mini solved 94% of 408 text-only introductory physics problems, with 100% in many mechanics chapters and clear drops in waves and thermodynamics.
desk verdict A useful but limited empirical snapshot: o3-mini solves 94% of text-only Halliday problems, yet training-data contamination and small samples keep the result from being a clean test of reasoning. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is a repeated-sampling reliability protocol. Each problem statement is submitted to o3-mini exactly as written, five times, and a problem is scored as solved only when every one of the five final answers agrees with the textbook key. That all-five agreement criterion is what turns a single answer into a claim about reliability, and the resulting per-chapter percentages are the paper's main evidence. A secondary mechanism is the failure taxonomy: every failed solution was read by a human expert and assigned to one of two categories, unchecked verbal reasoning or uncorrected calculation error, which then explains why the model behaves well in some topics and poorly in ot
What would settle it
Take the 24 failed problems from this study and rerun each one on o3-mini with a single added instruction: check the final answer against the problem's physical conditions before responding. If most of the 24 become consistently correct, the paper's explanation of the failures—missing self-evaluation—is supported; if they still fail, the mechanism is not the one described and the reliability estimate needs a different explanation.
Extended reading notes
Core claim
On the paper's own terms, the discovery is a measured reliability profile, not a physics result. For each of 408 text-only problems copied verbatim from the textbook, five separate o3-mini solutions were generated; 384 problems produced the answer-key result in all five attempts, and 24 did not. The per-chapter table shows the performance is topic-dependent: early chapters on measurement, straight-line motion, vectors, force, kinetic energy, center of mass, rotation, rolling, fluids, and oscillations are at or near 100%, while Waves–I is 87%, Waves–II is 76%, and the three chapters leading to entropy and the second law of thermodynamics decline from 96% to 92% to 88%. Examining the 24 failur
Load-bearing premise
The load-bearing premise is that the 408 text-only odd-numbered problems are a fair sample of every odd-numbered problem in the textbook, so the per-chapter percentages actually describe each chapter's difficulty; if problems with figures or tables differ systematically, the 94% and the topic trend are specific to text-only problems only.
Editorial extensions
If this is right
- In an introductory mechanics course, an instructor can expect o3-mini's output to be correct on nearly every text-only story problem it is asked, so its main risk there is misuse (copying) rather than wrong answers.
- For waves and, to a lesser extent, thermodynamics, the same model should be treated as a draft solver: its per-chapter failure rates are 13%, 24%, 8%, 12%, and 4% in the hardest chapters, so answers need independent checking.
- Because the all-five rule is stricter than 'the first answer was right,' the 94% figure is a conservative estimate of single-attempt accuracy; the model is likely right even more often on the problems it 'solved' by this definition.
- The identified failure modes imply an engineering remedy: augmenting a reasoning model with a calculator and a physics simulator should remove most of the 24 failures, since they stem from unverified numerical and physical steps.
- The educational payoff depends on how the output is used: if students learn from multiple generated examples, the worked-example literature cited in the paper says the examples must be coordinated and used with self-explanation; unchecked use invites cheating.
Reading between the lines
- Because the dataset deliberately excludes every problem containing a figure or table, the 94% and the topic-level dips are estimates for text-only problems only; figure-heavy problems in waves and thermodynamics could make the real per-topic reliability different.
- The decline across later chapters could be a training-data frequency effect, but the paper does not establish that; a comparison using volume 2 topics or controlled problems with matched mathematical complexity would separate frequency from difficulty.
- The strict all-five pass rule measures consistency with the answer key, not correctness of the key itself; since the authors do not verify the textbook answers, some 'failures' could be erroneous keys and some successes the same error five times.
- A cheap pedagogical experiment follows from the paper's error taxonomy: ask the model to check its own final answer before outputting it; if most of the 24 failures vanish, self-verification prompting is a sufficient safeguard, and if not, external computation tools are needed.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper evaluates OpenAI's o3-mini on 408 text-only odd-numbered end-of-chapter problems from Halliday and Resnick's Fundamentals of Physics, Vol. 1, spanning 20 chapters. For each problem, five independent solutions were generated; a problem was counted as solved only if all five final answers matched the textbook answer key. The authors report an overall success rate of 94% (384/408), with visibly lower rates in waves (Chs. 16–17) and thermodynamics (Chs. 18–20), and they illustrate failure modes with three worked examples. They conclude that the model can solve most introductory mechanics story problems reliably, but they also note that the question of reliability remains open and that training-data exposure may explain the high accuracy.
Significance. If interpreted narrowly—'o3-mini reproduces the answer-key final answers for 94% of these exact text-only problems'—the study is a useful and easily repeatable empirical data point for physics educators. The all-five-correct success criterion is transparent and conservative; the use of a standard textbook answer key and manual expert checking gives a clear operational definition of correctness. The per-topic breakdown addresses a practically important question, and the paper is refreshingly candid about its limitations. However, the broader inference that o3-mini is a 'reliable' physics problem-solver is not supported because the exact problems come from a decades-old, widely distributed textbook and the authors themselves concede they are likely in the model's training data. The text-only selection rule further limits generalizability. These are fixable with additional experiments and a more cautious framing.
major comments (3)
- [§IV and Abstract] The central inference is undercut by training-data contamination. The paper tests exact problems from a standard textbook that has been widely circulated for decades and is almost certainly represented in o3-mini's training corpus. The authors themselves write in §IV: 'For the kind of data that was the subject of this study—story problems from introductory physics—there may be enough of it in the training of these models. Then the high accuracy in solving them is expected,' and later offer 'underrepresented on the Internet' as an alternative explanation for the topic trend. As written, the 94% figure does not distinguish reasoning from memorization. To make the central claim load-bearing, the authors should add held-out or modified problems (e.g., systematically changed numbers, names, or contexts), a contamination audit, or at minimum re-state the abstract to describe performance on the
- [§II Methods / Table I] The analysis excludes every problem containing a figure or table, leaving only 408 of 629 odd-numbered problems. These 408 are not shown to be representative: figure/table problems often test graphical interpretation and may differ systematically in difficulty. This selection directly affects the per-chapter percentages and the topic-trend conclusion. For example, Chapter 12 has only 7 text-only problems, so its 86% rate is statistically indistinguishable from 100%; even the overall 94% should be accompanied by a confidence interval. The authors should either analyze the excluded problems (e.g., by using a multimodal model or manual transcription), or explicitly limit all conclusions to text-only problems and quantify the uncertainty.
- [§III Results / success criterion] The binary all-five-success criterion is conservative for a reliability claim, but it conflates stochasticity with lack of competence. Chapter 19, Problem 11 is counted as unsolved even though four of five runs produced the correct answer. This makes the reported success rate highly sensitive to n=5; a single anomalous run moves a problem from solved to unsolved. The authors should report the distribution of per-problem success counts (5/5, 4/5, ...) and show whether the overall and per-chapter conclusions are robust to using a 4-of-5 criterion. Without this, the claim that o3-mini 'struggles' with certain topics is not well quantified.
minor comments (5)
- [§I and §III] Typos: 'calleduser prompt' lacks a space; 'plausible and and for what reasons' has a duplicated 'and'; in §III, 'it made did not make this assumption' is ungrammatical.
- [§II Methods] The manuscript says problems were copied into LaTeX, but it does not state how final-answer matching handled equivalent forms (e.g., 5.6 kJ vs 5600 J, sign conventions, significant figures). A short statement of the equivalence rule would improve reproducibility.
- [Table I] The 'Problems Solved' column shows only percentages. Adding 'n/N (%)' would make the raw denominators transparent and avoid the impression that all percentages are equally precise.
- [§IV] The paragraph on the worked-example effect is interesting but not connected to the data; it reads as a separate literature review. Consider moving details to a 'Pedagogical Implications' section or shortening.
- [Data availability] No data availability statement or supplementary materials are provided. Given the modest dataset (2,040 outputs), releasing the problem texts, model outputs, and answer-key labels would allow independent verification.
Circularity Check
No circularity: the 94% figure is an externally benchmarked measurement, not a derivation from fitted inputs or self-citations.
full rationale
The paper's only quantitative claim is an observed success rate: N=408 text-only odd-numbered problems from Halliday/Resnick were sent verbatim to o3-mini, five solutions per problem were generated, and a human expert compared final answers against the textbook answer key (§II). The criterion that all five solutions must match is an explicit operational definition of 'successfully solved' (§II), not a parameter fitted to the data. The result is therefore not equivalent to any input by construction: the answer key is an outside standard, and the problems were not generated from the model's own outputs. No load-bearing step depends on a citation to the authors' previous work; the reference list contains no Bralin/Rebello items. No uniqueness theorem or ansatz is imported from prior work. The paper itself flags a real limitation in §IV—'there may be enough of it in the training of these models. Then the high accuracy in solving them is expected'—and §V notes the topic-trend explanation is open. That is a threat to generalization/construct validity (possible memorization), not circularity: the measured accuracy could be inflated by training-data overlap, but the measurement is still an external empirical observation rather than a definitional consequence. The sampling restriction to text-only problems (Methods) likewise affects representativeness, not circularity. Hence no circular step is present.
Assumptions & free parameters
free parameters (1)
- n_sample (solutions per problem) =
5
assumptions (4)
- domain assumption The textbook answer key is correct and was not independently verified.
- domain assumption The odd-numbered problems with answer keys are representative of all problems in each chapter.
- ad hoc to paper The 408 text-only problems are representative of the full problem set despite excluding all figure/table problems.
- domain assumption Five samples per problem are sufficient to measure reliability, and a single incorrect run means the problem is unsolved.
Cite this review
Pith. "Pith review of AI Reasoning Models for Problem Solving in Physics." pith.science (2026). https://pith.science/paper/54BFGS3I
@misc{pith2026250820941,
author = {Pith},
title = {Pith review of: AI Reasoning Models for Problem Solving in Physics},
year = {2026},
howpublished = {\url{https://pith.science/paper/54BFGS3I}},
note = {Machine review of arXiv:2508.20941}
}
read the original abstract
Reasoning models are the new generation of Large Language Models (LLMs) capable of complex problem solving. Their reliability in solving introductory physics problems was tested by evaluating a sample of n = 5 solutions generated by one such model -- OpenAI's o3-mini -- per each problem from 20 chapters of a standard undergraduate textbook. In total, N = 408 problems were given to the model and N x n = 2,040 generated solutions examined. The model successfully solved 94% of the problems posed, excelling at the beginning topics in mechanics but struggling with the later ones such as waves and thermodynamics.
Reference graph
Works this paper leans on
-
[1]
How reliable are AI reasoning models when solving story problems in physics?
-
[2]
What is the distribution of the story problem-solving ability of AI reasoning models across standard topics in physics? As presented in the following sections, the state-of-the-art AI reasoning models may be reliable problem-solvers in the be- ginning topics of a typical introductory physics course, but still struggle with solving problems in later topics...
-
[3]
Measurement 16 13 13 (100%)
-
[4]
Motion Along a Straight Line 35 25 25 (100%)
-
[5]
Vectors 22 17 17 (100%)
-
[6]
Motion in Two and Three Dimensions 41 32 30 (94%)
-
[7]
Force and Motion–I 34 18 16 (89%)
-
[8]
Force and Motion–II 30 15 15 (100%)
Show all 43 references
-
[9]
Kinetic Energy and Work 26 17 17 (100%)
-
[10]
Potential Energy and Conservation of Energy 33 12 11 (92%)
-
[11]
Center of Mass and Linear Momentum 40 22 22 (100%)
-
[12]
Rotation 34 25 24 (96%)
-
[13]
Rolling, Torque, and Angular Momentum 35 18 18 (100%)
-
[14]
Equilibrium and Elasticity 26 7 6 (86%)
-
[15]
Gravitation 35 26 25 (96%)
-
[16]
Fluids 36 25 25 (100%)
-
[17]
Oscillations 32 20 19 (95%)
-
[18]
Waves–I 30 23 20 (87%)
-
[19]
Waves–II 35 29 22 (76%)
-
[20]
Temperature, Heat, and the First Law of Thermodynamics 33 23 22 (96%)
-
[21]
The Kinetic Theory of Gases 32 25 23 (92%)
-
[22]
Tak- ing the square root: v = √ 3.335 × 1014 ≈ 5.78 × 105 m/s
Entropy and the Second Law of Thermodynamics 24 16 14 (88%) 629 408 384 (94%) Ch. 13, Problem 41 : Two neutron stars are sep- arated by a distance of 1.0 × 1010 m. They each have a mass of 1.0 × 1030 kg and a radius of 1.0×105 m. They are initially at rest with respect to each...
2025
-
[23]
D. H. Jonassen, Designing Research-Based Instruction for Story Problems, Educational Psychology Review 15, 267 (2003)
2003
-
[24]
L. Ding, N. Reay, A. Lee, and L. Bao, Exploring the role of conceptual scaffolding in solving synthesis problems, Phys. Rev. ST Phys. Educ. Res. 7 (2011)
2011
-
[25]
K. D. Wang, E. Burkholder, C. Wieman, S. Salehi, and N. Haber, Examining the potential and pitfalls of ChatGPT in science and engineering problem-solving, Frontiers in Educa- tion 8 (2024)
2024
-
[26]
Kieser and P
F. Kieser and P. Wulff, Using large language models to probe cognitive constructs, augment data, and design instructional materials, in Machine Learning in Educational Sciences: Ap- proaches, Applications and Advances , edited by M. S. Khine (Springer Nature Singapore, Singapo...
2024
-
[27]
B. A. Huberman and T. Hogg, Phase transitions in artificial intelligence systems, Artif. Intell. 33, 155 (1987)
1987
-
[28]
Bengio, Y
Y . Bengio, Y . Lecun, and G. Hinton, Deep learning for AI, Commun. ACM 64, 58 (2021)
2021
-
[29]
OpenAI, Introducing ChatGPT, https://openai.com/index/ chatgpt/ (2022)
2022
-
[30]
J. Yang, Z. Wang, Y . Lin, and Z. Zhao, Problematic To- kens: Tokenizer Bias in Large Language Models (2024), arXiv:2406.11214 [cs.CL]
2024 arXiv
-
[31]
Glazer, E
E. Glazer, E. Erdil, T. Besiroglu, D. Chicharro, E. Chen, A. Gunning, C. F. Olsson, J.-S. Denain, A. Ho, E. de Oliveira Santos, O. Järviniemi, M. Barnett, R. San- dler, M. Vrzala, J. Sevilla, Q. Ren, E. Pratt, L. Levine, G. Barkley, N. Stewart, B. Grechuk, T. Grechuk, S. V . E...
2024 arXiv
-
[32]
D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y . Pang, J. Di- rani, J. Michael, and S. R. Bowman, GPQA: A Graduate-Level Google-Proof Q&A Benchmark (2023), arXiv:2311.12022 [cs.AI]
2023 arXiv
-
[33]
OpenAI, Reasoning models: Explore advanced reason- ing and problem-solving models, https://platform.openai.com/ docs/guides/reasoning?api-mode=responses (n.d.), Accessed: May 31, 2025
2025
-
[34]
OpenAI, Introducing OpenAI o1-preview, https://openai.com/ index/introducing-openai-o1-preview/ (2024)
2024
-
[35]
OpenAI, OpenAI o3-mini: Pushing the frontier of cost- effective reasoning, https://openai.com/index/openai-o3-mini/ (2025)
2025
-
[36]
Halliday, R
D. Halliday, R. Resnick, and J. Walker, Fundamentals of Physics, 12th ed. (John Wiley & Sons, Inc, 2022)
2022
-
[37]
OpenAI, Reasoning models: Advice on prompt- ing, https://platform.openai.com/docs/guides/reasoning? api-mode=responses#advice-on-prompting (n.d.), Accessed: June 1, 2025
2025
-
[38]
Mialon, R
G. Mialon, R. Dessì, M. Lomeli, C. Nalmpantis, R. Pasunuru, R. Raileanu, B. Rozière, T. Schick, J. Dwivedi-Yu, A. Celiky- ilmaz, E. Grave, Y . LeCun, and T. Scialom, Augmented Lan- guage Models: a Survey, Transactions on Machine Learning Research (2023)
2023
-
[39]
R. K. Atkinson, S. J. Derry, A. Renkl, and D. Wortham, Learn- ing from Examples: Instructional Principles from the Worked Examples Research, Review of Educational Research 70, 181 (2000)
2000
-
[40]
Sweller, Cognitive Load During Problem Solving: Effects on Learning, Cognitive Science 12, 257 (1988)
J. Sweller, Cognitive Load During Problem Solving: Effects on Learning, Cognitive Science 12, 257 (1988)
1988
-
[41]
M. T. Chi, M. Bassok, M. W. Lewis, P. Reimann, and R. Glaser, Self-Explanations: How Students Study and Use Examples in Learning to Solve Problems, Cognitive Science13, 145 (1989)
1989
-
[42]
Ouyang, J
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schul- man, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe, Training language models to follow instruction...
2022 arXiv
-
[43]
com/index/introducing-o3-and-o4-mini/ (2025)
OpenAI, Introducing OpenAI o3 and o4-mini, https://openai. com/index/introducing-o3-and-o4-mini/ (2025)
2025
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.