Pith. sign in

REVIEW 3 major objections 5 minor 43 references

AI Reasoning Models for Problem Solving in Physics

T0 review · 3 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Five independent runs of OpenAI's o3-mini solved 94% of 408 text-only introductory physics problems, with 100% in many mechanics chapters and clear drops in waves and thermodynamics.

desk verdict A useful but limited empirical snapshot: o3-mini solves 94% of text-only Halliday problems, yet training-data contamination and small samples keep the result from being a clean test of reasoning. read the letter →

arxiv 2508.20941 v1 pith:54BFGS3I submitted 2025-08-28 physics.ed-ph

classification physics.ed-ph PACS 01.40.-d
keywords physicseducationlargelanguagemodelsreasoningo3-miniintroductoryproblemsolvingworkedexamplesreliability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to measure how reliably a new class of large language models—'reasoning models'—can solve ordinary end-of-chapter physics story problems, and whether that reliability is the same across topics. The authors take OpenAI's o3-mini, feed it 408 text-only odd-numbered problems from a standard mechanics-and-thermodynamics textbook with no extra instructions, and generate five solutions per problem. A problem is counted as solved only if all five final answers match the textbook's answer key. By that strict rule the model solved 384 of 408 problems (94%), with perfect or near-perfect performance across most early mechanics chapters, a sharp drop on the two waves chapters (87% and 76%), and a gradual slide across the three thermodynamics chapters (96%, 92%, 88%). The authors conclude that reasoning models are already reliable enough to be useful as generators of worked examples for introductory mechanics, but not yet trustworthy without checking for later topics.

What carries the argument

The carrying mechanism is a repeated-sampling reliability protocol. Each problem statement is submitted to o3-mini exactly as written, five times, and a problem is scored as solved only when every one of the five final answers agrees with the textbook key. That all-five agreement criterion is what turns a single answer into a claim about reliability, and the resulting per-chapter percentages are the paper's main evidence. A secondary mechanism is the failure taxonomy: every failed solution was read by a human expert and assigned to one of two categories, unchecked verbal reasoning or uncorrected calculation error, which then explains why the model behaves well in some topics and poorly in ot

What would settle it

Take the 24 failed problems from this study and rerun each one on o3-mini with a single added instruction: check the final answer against the problem's physical conditions before responding. If most of the 24 become consistently correct, the paper's explanation of the failures—missing self-evaluation—is supported; if they still fail, the mechanism is not the one described and the reliability estimate needs a different explanation.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is a measured reliability profile, not a physics result. For each of 408 text-only problems copied verbatim from the textbook, five separate o3-mini solutions were generated; 384 problems produced the answer-key result in all five attempts, and 24 did not. The per-chapter table shows the performance is topic-dependent: early chapters on measurement, straight-line motion, vectors, force, kinetic energy, center of mass, rotation, rolling, fluids, and oscillations are at or near 100%, while Waves–I is 87%, Waves–II is 76%, and the three chapters leading to entropy and the second law of thermodynamics decline from 96% to 92% to 88%. Examining the 24 failur

Load-bearing premise

The load-bearing premise is that the 408 text-only odd-numbered problems are a fair sample of every odd-numbered problem in the textbook, so the per-chapter percentages actually describe each chapter's difficulty; if problems with figures or tables differ systematically, the 94% and the topic trend are specific to text-only problems only.

Editorial extensions

If this is right

  • In an introductory mechanics course, an instructor can expect o3-mini's output to be correct on nearly every text-only story problem it is asked, so its main risk there is misuse (copying) rather than wrong answers.
  • For waves and, to a lesser extent, thermodynamics, the same model should be treated as a draft solver: its per-chapter failure rates are 13%, 24%, 8%, 12%, and 4% in the hardest chapters, so answers need independent checking.
  • Because the all-five rule is stricter than 'the first answer was right,' the 94% figure is a conservative estimate of single-attempt accuracy; the model is likely right even more often on the problems it 'solved' by this definition.
  • The identified failure modes imply an engineering remedy: augmenting a reasoning model with a calculator and a physics simulator should remove most of the 24 failures, since they stem from unverified numerical and physical steps.
  • The educational payoff depends on how the output is used: if students learn from multiple generated examples, the worked-example literature cited in the paper says the examples must be coordinated and used with self-explanation; unchecked use invites cheating.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the dataset deliberately excludes every problem containing a figure or table, the 94% and the topic-level dips are estimates for text-only problems only; figure-heavy problems in waves and thermodynamics could make the real per-topic reliability different.
  • The decline across later chapters could be a training-data frequency effect, but the paper does not establish that; a comparison using volume 2 topics or controlled problems with matched mathematical complexity would separate frequency from difficulty.
  • The strict all-five pass rule measures consistency with the answer key, not correctness of the key itself; since the authors do not verify the textbook answers, some 'failures' could be erroneous keys and some successes the same error five times.
  • A cheap pedagogical experiment follows from the paper's error taxonomy: ask the model to check its own final answer before outputting it; if most of the 24 failures vanish, self-verification prompting is a sufficient safeguard, and if not, external computation tools are needed.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper evaluates OpenAI's o3-mini on 408 text-only odd-numbered end-of-chapter problems from Halliday and Resnick's Fundamentals of Physics, Vol. 1, spanning 20 chapters. For each problem, five independent solutions were generated; a problem was counted as solved only if all five final answers matched the textbook answer key. The authors report an overall success rate of 94% (384/408), with visibly lower rates in waves (Chs. 16–17) and thermodynamics (Chs. 18–20), and they illustrate failure modes with three worked examples. They conclude that the model can solve most introductory mechanics story problems reliably, but they also note that the question of reliability remains open and that training-data exposure may explain the high accuracy.

Significance. If interpreted narrowly—'o3-mini reproduces the answer-key final answers for 94% of these exact text-only problems'—the study is a useful and easily repeatable empirical data point for physics educators. The all-five-correct success criterion is transparent and conservative; the use of a standard textbook answer key and manual expert checking gives a clear operational definition of correctness. The per-topic breakdown addresses a practically important question, and the paper is refreshingly candid about its limitations. However, the broader inference that o3-mini is a 'reliable' physics problem-solver is not supported because the exact problems come from a decades-old, widely distributed textbook and the authors themselves concede they are likely in the model's training data. The text-only selection rule further limits generalizability. These are fixable with additional experiments and a more cautious framing.

major comments (3)
  1. [§IV and Abstract] The central inference is undercut by training-data contamination. The paper tests exact problems from a standard textbook that has been widely circulated for decades and is almost certainly represented in o3-mini's training corpus. The authors themselves write in §IV: 'For the kind of data that was the subject of this study—story problems from introductory physics—there may be enough of it in the training of these models. Then the high accuracy in solving them is expected,' and later offer 'underrepresented on the Internet' as an alternative explanation for the topic trend. As written, the 94% figure does not distinguish reasoning from memorization. To make the central claim load-bearing, the authors should add held-out or modified problems (e.g., systematically changed numbers, names, or contexts), a contamination audit, or at minimum re-state the abstract to describe performance on the
  2. [§II Methods / Table I] The analysis excludes every problem containing a figure or table, leaving only 408 of 629 odd-numbered problems. These 408 are not shown to be representative: figure/table problems often test graphical interpretation and may differ systematically in difficulty. This selection directly affects the per-chapter percentages and the topic-trend conclusion. For example, Chapter 12 has only 7 text-only problems, so its 86% rate is statistically indistinguishable from 100%; even the overall 94% should be accompanied by a confidence interval. The authors should either analyze the excluded problems (e.g., by using a multimodal model or manual transcription), or explicitly limit all conclusions to text-only problems and quantify the uncertainty.
  3. [§III Results / success criterion] The binary all-five-success criterion is conservative for a reliability claim, but it conflates stochasticity with lack of competence. Chapter 19, Problem 11 is counted as unsolved even though four of five runs produced the correct answer. This makes the reported success rate highly sensitive to n=5; a single anomalous run moves a problem from solved to unsolved. The authors should report the distribution of per-problem success counts (5/5, 4/5, ...) and show whether the overall and per-chapter conclusions are robust to using a 4-of-5 criterion. Without this, the claim that o3-mini 'struggles' with certain topics is not well quantified.
minor comments (5)
  1. [§I and §III] Typos: 'calleduser prompt' lacks a space; 'plausible and and for what reasons' has a duplicated 'and'; in §III, 'it made did not make this assumption' is ungrammatical.
  2. [§II Methods] The manuscript says problems were copied into LaTeX, but it does not state how final-answer matching handled equivalent forms (e.g., 5.6 kJ vs 5600 J, sign conventions, significant figures). A short statement of the equivalence rule would improve reproducibility.
  3. [Table I] The 'Problems Solved' column shows only percentages. Adding 'n/N (%)' would make the raw denominators transparent and avoid the impression that all percentages are equally precise.
  4. [§IV] The paragraph on the worked-example effect is interesting but not connected to the data; it reads as a separate literature review. Consider moving details to a 'Pedagogical Implications' section or shortening.
  5. [Data availability] No data availability statement or supplementary materials are provided. Given the modest dataset (2,040 outputs), releasing the problem texts, model outputs, and answer-key labels would allow independent verification.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the 94% figure is an externally benchmarked measurement, not a derivation from fitted inputs or self-citations.

full rationale

The paper's only quantitative claim is an observed success rate: N=408 text-only odd-numbered problems from Halliday/Resnick were sent verbatim to o3-mini, five solutions per problem were generated, and a human expert compared final answers against the textbook answer key (§II). The criterion that all five solutions must match is an explicit operational definition of 'successfully solved' (§II), not a parameter fitted to the data. The result is therefore not equivalent to any input by construction: the answer key is an outside standard, and the problems were not generated from the model's own outputs. No load-bearing step depends on a citation to the authors' previous work; the reference list contains no Bralin/Rebello items. No uniqueness theorem or ansatz is imported from prior work. The paper itself flags a real limitation in §IV—'there may be enough of it in the training of these models. Then the high accuracy in solving them is expected'—and §V notes the topic-trend explanation is open. That is a threat to generalization/construct validity (possible memorization), not circularity: the measured accuracy could be inflated by training-data overlap, but the measurement is still an external empirical observation rather than a definitional consequence. The sampling restriction to text-only problems (Methods) likewise affects representativeness, not circularity. Hence no circular step is present.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central quantitative claim (94% success) rests on several unverified assumptions: the correctness of the textbook answer key, the representativeness of odd-numbered text-only problems, and the adequacy of n=5 as a reliability threshold. None of these are tested, and the last two are acknowledged only implicitly.

free parameters (1)
  • n_sample (solutions per problem) = 5
    Hand-chosen threshold: a problem counts as solved only if all 5 generated solutions match the answer key. This directly determines the reported 94% success rate; using fewer samples would likely raise the apparent success rate.
assumptions (4)
  • domain assumption The textbook answer key is correct and was not independently verified.
    Methods: 'The textbook answers themselves were not evaluated for correctness. We assume that a book with so many editions has identified any potential errors in its answer key.' If the answer key contains errors, the 94% figure is affected.
  • domain assumption The odd-numbered problems with answer keys are representative of all problems in each chapter.
    Methods: only odd-numbered problems were selected; no comparison to even-numbered problems is made, so representativeness is assumed.
  • ad hoc to paper The 408 text-only problems are representative of the full problem set despite excluding all figure/table problems.
    Methods: 'o3-mini only handles text... all problems that contain a figure or a table were ignored.' This exclusion is a practical constraint, not a sampling justification.
  • domain assumption Five samples per problem are sufficient to measure reliability, and a single incorrect run means the problem is unsolved.
    Methods: n_sample = 5 and success requires all 5 correct. The paper does not justify this threshold or test its sensitivity, yet the central percentage depends on it.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AI Reasoning Models for Problem Solving in Physics." pith.science (2026). https://pith.science/paper/54BFGS3I

@misc{pith2026250820941,
  author       = {Pith},
  title        = {Pith review of: AI Reasoning Models for Problem Solving in Physics},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/54BFGS3I}},
  note         = {Machine review of arXiv:2508.20941}
}
read the original abstract

Reasoning models are the new generation of Large Language Models (LLMs) capable of complex problem solving. Their reliability in solving introductory physics problems was tested by evaluating a sample of n = 5 solutions generated by one such model -- OpenAI's o3-mini -- per each problem from 20 chapters of a standard undergraduate textbook. In total, N = 408 problems were given to the model and N x n = 2,040 generated solutions examined. The model successfully solved 94% of the problems posed, excelling at the beginning topics in mechanics but struggling with the later ones such as waves and thermodynamics.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

43 extracted references · 39 canonical work pages

  1. [1]

    How reliable are AI reasoning models when solving story problems in physics?

  2. [2]

    o-series

    What is the distribution of the story problem-solving ability of AI reasoning models across standard topics in physics? As presented in the following sections, the state-of-the-art AI reasoning models may be reliable problem-solvers in the be- ginning topics of a typical introductory physics course, but still struggle with solving problems in later topics...

  3. [3]

    Measurement 16 13 13 (100%)

  4. [4]

    Motion Along a Straight Line 35 25 25 (100%)

  5. [5]

    Vectors 22 17 17 (100%)

  6. [6]

    Motion in Two and Three Dimensions 41 32 30 (94%)

  7. [7]

    Force and Motion–I 34 18 16 (89%)

  8. [8]

    Force and Motion–II 30 15 15 (100%)

Show all 43 references
  1. [9]

    Kinetic Energy and Work 26 17 17 (100%)

  2. [10]

    Potential Energy and Conservation of Energy 33 12 11 (92%)

  3. [11]

    Center of Mass and Linear Momentum 40 22 22 (100%)

  4. [12]

    Rotation 34 25 24 (96%)

  5. [13]

    Rolling, Torque, and Angular Momentum 35 18 18 (100%)

  6. [14]

    Equilibrium and Elasticity 26 7 6 (86%)

  7. [15]

    Gravitation 35 26 25 (96%)

  8. [16]

    Fluids 36 25 25 (100%)

  9. [17]

    Oscillations 32 20 19 (95%)

  10. [18]

    Waves–I 30 23 20 (87%)

  11. [19]

    Waves–II 35 29 22 (76%)

  12. [20]

    Temperature, Heat, and the First Law of Thermodynamics 33 23 22 (96%)

  13. [21]

    The Kinetic Theory of Gases 32 25 23 (92%)

  14. [22]

    Tak- ing the square root: v = √ 3.335 × 1014 ≈ 5.78 × 105 m/s

    Entropy and the Second Law of Thermodynamics 24 16 14 (88%) 629 408 384 (94%) Ch. 13, Problem 41 : Two neutron stars are sep- arated by a distance of 1.0 × 1010 m. They each have a mass of 1.0 × 1030 kg and a radius of 1.0×105 m. They are initially at rest with respect to each...

  15. [23]

    D. H. Jonassen, Designing Research-Based Instruction for Story Problems, Educational Psychology Review 15, 267 (2003)

  16. [24]

    L. Ding, N. Reay, A. Lee, and L. Bao, Exploring the role of conceptual scaffolding in solving synthesis problems, Phys. Rev. ST Phys. Educ. Res. 7 (2011)

  17. [25]

    K. D. Wang, E. Burkholder, C. Wieman, S. Salehi, and N. Haber, Examining the potential and pitfalls of ChatGPT in science and engineering problem-solving, Frontiers in Educa- tion 8 (2024)

  18. [26]

    Kieser and P

    F. Kieser and P. Wulff, Using large language models to probe cognitive constructs, augment data, and design instructional materials, in Machine Learning in Educational Sciences: Ap- proaches, Applications and Advances , edited by M. S. Khine (Springer Nature Singapore, Singapo...

  19. [27]

    B. A. Huberman and T. Hogg, Phase transitions in artificial intelligence systems, Artif. Intell. 33, 155 (1987)

  20. [28]

    Bengio, Y

    Y . Bengio, Y . Lecun, and G. Hinton, Deep learning for AI, Commun. ACM 64, 58 (2021)

  21. [29]

    OpenAI, Introducing ChatGPT, https://openai.com/index/ chatgpt/ (2022)

  22. [30]

    J. Yang, Z. Wang, Y . Lin, and Z. Zhao, Problematic To- kens: Tokenizer Bias in Large Language Models (2024), arXiv:2406.11214 [cs.CL]

  23. [31]

    Glazer, E

    E. Glazer, E. Erdil, T. Besiroglu, D. Chicharro, E. Chen, A. Gunning, C. F. Olsson, J.-S. Denain, A. Ho, E. de Oliveira Santos, O. Järviniemi, M. Barnett, R. San- dler, M. Vrzala, J. Sevilla, Q. Ren, E. Pratt, L. Levine, G. Barkley, N. Stewart, B. Grechuk, T. Grechuk, S. V . E...

  24. [32]

    D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y . Pang, J. Di- rani, J. Michael, and S. R. Bowman, GPQA: A Graduate-Level Google-Proof Q&A Benchmark (2023), arXiv:2311.12022 [cs.AI]

  25. [33]

    OpenAI, Reasoning models: Explore advanced reason- ing and problem-solving models, https://platform.openai.com/ docs/guides/reasoning?api-mode=responses (n.d.), Accessed: May 31, 2025

  26. [34]

    OpenAI, Introducing OpenAI o1-preview, https://openai.com/ index/introducing-openai-o1-preview/ (2024)

  27. [35]

    OpenAI, OpenAI o3-mini: Pushing the frontier of cost- effective reasoning, https://openai.com/index/openai-o3-mini/ (2025)

  28. [36]

    Halliday, R

    D. Halliday, R. Resnick, and J. Walker, Fundamentals of Physics, 12th ed. (John Wiley & Sons, Inc, 2022)

  29. [37]

    OpenAI, Reasoning models: Advice on prompt- ing, https://platform.openai.com/docs/guides/reasoning? api-mode=responses#advice-on-prompting (n.d.), Accessed: June 1, 2025

  30. [38]

    Mialon, R

    G. Mialon, R. Dessì, M. Lomeli, C. Nalmpantis, R. Pasunuru, R. Raileanu, B. Rozière, T. Schick, J. Dwivedi-Yu, A. Celiky- ilmaz, E. Grave, Y . LeCun, and T. Scialom, Augmented Lan- guage Models: a Survey, Transactions on Machine Learning Research (2023)

  31. [39]

    R. K. Atkinson, S. J. Derry, A. Renkl, and D. Wortham, Learn- ing from Examples: Instructional Principles from the Worked Examples Research, Review of Educational Research 70, 181 (2000)

  32. [40]

    Sweller, Cognitive Load During Problem Solving: Effects on Learning, Cognitive Science 12, 257 (1988)

    J. Sweller, Cognitive Load During Problem Solving: Effects on Learning, Cognitive Science 12, 257 (1988)

  33. [41]

    M. T. Chi, M. Bassok, M. W. Lewis, P. Reimann, and R. Glaser, Self-Explanations: How Students Study and Use Examples in Learning to Solve Problems, Cognitive Science13, 145 (1989)

  34. [42]

    Ouyang, J

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schul- man, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe, Training language models to follow instruction...

  35. [43]

    com/index/introducing-o3-and-o4-mini/ (2025)

    OpenAI, Introducing OpenAI o3 and o4-mini, https://openai. com/index/introducing-o3-and-o4-mini/ (2025)

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.