Pith. sign in

REVIEW 3 major objections 4 minor 113 references

The paper reports that the o4-mini model solves about 90% of standard introductory physics problems, but accuracy drops from 96% on text-only problems to 79% when images must be interpreted, and declines with problem difficulty.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 02:28 UTC pith:K2FBRLGI

load-bearing objection A useful large benchmark of o4-mini on Halliday & Resnick, with plausible modality and difficulty effects, but sloppy reporting (effort contradiction, odds/probability mix-up) and a grading-validation gap limit the current credibility of the exact percentages. the 3 major comments →

arxiv 2607.14303 v1 pith:K2FBRLGI submitted 2026-07-15 physics.ed-ph cs.AI

Assessing AI in Introductory Physics Problem Solving

classification physics.ed-ph cs.AI PACS 01.40.-d
keywords AI problem solvingphysics educationlarge language modelsmultimodal reasoningproblem difficultyo4-minitextbook problemsphysics education research
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper attempts to quantify how well a state-of-the-art reasoning model, o4-mini, handles traditional end-of-chapter problems from a standard introductory physics textbook. It argues that the model solves the large majority of such problems, about 90% overall, but that this capability is uneven: text-only problems are solved at 96%, while problems requiring joint interpretation of text and images fall to 79%. It also claims accuracy decreases with difficulty, from 94% easy to 84% hard, and that this difficulty gradient is statistically significant. A sympathetic reader would care because it maps the current frontier of AI physics problem solving and identifies modality as the main remaining bottleneck.

Core claim

The model o4-mini attains an overall accuracy of 0.90±0.03 on 1,203 odd-numbered problems from a standard introductory physics textbook, with a marked modality gap (0.96 text-only vs 0.79 text+image) and a monotonic decline across difficulty levels (0.94 easy, 0.88 medium, 0.84 hard). Logistic regression and generalized estimating equations confirm the difficulty effect, with odds of a correct response reduced by 51% for medium and 64% for hard relative to easy. The model's output effort roughly doubles for medium and hard problems. These results indicate that a current reasoning model can solve most standard introductory physics problems, but performance remains strongly constrained by visu

What carries the argument

The evaluation pipeline: 1,203 problem-answer pairs extracted from the textbook, formatted in LaTeX with unified units, images as PNG screenshots; each problem run five times through o4-mini; solutions scored right/wrong by a large language model (GPT-5) against the textbook answer key, with a 600-item human audit of the grader. Statistical analysis uses logistic regression and generalized estimating equations to assess the difficulty-accuracy relationship.

Load-bearing premise

Every accuracy figure depends on the AI grader being reliable: only 600 of 6,015 scores were spot-checked, and the small error rate's direction is unknown, so if the grader systematically credits flawed solutions the 90%, 96%, and 79% numbers are too high.

What would settle it

Take a random sample of, say, 300 problems from the dataset, have two independent human physics instructors score the o4-mini solutions blind to the AI grader's scores, and compare the resulting accuracy estimates to the paper's 0.90, 0.96, and 0.79 figures; a substantial disagreement would overturn the central claim.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Students using such models on text-only homework can expect mostly correct solutions; image-based problems are much less reliable.
  • The difficulty gradient implies AI could serve as an objective difficulty classifier for problem banks, as the paper itself suggests.
  • Accuracy is stable across mechanics, electromagnetism, quantum theory, and other topics, so the model's weakness is not tied to specific content.
  • The modality gap suggests that improving multimodal grounding is the next barrier for AI physics problem solving.
  • Teachers and students should treat the high overall accuracy with caution because it masks systematic weaknesses on visual and hard problems.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • One implication the paper leaves implicit is that the proposed 'computational difficulty measure' is only as meaningful as the textbook's own difficulty labels; a natural test is to compare model accuracy against student success rates on the same problems.
  • Because only odd-numbered problems with provided answers were used, the 90% figure excludes open-ended, drawing, or explanation problems; the result likely overstates the model's ability on the full range of textbook tasks.
  • The direction of the 14 grader errors in the 600-item audit is not reported; if most were cases where GPT-5 marked an incorrect solution as correct, the true accuracy could be meaningfully lower than 90%.
  • Re-running this exact 1,203-problem set on successor reasoning models would yield a direct longitudinal measure of whether the modality gap narrows over time.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The manuscript evaluates OpenAI's o4-mini on N=1203 odd-numbered end-of-chapter problems from Halliday and Resnick's Fundamentals of Physics, with each problem solved five times (6,015 runs). The authors report overall accuracy 0.90±0.03, text-only accuracy 0.96±0.03 versus 0.79±0.04 for problems requiring images, and accuracy declining from 0.94 (easy) to 0.88 (medium) to 0.84 (hard). They test the difficulty trend with logistic regression and generalized estimating equations, and also report model effort in output tokens. The conclusions are that current reasoning LLMs can solve most standard introductory problems but remain materially weaker on multimodal and harder problems.

Significance. If valid, this is a useful, large-scale benchmark of a reasoning model on a canonical introductory textbook: the design includes repeated runs per problem, explicit model checkpoint and prompting details, an external answer-key ground truth, and clustered regression for the difficulty analysis. The reported 17-point text-versus-image gap and the difficulty gradient are plausible and of practical interest to PER and to AI-in-education researchers. However, the central accuracy estimates depend on LLM grading whose validation is underreported, and the paper contains unresolved arithmetic inconsistencies in the dataset counts. Because the scoring disagreement rate is comparable to the reported difficulty steps, the grading issue must be resolved before the quantitative claims can be accepted as stated.

major comments (3)
  1. [Section II (AI-based scoring)] All headline figures (0.90 overall, 0.96 text, 0.79 image, 0.94/0.88/0.84 by difficulty) are produced by GPT-5 grading. The 600-run validation is not described (who re-scored, using what rubric?), is not stratified by modality or difficulty, and only reports 14 disagreements without a confusion matrix or error direction. Absorbing 0.023 as a 'systematic error' into the error bars does not correct a possible unidirectional bias. Since the modality gap (0.17) and adjacent difficulty steps (0.06, 0.04) are of the same scale as the raw disagreement rate, the authors must report false-positive and false-negative rates by condition and either correct the point estimates or provide a quantitative bound on grading bias.
  2. [Section II, Fig. 1 and Tables I, II, VI] The dataset arithmetic is inconsistent. Fig. 1 gives 1,225 problem–answer pairs and then N=1,203; summing the two volume tables yields 1,205 problems (410+380 text, 208+207 image-based) with difficulty totals 507/608/90, while Table VI reports 507/607/89. Additionally, the text says 411 images were extracted, whereas Fig. 1 counts 411 figures + 12 tables = 423 PNG images, with Nimg=415. Because every accuracy and regression uses these denominators, the exclusion rules need to be stated explicitly and the tables corrected accordingly.
  3. [Appendix B and Conclusion] The text interprets odds ratios as probability reductions: 'the probability of solving a Medium-level problem is smaller ... by 0.51' and by 0.27 for Hard. These are reductions in odds, not in probability: an odds ratio of 0.49 does not imply a 0.51 decrease in probability. This overstates the effect sizes. The authors should restate the results in terms of odds ratios with confidence intervals, or convert to predicted probabilities at a specified baseline.
minor comments (4)
  1. [Section III and Conclusion] Figure 4 discussion says the output size 'decreases correspondingly' as difficulty increases, while the Conclusion says the effort 'almost doubled' for Medium and Hard problems. These statements contradict each other and should be reconciled.
  2. [Tables III–IV captions] The captions do not specify whether the bottom-row averages are unweighted means of chapter accuracies or pooled accuracies over all problems in the volume. Please clarify.
  3. [Reference [29]] The author string 'D. E. Trowbridge1981' appears malformed; it should be 'D. E. Trowbridge and L. C. McDermott' (or similar).
  4. [Overall] No data or code availability statement is provided. Given the novelty of using an LLM grader at scale, the problem-level scores and scoring prompts should be released to make the quantitative claims checkable.

Circularity Check

0 steps flagged

No significant circularity: accuracies are measured against an external textbook answer key, and no prediction is constructed from a fitted input.

full rationale

The paper is an empirical benchmark, not a derivation. The central claims—overall accuracy 0.90, text-only 0.96 vs. image-based 0.79, and difficulty gradient—are defined by Equation (2) as averages of binary scores assigned against Halliday & Resnick's textbook answers. The textbook answer key is external to the model and to the authors, so the accuracy numbers are not fitted parameters renamed as predictions. The difficulty analysis regresses model scores on the textbook authors' difficulty labels; the predictor is independent of the outcome, so the regression is a measurement of a correlation, not a self-defined prediction. The only potentially load-bearing auxiliary element is the use of GPT-5 as grader, but grading is anchored to the external answer key and a 600-sample validation check with 14 disagreements is reported; this converts the concern into a grader-reliability / measurement-validity issue, not a circularity of construction. There are no self-citations by the authors, no self-referential uniqueness theorem, and no ansatz smuggled in via citation. The paper's own limitation statement about assuming the answer key has no errors is an explicit assumption, not a circular step. Accordingly, no circular step can be exhibited by quoting equations that reduce to their inputs, and the appropriate score is 0.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 0 invented entities

The paper introduces no free physics parameters, axioms, or invented entities: it measures an existing model on existing problems. The ledger records the measurement assumptions (answer-key ground truth, AI grading, difficulty labels) and the fitted regression coefficients, because the difficulty claim is carried by those fits.

free parameters (3)
  • β1 (Medium vs Easy difficulty coefficient) = -0.7072
    Fitted in logistic regression (Table VIII) to the 1,203 problem outcomes; carries the claim that accuracy drops with difficulty.
  • β2 (Hard vs Easy difficulty coefficient) = -1.0294
    Fitted in logistic regression (Table VIII); also used to compute the Medium→Hard odds ratio 0.73.
  • Systematic scoring error = 0.023
    Set by the 14/600 wrong GPT-5 grades; assumed fixed and symmetric and added to the statistical error in reported accuracies.
axioms (5)
  • domain assumption Textbook answer key is error-free and serves as ground truth.
    Section II states this explicitly; any key errors would shift all accuracy estimates.
  • ad hoc to paper GPT-5 can reliably grade o4-mini solutions against the answer key, with error rate fixed at 0.023.
    Section II; only 600 of 6015 scores verified and no direction of grader bias reported.
  • domain assumption Difficulty labels from Halliday & Resnick (Easy, Medium, Hard) are meaningful ordinal categories.
    Labels are taken from the textbook and used as the predictor; no independent difficulty measure.
  • standard math In tests 1–2 difficulty is equally spaced on the logit scale (1,2,3).
    Appendix B; relaxed in test 3, but the headline 'drop by 0.06 then 0.04' interpretation uses the raw grouping.
  • domain assumption The textbook's odd-numbered problems with answers are representative of the curriculum.
    Section II; even-numbered problems and draw/sketch problems were excluded; if these differ, the benchmark is biased.

pith-pipeline@v1.3.0-alltime-deepseek · 12189 in / 13577 out tokens · 137313 ms · 2026-08-02T02:28:54.136605+00:00 · methodology

0 comments
read the original abstract

Reasoning or inference-scaling models are the new generation of Large Language Models (LLMs) capable of complex problem solving. To investigate their problem-solving capability in physics, we evaluated model o4-mini by OpenAI on solving traditional, end-of-chapter problems from Halliday and Resnick's "Fundamentals of Physics," spanning core topics in the undergraduate physics curriculum. Performance was analyzed across modality and problem difficulty. The model solved the problems with overall accuracy of about 90%, but performance depended strongly on representation: accuracy was much higher on text-only problems (96%) than on problems requiring coordinated interpretation of text and images (79%). Accuracy also declined significantly as the problem difficulty increased from low to medium to high. These results show that state-of-the-art LLMs can solve much of the standard introductory physics problems, but that their performance remains uneven and constrained by problem modality and problem difficulty.

Figures

Figures reproduced from arXiv: 2607.14303 by Amir Bralin, N. Sanjay Rebello.

Figure 1
Figure 1. Figure 1: FIG. 1. Process of converting textbook end-of-chapter problems into AI model input [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2. Accuracy results by physics topics. [PITH_FULL_IMAGE:figures/full_fig_p014_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: FIG. 3. Accuracy results by problem difficulty. [PITH_FULL_IMAGE:figures/full_fig_p015_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: FIG. 4. Average output tokens by problem difficulty. [PITH_FULL_IMAGE:figures/full_fig_p016_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

113 extracted references · 14 canonical work pages

  1. [1]

    What is the problem-solving capability of AI models on standard topics in the intro- ductory physics curriculum?

  2. [2]

    How does this capability depend on modality—language and vision?

  3. [3]

    o4-mini” by OpenAI was selected due to its design and affordability. Together with a larger model “o3,

    How does this capability depend on different levels of difficulty in the problems solved? An additional property of a model when solving analytical problems is the amount of “effort” it exerts while engaging in this process. This quantity may be defined as the number of tokensproduced by the model in generating the full solution to a given problem. It pro...

  4. [4]

    Measurement 13 2 8 5 2

  5. [5]

    Motion Along a Straight Line 25 10 15 16 4

  6. [6]

    Motion in Two and Three Dimensions 33 8 13 22 6

  7. [7]

    Force and Motion–I 18 16 11 20 3

  8. [8]

    Force and Motion–II 15 14 10 16 3

  9. [9]

    Kinetic Energy and Work 17 9 9 15 2

  10. [10]

    Potential Energy and Conservation of Energy 12 18 9 15 6

  11. [11]

    Center of Mass and Linear Momentum 22 18 13 23 4

  12. [12]

    Rotation 25 9 14 16 4

  13. [13]

    Rolling, Torque, and Angular Momentum 18 17 16 14 5

  14. [14]

    Equilibrium and Elasticity 6 18 10 11 3

  15. [15]

    Gravitation 26 8 18 13 3

  16. [16]

    Oscillations 20 12 17 12 3

  17. [17]

    Waves–I 23 7 13 15 2

  18. [18]

    Waves–II 29 6 16 15 4

  19. [19]

    Temperature, Heat, and the First Law of Thermodynamics 23 10 14 17 2

  20. [20]

    The Kinetic Theory of Gases 26 6 13 17 2

  21. [21]

    The additional errors coming from run variation within each chapter problem were propagated, adding to the total statistical error in our chapter ac- curacy results

    Entropy and the Second Law of Thermodynamics 16 6 8 11 3 410 208 258 299 61 culated by SEMσ chap/√Ni. The additional errors coming from run variation within each chapter problem were propagated, adding to the total statistical error in our chapter ac- curacy results. Finally, to report the overall accuracy for all textbook chapters considered together, we...

  22. [22]

    Coulomb’s Law 11 8 6 10 3

  23. [23]

    Electric Fields 15 13 10 17 1

  24. [24]

    Gauss’ Law 16 12 12 13 3

  25. [25]

    Electric Potential 22 12 12 19 3

  26. [26]

    Capacitance 15 7 11 16 1

  27. [27]

    Current and Resistance 24 3 10 16 1

  28. [28]

    Circuits 13 20 12 18 3

  29. [29]

    Magnetic Fields 23 9 17 14 1

  30. [30]

    Magnetic Fields Due to Currents 12 20 13 16 3

  31. [31]

    Induction and Inductance 19 19 18 17 3

  32. [32]

    Electromagnetic Oscillations and Alternating Current 20 10 15 15

  33. [33]

    Maxwell’s Equations; Magnetism of Matter 16 9 9 15 1

  34. [34]

    Electromagnetic Waves 19 14 15 17 1

  35. [35]

    Interference 13 8 12 18 1

  36. [36]

    Diffraction 29 6 21 14

  37. [37]

    Relativity 26 3 11 17 1

  38. [38]

    Photons and Matter Waves 33 2 10 23 2

  39. [39]

    More About Matter Waves 17 6 9 14

  40. [40]

    RESUL TS Tables III and IV show the scores achieved by model o4-mini on all problems from 40 chapters from Hallidayet al.as graded by model GPT-5

    All About Atoms 25 4 18 11 380 207 249 309 29 III. RESUL TS Tables III and IV show the scores achieved by model o4-mini on all problems from 40 chapters from Hallidayet al.as graded by model GPT-5. The column Text indicates the model accuracy on text-only problems, while column Text + Image indicates the accuracy 10 on problems with images when the model ...

  41. [41]

    Measurement 0.97(0.07) 1.00(0.00) 0.97(0.07)

  42. [42]

    Motion Along a Straight Line 1.00(0.00) 0.46(0.25) 0.85(0.10)

  43. [43]

    Vectors 0.98(0.06) 0.48(0.32) 0.86(0.12)

  44. [44]

    Motion in Two and Three Dimensions 0.98(0.05) 0.75(0.19) 0.94(0.07)

  45. [45]

    Force and Motion–I 0.94(0.08) 0.82(0.17) 0.89(0.11)

  46. [46]

    Force and Motion–II 1.00(0.00) 0.94(0.10) 0.97(0.06)

  47. [47]

    Kinetic Energy and Work 0.99(0.05) 0.78(0.20) 0.92(0.09)

  48. [48]

    Potential Energy and CoE 1.00(0.00) 0.84(0.16) 0.91(0.11)

  49. [49]

    Center of Mass and Linear Momentum 1.00(0.00) 0.82(0.14) 0.92(0.08)

  50. [50]

    Rotation 0.94(0.08) 0.71(0.20) 0.88(0.09)

  51. [51]

    Rolling, Torque, and AM 0.91(0.11) 0.75(0.15) 0.83(0.11)

  52. [52]

    Equilibrium and Elasticity 0.93(0.13) 0.59(0.23) 0.68(0.19)

  53. [53]

    Gravitation 0.99(0.04) 1.00(0.00) 0.99(0.03)

  54. [54]

    Fluids 0.98(0.05) 0.87(0.16) 0.95(0.07)

  55. [55]

    Oscillations 0.98(0.06) 0.97(0.08) 0.98(0.06)

  56. [56]

    Waves–I 0.93(0.09) 0.17(0.19) 0.74(0.12)

  57. [57]

    Waves–II 0.90(0.11) 1.00(0.00) 0.92(0.10)

  58. [58]

    Temperature, Heat, and the FLT 0.97(0.06) 0.84(0.18) 0.93(0.08)

  59. [59]

    The Kinetic Theory of Gases 0.93(0.10) 0.83(0.19) 0.91(0.10)

  60. [60]

    Entropy and the SLT 0.94(0.09) 0.93(0.13) 0.94(0.09) Average:0.96(0.05) 0.79(0.10) 0.90(0.06) 12 TABLE IV. Vol. 2 scores by chapter and modality. Some chapter titles were abbreviated for visual purposes. (AC: Alternating Current, MoM: Magnetism of Matter) Chapter T ext T ext + Image Overall

  61. [61]

    Coulomb’s Law 0.96(0.08) 0.95(0.11) 0.96(0.08)

  62. [62]

    Electric Fields 0.95(0.09) 0.82(0.20) 0.89(0.13)

  63. [63]

    Gauss’ Law 0.94(0.10) 0.80(0.17) 0.88(0.11)

  64. [64]

    Electric Potential 0.97(0.07) 0.92(0.11) 0.95(0.07)

  65. [65]

    Capacitance 1.00(0.00) 0.78(0.20) 0.90(0.11)

  66. [66]

    Current and Resistance 1.00(0.00) 1.00(0.00) 1.00(0.00)

  67. [67]

    Circuits 1.00(0.00) 0.88(0.12) 0.93(0.08)

  68. [68]

    Magnetic Fields 0.94(0.09) 0.87(0.17) 0.92(0.10)

  69. [69]

    Magnetic Fields Due to Currents 0.97(0.08) 0.58(0.22) 0.72(0.17)

  70. [70]

    Induction and Inductance 0.95(0.08) 0.89(0.12) 0.92(0.08)

  71. [71]

    Electromagnetic Oscillations and AC 1.00(0.00) 0.96(0.09) 0.99(0.04)

  72. [72]

    Maxwell’s Equations; MoM 0.90(0.11) 0.76(0.22) 0.85(0.12)

  73. [73]

    Electromagnetic Waves 1.00(0.00) 0.73(0.20) 0.88(0.10)

  74. [74]

    Images 1.00(0.00) 0.67(0.27) 0.89(0.11)

  75. [75]

    Interference 0.97(0.07) 0.86(0.13) 0.90(0.10)

  76. [76]

    Diffraction 0.90(0.10) 0.33(0.23) 0.81(0.11)

  77. [77]

    Relativity 0.95(0.07) 1.00(0.00) 0.96(0.06)

  78. [78]

    Photons and Matter Waves 0.84(0.12) 0.80(0.22) 0.84(0.13)

  79. [79]

    More About Matter Waves 0.98(0.06) 0.83(0.19) 0.94(0.08)

  80. [80]

    it appears that grouping 17 the items in terms of test objectives does not provide novel meaningful insights into the strengths and weaknesses of individual chatbots

    All About Atoms 0.93(0.09) 0.90(0.18) 0.92(0.09) Average:0.95(0.05) 0.81(0.10) 0.90(0.06) 13 TABLE V. Accuracy results by physics topics. #T opic ChaptersN problems Accuracy 1 Mechanics 1–14 434 0.90(0.06) 2 Oscillations and Waves 15–17 96 0.89(0.06) 3 Thermodynamics and Kinetic Theory 18–20 87 0.93(0.06) 4 Electromagnetism 21–32 353 0.91(0.06) 5 Electrom...

Showing first 80 references.