REVIEW 4 major objections 4 minor 3 cited by
GPT-4o can score introductory-physics argument essays in line with multiple-choice performance, and students find its feedback useful and accurate.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Students who answered a physics multiple-choice question correctly received higher GPT-4o essay scores, and most students rated the AI feedback as mostly useful and mostly accurate.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection A modest, honest exploratory study: student perception data are solid, but the 'LLM scoring works' claim rests on an unvalidated MC-correctness proxy. the 4 major comments →
Students' Perceptions to a Large Language Model's Generated Feedback and Scores of Argumentation Essays
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The central claim is that a large language model can produce scores and feedback for students' written physics arguments that behave the way a useful grader's would: on the hard quiz the LLM gave students with correct multiple-choice answers an average 2.14/5 versus 1.12/5 for incorrect responders, and on the easy quiz 2.77 versus 2.53, both differences statistically significant. In the survey, the majority of students selected 'mostly useful' and 'mostly accurate' on Likert scales. The authors take this as evidence that the LLM scoring works for both high- and low-difficulty problems and that LLM feedback is a viable path to individualized feedback at scale.
What carries the argument
The machinery is a GPT-4o grading pipeline. Each essay goes to the model with the original problem, an example of an ideal answer, a five-point rubric built on the claim-evidence-reasoning framework, and a prompt that restricts feedback to the physics content in at most 100 words. The rubric allots two to three points to identifying the relevant physics claim and two to three points to appropriate evidence, mirroring the scaffolding students practiced in recitation. The validation instrument is the comparison of LLM scores for students who answered the underlying multiple-choice question correctly versus incorrectly, and the outcome measure is the student survey of perceived usefulness and a
Load-bearing premise
The load-bearing assumption is that getting the multiple-choice question right is a good stand-in for writing a good argument, so an LLM that scores correct-answer essays higher counts as 'working'—but the paper never checks LLM scores against human expert scores.
What would settle it
Score the same essays with two independent human raters using the same five-point CER rubric and compare their scores to GPT-4o's, for example with quadratic weighted kappa. If human–LLM agreement is below roughly 0.5, or if human raters agree with each other far more than with the LLM, the claim that the LLM scoring works is falsified. A second check: feed the LLM essays that contain no physics content but are padded with CER-sounding sentences; if those score at or above the average student essay, the scores are not tracking argument quality.
If this is right
- Large-enrollment physics courses could give every student individualized written feedback on argumentation essays without additional graders.
- LLM scores carry signal about essay quality, at least on hard problems, since they separate correct and incorrect responders with a large effect size.
- Feedback reaches struggling students too: students who answered incorrectly rated it as useful and accurate as those who answered correctly.
- Moving from delayed end-of-semester feedback to feedback delivered within days of each quiz is a natural next deployment step.
- The same prompt-and-rubric recipe can be reused across quizzes and, with modest changes, across different physics topics.
Where Pith is reading between the lines
- The paper's validation shortcut, comparing LLM scores to multiple-choice correctness, is never checked against human expert scoring, so an obvious extension is to compare LLM and human rubric scores on the same essays.
- Because the survey measures perception, not learning, a future study should measure whether students who act on the feedback improve their later argumentation essays.
- A possible confound the authors do not test is essay length or verbosity: if the LLM implicitly rewards longer essays, the score difference could partly reflect effort rather than argument quality.
- The approach could be tested across other STEM argumentation genres, such as lab conclusions or design justifications, to see whether the CER rubric transfers.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript reports an exploratory classroom study in which GPT-4o generated rubric-based scores and delayed feedback for students' words-only CER-style argumentation essays in an introductory calculus-based physics course, for two quizzes (Quiz 08 and Quiz 09). The authors compare average LLM scores for students who answered the associated multiple-choice question correctly versus incorrectly (Table I) and report Likert ratings of perceived usefulness and accuracy of the LLM feedback. They conclude that the LLM scoring 'works' for both a high-difficulty and a low-difficulty problem, and that students perceive the feedback as mostly useful and accurate, suggesting that LLM-based feedback is a viable option in large-enrollment courses.
Significance. If the conclusions were fully supported, the paper would be a useful proof-of-concept for scalable, individualized feedback in large introductory physics courses. The study has real strengths: it uses a genuine classroom setting, has large score samples (N=730 and N=565), uses simple and transparent statistical analyses, and explicitly acknowledges several limitations, including the delayed timing of feedback and the absence of a human-grader comparison. The perception data are direct and relevant. However, the score-validity evidence is indirect, because it relies entirely on multiple-choice correctness as a proxy for written-argument quality, and survey response rates are not reported. The contribution is therefore suggestive rather than confirmatory.
major comments (4)
- [Section IV, Table I and Section V] The claim that 'LLM scored students as expected... which indicates that the LLM scoring works for both high and low difficulty problems' is not supported by the evidence presented. The only external criterion used is whether the student answered the associated multiple-choice question correctly. The MC item tests selection of a final answer, while the LLM scores a words-only CER argument; these are different constructs. The observed score gap could be driven by essay length, keyword usage, writing fluency, or general course ability rather than argumentation quality. No human-expert scores are reported, and the paper itself lists human-grader comparison as future work (Section V). The wording should be restrained to 'scores differed by MC correctness' unless criterion validity is established.
- [Table I, Quiz 09 row; Section IV] The text states 'A very small p-value and large Cohen’s d' for Quiz 09, but Table I reports d = 0.22, which is conventionally a small effect size. This is not merely a wording issue: the small effect undermines the claim that LLM scoring 'works' equally for low-difficulty problems. The authors should report the effect size accurately and interpret the Quiz 09 result as weak evidence of a scoring difference, not as confirmation.
- [Section III.B] No reliability information is given for the LLM-generated scores. The API temperature was set to 0.80, so the model is nondeterministic, yet it appears each essay was scored only once. Without repeating the scoring or reporting a stability metric (e.g., agreement across runs), the reported differences in Table I may partly reflect prompt/temperature noise. This matters for the conclusion that LLM scoring 'works.'
- [Section III.C and Figure 1] The survey response rate is not reported. Students received bonus points for completing the delayed post-survey, but the manuscript does not state how many of the 730 (Quiz 08) or 565 (Quiz 09) students actually responded, nor does Figure 1 show sample sizes. If response rates were low, self-selection could bias the perception ratings, which are central to the paper's positive conclusion about student perceptions. Please report response counts and, ideally, compare respondents to nonrespondents.
minor comments (4)
- [Table I] Typographical and table-formatting issues: 'Correct A verage' is missing a space, and the column headers could be more readable. Adding sample sizes per subgroup (correct/incorrect) would also help.
- [Section IV] The sentence 'Due to the very small p-value and large Cohen’s d' appears in the Quiz 08 paragraph and again with 'large Cohen’s d' in the Quiz 09 paragraph. The Quiz 09 wording contradicts the reported d=0.22 and should be corrected.
- [Figure 1] The panels are described only in the caption as (a) and (b) etc., but the axis labels, legend, sample sizes, and exact Likert anchors are not visible in the text. Please include these details for readability.
- [Section III.C] The manuscript states that one question was a six-point Likert scale and another a four-point Likert scale, but the exact question wording is not provided. Reporting the item stems would strengthen reproducibility.
Circularity Check
No significant circularity; LLM scores are compared to an external criterion (MC correctness) and perception ratings are self-reported survey data.
full rationale
The paper's central empirical claims are (1) LLM-generated scores differ between students who answered the underlying multiple-choice question correctly versus incorrectly, and (2) students mostly perceived the LLM feedback as useful and accurate. Neither claim reduces to the paper's inputs by construction. The LLM scores were generated by prompting GPT-4o with a rubric, problem statement, and an ideal example; the multiple-choice correctness variable was not fed into the scoring prompt, so the observed score differences are an external comparison rather than a fitted parameter being renamed a prediction. The rubric design choices may influence scores, but they do not force the reported t-test results or effect sizes. The perception data are self-reported Likert ratings and are not derived from the scoring model. The limitations section openly acknowledges the absence of human-grader comparison and delayed feedback, and lists human-grader comparison as future work; this is a validity limitation, not an indication of circularity. Self-citations (e.g., Rebello et al. for argumentation effectiveness, Allen et al. for the scaffolding implementation) appear only in background or methodological context and are not load-bearing for the paper's main conclusions. No circular step could be identified from the text.
Axiom & Free-Parameter Ledger
free parameters (1)
- GPT-4o temperature =
0.80
axioms (4)
- domain assumption Students who answer the MC question correctly tend to write better scientific arguments.
- domain assumption The CER (claim, evidence, reasoning) framework is an appropriate and valid model for scientific argumentation in this context.
- domain assumption The survey respondents are representative of the full class.
- standard math The t-test assumptions (independence, approximate normality) hold for the LLM scores.
Cite this review
Pith. "Pith review of Students' Perceptions to a Large Language Model's Generated Feedback and Scores of Argumentation Essays." pith.science (2026). https://pith.science/paper/FA4TB6RU
@misc{pith2026250814759,
author = {Pith},
title = {Pith review of: Students' Perceptions to a Large Language Model's Generated Feedback and Scores of Argumentation Essays},
year = {2026},
howpublished = {\url{https://pith.science/paper/FA4TB6RU}},
note = {Machine review of arXiv:2508.14759}
}
read the original abstract
Students in introductory physics courses often rely on ineffective strategies, focusing on final answers rather than understanding underlying principles. Integrating scientific argumentation into problem-solving fosters critical thinking and links conceptual knowledge with practical application. By facilitating learners to articulate their scientific arguments for solving problems, and by providing real-time feedback on students' strategies, we aim to enable students to develop superior problem-solving skills. Providing timely, individualized feedback to students in large-enrollment physics courses remains a challenge. Recent advances in Artificial Intelligence (AI) offer promising solutions. This study investigates the potential of AI-generated feedback on students' written scientific arguments in an introductory physics class. Using Open AI's GPT-4o, we provided delayed feedback on student written scientific arguments and surveyed them about the perceived usefulness and accuracy of this feedback. Our findings offer insights into the viability of implementing real-time AI feedback to enhance students' problem-solving and metacognitive skills in large-enrollment classrooms.
Figures
Forward citations
Cited by 3 Pith papers
-
Students' Epistemological Beliefs and their Chatbot Preferences in AI-mediated Physics Learning
Students who chose a combination chatbot (guided inquiry then answers) scored slightly higher on an epistemological-beliefs survey than answer-preference students, but the difference was not robust to multiple-testing...
-
Feedback That Clicks: Introductory Physics Students' Valued Features in AI Feedback Generated From Self-Crafted and Engineered Prompts
Introductory physics students rarely use prompt engineering on their own, but they rate AI feedback as most useful when the prompt explicitly requests evaluation, the correct answer, and improvement suggestions.
-
Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering
LLM graders achieve substantial human agreement on math and science MCAS items but vary on ELA, performing best as sources of formative narrative feedback rather than summative numerical scores.
Reference graph
Works this paper leans on
-
[1]
National Research Council, Next Generation Science Stan- dards: For States, By States (The National Academies Press, Washington, DC, 2013)
work page 2013
-
[2]
W. J. Leonard, R. J. Dufresne, and J. P. Mestre, Using quali- tative problem-solving strategies to highlight the role of con- ceptual knowledge in solving problems, American Journal of Physics 64, 1495 (1996)
work page 1996
-
[3]
D. P. Maloney, An overview of physics education research on problem solving, Getting started in PER 2, 1 (2011)
work page 2011
-
[4]
R. J. Dufresne, W. J. Gerace, and W. J. Leonard, Solv- ing physics problems with multiple representations, Physics Teacher 35, 270 (1997)
work page 1997
-
[5]
Tuminaro and E
J. Tuminaro and E. F. Redish, Elements of a cognitive model of physics problem solving: Epistemic games, Phys. Rev. ST Phys. Educ. Res. 3, 020101 (2007)
2007
-
[6]
C. M. Rebello, Using a hybrid of argumentation and problem solving prompts to facilitate undergraduates’ problem solving performance and confidence, in The 13th Conference of the European Science Education Research Association (ESERA) (2019)
work page 2019
-
[7]
E. A. Siverling, T. J. Moore, E. Suazo-Flores, C. A. Mathis, and S. S. Guzey, What initiates evidence-based reasoning?: Situ- ations that prompt students to support their design ideas and decisions, Journal of Engineering Education 110, 294 (2021)
work page 2021
-
[8]
G. Kortemeyer, J. Nöhl, and D. Onishchuk, Grading assistance for a handwritten thermodynamics exam using artificial intelli- gence: An exploratory study, Phys. Rev. Phys. Educ. Res. 20, 020144 (2024)
work page 2024
-
[9]
O. Henkel, A. Boxer, L. Hills, and B. Roberts, Can large lan- guage models make the grade? an empirical study evaluating llms ability to mark short answer questions in k-12 education (2024), arXiv:2405.02985 [cs.CL]
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[10]
Z. Chen and T. Wan, Achieving human level partial credit grad- ing of written responses to physics conceptual question using gpt-3.5 with only prompt engineering, in Physics Education Research Conference 2024 , PER Conference (Boston, MA,
work page 2024
-
[11]
L. K. Berland and B. J. Reiser, Making sense of argumentation and explanation, Science education 93, 26 (2009)
work page 2009
-
[12]
D. Kuhn, Science as argument: Implications for teaching and learning scientific thinking, Science education 77, 319 (1993)
work page 1993
-
[13]
L. K. Berland and K. L. McNeill, A learning progression for scientific argumentation: Understanding student work and de- signing supportive instructional contexts, Science Education 94, 765 (2010)
work page 2010
-
[14]
Kuhn, Teaching and learning science as argument, Science Education 94, 810 (2010)
D. Kuhn, Teaching and learning science as argument, Science Education 94, 810 (2010)
work page 2010
-
[15]
you’re going to want to find out which and prove it
E. A. Forman, J. Larreamendy-Joerns, M. K. Stein, and C. A. Brown, “you’re going to want to find out which and prove it”: Collective argumentation in a mathematics classroom, Learn- ing and instruction 8, 527 (1998)
work page 1998
-
[16]
M. P. Jiménez-Aleixandre, A. Bugallo Rodríguez, and R. A. Duschl, “doing the lesson” or “doing science”: Argument in high school genetics, Science education 84, 757 (2000)
work page 2000
-
[17]
D. H. Jonassen and B. Kim, Arguing to learn and learning to argue: Design justifications and guidelines, Educational Tech- nology Research and Development 58, 439 (2010)
work page 2010
- [18]
- [19]
-
[20]
A. Christodoulou and J. Osborne, The science classroom as a site of epistemic talk: A case study of a teacher’s attempts to teach science based on argument, Journal of Research in Sci- ence Teaching 51, 1275 (2014)
work page 2014
-
[21]
K. L. McNeill and D. S. Pimentel, Scientific discourse in three urban classrooms: The role of the teacher in engaging high school students in argumentation, Science Education 94, 203 (2010)
work page 2010
-
[22]
S. Schworm and A. Renkl, Learning argumentation skills through the use of prompts for self-explaining examples., Jour- nal of Educational Psychology 99, 285 (2007)
work page 2007
-
[23]
W. N. Wampler, The relationship between students’ problem solving frames and epistemological beliefs , Ph.D. thesis, Pur- due University (2013)
work page 2013
-
[24]
C. M. Rebello, Scaffolding evidence-based reasoning in a tech- nology supported engineering design activity, inThe 13th Con- ference of the European Science Education Research Associa- tion (ESERA) (2019)
work page 2019
-
[25]
K. L. McNeill and J. S. Krajcik, Supporting grade 5-8 students in constructing explanations in science: The claim, evidence, and reasoning framework for talk and writing., Pearson (2011)
work page 2011
-
[26]
S. E. Toulmin, The uses of argument (Cambridge university press, 2003)
work page 2003
-
[27]
K. L. McNeill and J. Krajcik, Inquiry and scientific explana- tions: Helping students use evidence and reasoning, Science as inquiry in the secondary setting 121, 34 (2008)
work page 2008
-
[28]
J. Wang, Scrutinising the positions of students and teacher en- gaged in argumentation in a high school physics classroom, International Journal of Science Education 42, 25 (2020)
work page 2020
-
[29]
J. Hattie and H. Timperley, The power of feedback, Review of Educational Research 77, 81 (2007)
work page 2007
-
[30]
has shown that feedback is one of the important drivers of learning. Feedback can facilitate improvements in learn- ers’ understanding and skills [31–34] by informing learners about their progress, reinforcing good practice, and moti- vating them to engage in self-regulation [34, 35]. It facil- itates self-assessment and reflection on performance, which c...
-
[31]
M. Henderson, T. Ryan, D. Boud, P. Dawson, M. Phillips, E. Molloy, and P. Mahoney, The usefulness of feedback, Ac- tive Learning in Higher Education 22, 229 (2021)
work page 2021
-
[32]
A. Burgess, C. van Diggele, C. Roberts, and C. Mellis, Feed- back in the clinical setting, BMC Medical Education 20, 1 (2020)
work page 2020
-
[33]
R. Sadler, Beyond feedback: Developing student capability in complex appraisal, Assessment & Evaluation in Higher Educa- tion 35, 535 (2010)
work page 2010
-
[34]
D. Nicol and D. Macfarlane-Dick, Formative assessment and self-regulated learning: A model and seven principles of good feedback practice, Studies in Higher Education31, 199 (2006)
work page 2006
-
[35]
A. Burgess and C. Mellis, Receiving feedback from peers: medical students’ perceptions, Clinical Teacher12, 245 (2015)
work page 2015
-
[36]
L. A. Shepard, The role of assessment in a learning culture, Educational Researcher 29, 4 (2000)
2000
-
[37]
A. N. Kluger and A. DeNisi, The effects of feedback interven- tions on performance: A historical review, a meta-analysis, and a preliminary feedback intervention theory, Psychological Bul- letin 119, 254 (1996)
work page 1996
-
[38]
B. Wisniewski, K. Zierer, and J. Hattie, The power of feedback revisited: A meta-analysis of educational feedback research, Frontiers in Psychology 10, 3087 (2020)
work page 2020
- [39]
- [40]
-
[41]
D. R. Ferris, The influence of teacher commentary on student revision, TESOL Quarterly 31, 315 (1997)
work page 1997
-
[42]
W. T. Branch and A. Paranjape, Feedback and reflection: Teaching methods for clinical settings, Academic Medicine77, 1185 (2002)
work page 2002
-
[43]
S. Hepplestone and G. Chikwa, Understanding how students process and use feedback to support their learning, Practitioner Research in Higher Education 8, 41 (2014)
work page 2014
- [44]
-
[45]
T. Ryan, M. Henderson, and M. Phillips, Feedback modes mat- ter: Comparing student perceptions of digital and non-digital feedback modes in higher education, British Journal of Educa- tional Technology 50, 1507 (2019)
work page 2019
-
[46]
N. E. Winstone, R. A. Nash, M. Parker, and J. R. and, Support- ing learners’ agentic engagement with feedback: A systematic review and a taxonomy of recipience processes, Educational Psychologist 52, 17 (2017)
work page 2017
-
[47]
K. Court, Tutor feedback on draft essays: Developing students’ academic writing and subject knowledge, Journal of Further and Higher Education 38, 327 (2014)
work page 2014
-
[48]
D. Boud and E. Molloy, Feedback in Higher and Professional Education: Understanding it and doing it well (Routledge, 2013)
work page 2013
-
[49]
Orlando, How to effectively assess online learning (Magna Publications, 2011)
J. Orlando, How to effectively assess online learning (Magna Publications, 2011)
work page 2011
-
[50]
E. Pitt and L. N. and, ‘now that’s the feedback i want!’ stu- dents’ reactions to feedback on graded work and what they do with it, Assessment & Evaluation in Higher Education 42, 499 (2017)
work page 2017
-
[51]
C. Furnborough and M. T. and, Adult beginner distance lan- guage learner perceptions and use of assignment feedback, Dis- tance Education 30, 399 (2009)
work page 2009
-
[52]
T. D. Wolsey, E-feedback: An exploratory study of using email to provide feedback to students, Journal of Writing Assessment 4, 1 (2008)
work page 2008
-
[53]
K. L. McNeill and D. M. Martin, Claims, evidence, and rea- soning, Science and Children 48, 52 (2011)
work page 2011
-
[54]
M. Ortiz-Rodríguez, R. W. Telg, T. Irani, T. G. Roberts, and E. Rhoades, College students’perceptions of quality in distance education the importance of communication., Quarterly Re- view of Distance Education 6 (2005)
work page 2005
-
[55]
using its application programming interface (API). For both quizzes, the LLM was given the prompt: You are an ed- ucator. Your goal is to provide useful feedback on the physics aspect of the essay. Your feedback should be no more than 100 words long. Focus only on the physics ideas and con- cepts. Do not include salutations. Do not comment on the grammar ...
- [56]
-
[57]
OpenAI, Gpt-4o (2024), large multimodal model
work page 2024
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.