REVIEW 4 major objections 6 minor 1 cited by
Students overwhelmingly prefer AI feedback generated from engineered prompts, especially when the prompt embeds principles of effective feedback.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Introductory physics students rarely use prompt engineering on their own, but they rate AI feedback as most useful when the prompt explicitly requests evaluation, the correct answer, and improvement suggestions.
T0 review reviewed 2026-08-04 challenge →
load-bearing objection Valuable descriptive taxonomy and prompt data, but the headline preference result is undermined by the non-blind design the authors themselves acknowledge. the 4 major comments →
Feedback That Clicks: Introductory Physics Students' Valued Features in AI Feedback Generated From Self-Crafted and Engineered Prompts
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
On its own terms, the paper establishes that when students see three AI feedback versions—one from their self-crafted prompt, one from a prompt built on foundational prompt-engineering techniques, and one that adds explicit principles of effective feedback (feed up, feed back, feed forward)—they overwhelmingly prefer the third. The authors interpret this as evidence that structured prompting matters, and that grounding prompts in education research on feedback improves perceived usefulness. They also document that most student prompts contain only partial prompt-engineering features, and that AI's ability to fill in missing context, while usually correct, sometimes yields feedback that misse
What carries the argument
The comparison vehicle is a set of three prompts built from the same physics task: Feedback A comes from the student's own prompt; Feedback B embeds five foundational prompt-engineering techniques (instructions, clarity and precision, role-based prompting, delimiters, zero-shot framing); Feedback C adds to B an explicit request to highlight the correct answer, strengths and limitations of the student's answer, gaps, a 200-word cap, and three bulleted improvement suggestions—operationalizing the 'feed up / feed back / feed forward' principles of effective feedback. Student justifications are coded into 12 features under four themes.
Load-bearing premise
The central claim rests on the assumption that students' rankings reflect the actual quality of the feedback, even though students saw the prompts that generated each version and may have rated sophistication rather than usefulness.
What would settle it
Conduct the same ranking task with the three feedback texts presented without showing which prompt generated them; if the 61% majority shifts or disappears, the prompt-design preference is an artifact. A second check: have independent physics instructors rate the three feedback versions for accuracy and actionability, and compare those ratings to student rankings.
If this is right
- If the preference ranking reflects actual usefulness, instructors should not assume students' self-crafted prompting is enough; providing structured prompt templates changes which feedback students engage with.
- Feedback design for AI systems should embed principles of effective feedback, not just prompt-engineering tricks.
- The 12-feature taxonomy gives a vocabulary for evaluating AI feedback in physics education and for comparing future studies.
- Because AI can infer context from incomplete prompts, prompt sophistication may matter more for conceptual accuracy than for producing plausible-looking feedback.
- Since each preferred feedback type corresponded to a different feature profile, no single feedback feature satisfies all students.
Where Pith is reading between the lines
- Inference: The ranking may partly reflect a 'sophistication heuristic'—students saw longer, more detailed prompts and inferred the output would be better; a blind design could separate this from content quality.
- Inference: The prompt used for Feedback C explicitly demanded the features students then valued (strengths, gaps, bullets), so the ranking partly measures prompt-output compliance, not necessarily learning gains.
- Inference: A testable extension is to measure whether Feedback C actually improves subsequent performance on a similar physics task, rather than only perceived usefulness.
- Inference: Given the near-absence of role-based prompting in student prompts, novices may benefit more from scaffolds that teach when to assign roles and structure, not just content.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a study of 1,235 introductory physics students who used AI to obtain feedback on a written argument about the slope of a gravitational potential energy vs. distance graph. Students first wrote their own prompt (Feedback A), then used two instructor-provided prompts: one embedding foundational prompt-engineering techniques (Feedback B) and one additionally embedding principles of effective feedback (Feedback C). The authors analyze the sophistication of student-generated prompts, identify a 12-feature taxonomy (grouped into Evaluation, Content, Presentation, Depth) from students' justifications of their preferred feedback, and report preference rankings: 61.4% ranked C most useful, 29.7% B, and 8.9% A. They conclude that students overwhelmingly prefer feedback from structured prompts, especially those combining prompt engineering with principles of effective feedback, and discuss implications for AI feedback design and prompt-engineering instruction.
Significance. If the preference ranking reflected genuine feedback-quality judgments, the result would be practically important for AI-supported formative assessment in large-enrollment physics courses. The study has notable strengths: a large sample (1,235 valid responses), a clearly described task context, transparent emergent-coding procedures with exemplar quotes, and a taxonomy that usefully organizes student-valued feedback features. The paper also documents an interesting phenomenon—AI 'contextual intuitiveness' when student prompts omit context—and gives concrete exemplars. However, the central preference claim is weakened by a non-blind design, an uncontrolled AI platform, and the absence of inferential statistics. The feature taxonomy is valuable as a descriptive contribution, but the causal/interpretive claim that structured prompts produce more useful feedback requires additional support or substantial reframing.
major comments (4)
- [Section VI; Section IV.C.1; Figure 4] The central ranking result (61.4% C vs. 29.7% B vs. 8.9% A) is not secure as evidence about feedback quality. As the authors acknowledge in Section VI, 'by design, students were aware of the prompts used to generate the three types of feedback.' Students saw that Feedback C came from a longer, more elaborate prompt that explicitly asked for the correct answer, strengths/limitations, gaps, a 200-word limit, and three bulleted suggestions. Their rankings—and the justifications in Table IV that cite exactly these features—may reflect a demand characteristic or a prompt-sophistication heuristic rather than perceived usefulness of the feedback itself. A blind or randomized re-ranking, or an analysis that controls for prompt length/visibility, is needed to validate the interpretation. At minimum, the claims should be reframed as preferences for prompt types, not for feedback quality.
- [Section III.A; Section IV.C.1] The AI platform was not controlled: students used 'AI platforms of their choice' (Section III.A). Because different models and versions generate different feedback, the aggregate percentages in Figure 4 mix heterogeneous generation conditions. The paper does not report how many students used which platform, nor whether the C-over-A/B ordering is stable across platforms. Additionally, the percentages are reported without confidence intervals or any test against a chance/null model. Given 1,235 independent rankings, a multinomial confidence interval or a simple exact test would be straightforward and would strengthen the 'overwhelming majority' claim. Without this, the descriptive ranking is suggestive but not statistically grounded.
- [Section IV.B; Table III; Section IV.C.2] There is a partial circularity between the design of Feedback C's prompt and the features students are reported to value. The prompt for C explicitly instructs the AI to 'highlight the correct answer, my provided answer (including strengths and limitations), and gaps... within 200 words... three potential ways through a bulleted list.' It is therefore not surprising that students who preferred C cite Critique, Guidance, Structure, and Conciseness. The paper notes this consistency in Section IV.C.2, but it is not treated as a limitation. The feature taxonomy is still valuable descriptively, but the claim that these are independently 'valued features' of AI feedback needs to be tempered by the fact that the prompt itself engineered these features into the output. A prompt-agnostic version or a feature-blind comparison would disentangle prompt-driven expectations from genuine feedback evalu
- [Section IV.B.1 vs. Table III and Section IV.C.2] There is a direct numerical inconsistency in the reported prevalence of the Evaluation theme: Section IV.B.1 states 'A total of 65.2% of student descriptions evidenced at least one of these features,' while Table III and Section IV.C.2 report 69.4% for the same theme. Since these percentages are a core descriptive result, the correct figure must be identified and used consistently throughout the manuscript.
minor comments (6)
- [Figures 2 and 3] The y-axis labels are not specified. Please state explicitly that the values are percentages of student-generated prompts (or responses), and include the denominator N for each bar. Also clarify how the 6.8% rejected prompts are handled in the denominator.
- [Section III.B] Inter-rater reliability is described only as 'compared and discussed until full agreement was reached' on 25 responses. A formal agreement statistic (e.g., Cohen's kappa) and a description of how the 25 responses were sampled would strengthen the coding reliability argument.
- [Section II.A] The text lists seven foundational techniques but says 'we adopt five by omitting the fifth and the seventh.' Since the sixth (zero-shot) is described as being used without examples, the counting is confusing. Clarify the operationalization of each adopted technique in the prompt templates.
- [Section VI; first paragraph] The sentence beginning 'Finally, Detailed results...' is ungrammatical and appears to be a leftover fragment. Also, the paragraph contains two consecutive 'Finally' constructions; revise for clarity.
- [References] Reference [1] appears to be a preprint by the same authors with a similar title ('From self-crafted to engineered prompts...'). If this is a prior version or companion paper, the relationship should be clarified and the citation should be distinguished from the present manuscript.
- [Section V] The summary bullet 'For a given preferred feedback choice, there was no common second preferred candidate' is not quite accurate given Figure 5: among students preferring B, 90% chose A as second; among those preferring C, 95.5% chose B. The text in Section IV.C.1 already notes these patterns, so the summary should be reconciled.
Circularity Check
No significant circularity; the study is an empirical preference/feature analysis, and the non-blind design is an acknowledged validity limitation rather than a circular derivation.
full rationale
The paper contains no derivation chain in which an output quantity is defined in terms of the input quantity. RQ1 (prompt sophistication) is a descriptive coding result; RQ2 (valued features) is an emergent-coding analysis of open-ended student justifications; RQ3 (preference rankings) is a summary of students' stated rankings. Feedback C's prompt did explicitly instruct the AI to highlight correct answers, strengths/limitations, gaps, a 200-word limit, and three bulleted suggestions (Table I), and the paper later notes that students preferring C emphasized Evaluation and Presentation features 'consistent with the design of the associated prompt' (Section V). This is an alignment between intervention and outcome, not a circular reduction: the features were coded from student justifications, not read off the prompt, and the central claim is about students' perceived usefulness, not an objective prediction. The self-citations (e.g., Ref. [1]) are contextual and not load-bearing; no uniqueness theorem or fitted parameter is invoked. Section VI explicitly acknowledges the non-blind design as a possible bias: 'by design, students were aware of the prompts used to generate the three types of feedback. This may have introduced a bias in their preferences...' That is a legitimate methodological limitation affecting internal validity, but it is not an instance of a result being equivalent to its inputs by construction. Accordingly, no circularity is identified.
Axiom & Free-Parameter Ledger
free parameters (3)
- Dichotomous coding threshold for 'Being Clear and Precise' =
Both components required (explicit task context + explicit argument)
- Operationalized set of prompt engineering techniques =
Five of seven foundational techniques, zero-shot only
- Judgment of 'contextual intuitiveness' =
285 of 293 ambiguous prompts correctly interpreted
axioms (4)
- domain assumption Student self-reported rankings and justifications reflect genuine perceived usefulness of feedback
- domain assumption The AI platforms students chose are representative of current LLMs and produced comparable feedback
- domain assumption Observations from one extra-credit task in one introductory physics course generalize to other contexts
- domain assumption Emergent coding categories capture features of feedback usefulness
Cite this review
Pith. "Pith review of Feedback That Clicks: Introductory Physics Students' Valued Features in AI Feedback Generated From Self-Crafted and Engineered Prompts." pith.science (2026). https://pith.science/paper/22IKEKPQ
@misc{pith2026250908516,
author = {Pith},
title = {Pith review of: Feedback That Clicks: Introductory Physics Students' Valued Features in AI Feedback Generated From Self-Crafted and Engineered Prompts},
year = {2026},
howpublished = {\url{https://pith.science/paper/22IKEKPQ}},
note = {Machine review of arXiv:2509.08516}
}
read the original abstract
Since the advent of GPT-3.5 in 2022, Generative Artificial Intelligence (AI) has shown tremendous potential in STEM education, particularly in providing real-time, customized feedback to students in large-enrollment courses. A crucial skill that mediates effective use of AI is the systematic structuring of natural language instructions to AI models, commonly referred to as prompt engineering. This study has three objectives: (i) to investigate the sophistication of student-generated prompts when seeking feedback from AI on their arguments, (ii) to examine the features that students value in AI-generated feedback, and (iii) to analyze trends in student preferences for feedback generated from self-crafted prompts versus prompts incorporating prompt engineering techniques and principles of effective feedback. Results indicate that student-generated prompts typically reflect only a subset of foundational prompt engineering techniques. Despite this lack of sophistication, such as incomplete descriptions of task context, AI responses demonstrated contextual intuitiveness by accurately inferring context from the overall content of the prompt. We also identified 12 distinct features that students attribute the usefulness of AI-generated feedback, spanning four broader themes: Evaluation, Content, Presentation, and Depth. Finally, results show that students overwhelmingly prefer feedback generated from structured prompts, particularly those combining prompt engineering techniques with principles of effective feedback. Implications of these results such as integrating the principles of effective feedback in design and delivery of feedback through AI systems, and incorporating prompt engineering in introductory physics courses are discussed.
Figures
Forward citations
Cited by 1 Pith paper
-
Students' Epistemological Beliefs and their Chatbot Preferences in AI-mediated Physics Learning
Students who chose a combination chatbot (guided inquiry then answers) scored slightly higher on an epistemological-beliefs survey than answer-preference students, but the difference was not robust to multiple-testing...
Reference graph
Works this paper leans on
-
[1]
This technique involves designing directives that guide the AI model toward the intended and relevant output
Providing instructions. This technique involves designing directives that guide the AI model toward the intended and relevant output. For example, consider the objective of understanding the significance of the resistance R of a substance with respect to temperature ( T ). A prompt such as ‘ Relation between R and T ’ may yield broad and diverse results w...
-
[2]
Explain the meaning of R in science
Being clear and precise . Crafting clear and precise instructions helps the model generate accurate and relevant output. In contrast, general instructions often lead to overly broad or ambiguous responses given the model’s vast training dataset. For example, if the objective is to gain a basic understanding of electrical resistance, the prompt “Explain th...
-
[3]
You are an expert physics instructor tasked with grading student responses
Role-based prompting. The feature of assigning a role to the AI model to yield task-specific output. Within this prompting technique, two subtypes can be distinguished: static role prompting and dynamic role-play prompting. Static role prompting involves assigning a fixed role to the model to enhance contextual accuracy and task-specific performance. Expe...
-
[4]
When the task is complex and the prompt has multiple components, delimiters such as triple quotes or custom symbols (*) are used to separate the components
Use of delimiters. When the task is complex and the prompt has multiple components, delimiters such as triple quotes or custom symbols (*) are used to separate the components. Clear demarcation between the prompt components helps AI generate more coherent and structured output
-
[5]
T rying several times . Also referred to as resampling, this technique involves running the model multiple times and selecting the best candidate output, thereby increasing the likelihood of obtaining an optimal response. It is particularly beneficial given the nondeterministic nature of AI models in generating outputs
-
[6]
The feature of providing none (Zero-shot), single (One-shot) or multiple examples (Few-shot prompting) to draw from and generate responses
Zero-shot, One-shot, or F ew-shot prompting. The feature of providing none (Zero-shot), single (One-shot) or multiple examples (Few-shot prompting) to draw from and generate responses. In the context of zero-shot prompting, the model relies on its large training data for generating responses
-
[7]
Varying the parameter settings of the AI model such as ‘Temperature’ can also influence the outputs
V ariations of model settings . Varying the parameter settings of the AI model such as ‘Temperature’ can also influence the outputs. The Temperature parameter regulates the degree of randomness of the generated output. A lower temperature parameter can lead to higher deterministic outputs and vice-versa. Among the seven techniques described above, we adop...
-
[8]
““You are an expert physicist and your objective is to give feedback on my answer which is presented as an argument with a claim, evidence and reasoning about a physics problem
Desired performance (‘ Where is the learner going?’): This feature involves explicitly and clearly specifying the goals associated with a given task or performance. Clear communication of 4 TABLE I. The three types of feedback (first column) and corresponding prompt statements (second column) embedding the foundational prompt engineering techniques and pr...
-
[9]
Also referred to as ‘Feed Back’, it includes highlighting both strengths or well-executed parts of the task and gaps in student understanding through probing questions
Current performance (‘ How is the learner going?’): Clarifying the current state of performance, often in relation to the desired performance. Also referred to as ‘Feed Back’, it includes highlighting both strengths or well-executed parts of the task and gaps in student understanding through probing questions. Such a ‘praise and question’ approach can sup...
-
[10]
Being Clear and Precise
Concrete approaches to fill the gaps (‘Where to next?’) : Prescriptions about ways to address gaps between current and desired performances. You are an intern in an astrophysics laboratory, monitoring the motion of a satellite of mass 500 kg around a planet (with mass 6 .42 × 1023 kg) in a nearby galaxy. You receive data the following from the observatory...
2025
-
[11]
Please review and provide feedback for the argument
Providing clear instructions Overall, 90.2% responses evidenced the technique of providing clear instructions in their prompts by explicitly seeking feedback for their arguments. These instructions were framed as requests (“Please review and provide feedback for the argument ”), commands (“ Give me feedback to this argument”) , or questions (“ Can you let...
-
[12]
As the distance increases, the potential energy increases
Being clear and precise To capture the second technique, ‘Being clear and precise’, we analyzed students’ prompts across two components: (i) explicit description of the task’s context by specifying the slope of gravitational potential energy versus distance graph and (ii) explicit mention of arguments. If students evidenced both components, we coded the p...
-
[13]
Out of 1151 responses, only six demonstrated role-based prompting, and only three included explicit use of delimiters such as quotes and asterisks
Role-based prompting and Use of delimiters The third and fourth techniques (‘Role-based prompting’ and the ‘Use of delimiters’) were the least evidenced in students’ prompts. Out of 1151 responses, only six demonstrated role-based prompting, and only three included explicit use of delimiters such as quotes and asterisks. Exemplar prompts illustrating thes...
-
[14]
correct” solutions OR “The right approach
Evaluation (E) The first predominant theme corresponds to features of feedback that critically evaluate the strengths and shortcomings of student solutions along with a clear description of ways for improvement to facilitate better performance on similar tasks in future. Three features are associated with this theme: (E1) Affirmation (acknowledgment of th...
-
[15]
These include the elucidation of the ‘correct’ approach to solving the physics problem or the specification of the ‘correct answer’
Content (C) This theme reflects feedback features that predominantly focus on content knowledge. These include the elucidation of the ‘correct’ approach to solving the physics problem or the specification of the ‘correct answer’. The theme also encompasses the application of physics and mathematics principles, along with relevant equations. In addition, i...
-
[16]
Presentation (P) The third theme encompasses features related to the presentation of feedback that are independent of the task’s content. These include the clarity of the provided information, the external structure or formatting (e.g., use of an enumerated list), the feedback’s length or brevity, and its overall ease of comprehension. Four features captu...
-
[17]
This theme however does not correspond to the ‘quantity’ or ‘wordiness’ of the feedback but rather the quality of detailed discussion of the students’ solution
Depth (D) The last theme corresponds to the extent of detailed discussion of the student’s solution within or beyond the purview of the task. This theme however does not correspond to the ‘quantity’ or ‘wordiness’ of the feedback but rather the quality of detailed discussion of the students’ solution. Two features are associated with this theme: (D1) Deta...
-
[18]
Students then ranked the three feedback statements generated by AI based on their perceived usefulness along with relevant justifications
Feedback Ranking RQ3: What are the broad trends in students’ preferences for AI-generated feedback based on (a) their self-crafted prompts, (b) prompts entailing foundational prompt-engineering techniques, and (c) prompts combining the foundational prompt-engineering techniques with principles of effective feedback? As noted in Section III, students used ...
-
[19]
Fig 6 presents the trends among the 12 feedback features across the three feedback types as well as the entire data
Themes and features across the feedback types In addition to investigating students’ preferred choice of feedback type, we also explored feedback features identified in the previous subsection (RQ2) that students attribute across each feedback type. Fig 6 presents the trends among the 12 feedback features across the three feedback types as well as the ent...
-
[20]
Student-generated prompts rarely incorporate foundational prompting strategies in combination, as outlined in the prompt engineering literature
-
[21]
Despite this sophistication, the models may not always provide conceptually accurate feedback
The emerging large language models frequently demonstrate ‘contextual intuitiveness’, i.e., taking into account the missing information such as the task’s context from prompts. Despite this sophistication, the models may not always provide conceptually accurate feedback
-
[22]
We identified four themes capturing the features students attribute for the perceived usefulness of AI-feedback: Evaluation, Content, Presentation, and Depth. The specific features include:: (i) Affirmation, (ii) Critique, (iii) Guidance, (iv) Correctness, (v) Conceptual, (vi) Schema, (vii) Clarity, (viii) Structure, (ix) Conciseness, (x) Intelligibility,...
-
[23]
Those who preferred feedback generated from self-crafted prompts preferred conceptual-related information in feedback the most. Students who preferred feedback generated from prompt entailing prompt engineering techniques valued detailed discussion of conceptual-related information along with critical evaluation of their solutions. Students who preferred ...
-
[24]
Students least preferred the feedback generated using self-crafted prompts
An overwhelming fraction of students preferred feedback generated from structured prompts, particularly the one entailing the combination of prompt engineering techniques and the principles of effective feedback. Students least preferred the feedback generated using self-crafted prompts. For a given preferred feedback choice, there was no common second pr...
-
[25]
A. Sirnoorkar and N. S. Rebello, From self-crafted to engineered prompts: Student evaluations of ai-generated feedback in introductory physics, arXiv preprint arXiv:2508.09825 (2025)
Pith/arXiv arXiv 2025
-
[26]
Hattie and H
J. Hattie and H. Timperley, The power of feedback, Review of educational research 77, 81 (2007)
2007
-
[27]
Burke and J
D. Burke and J. Pieterick, Giving students effective written feedback (McGraw-Hill Education (UK), 2010)
2010
-
[28]
Azevedo and R
R. Azevedo and R. M. Bernard, A meta-analysis of the effects of feedback in computer-based instruction, Journal of Educational Computing Research 13, 111 (1995)
1995
-
[29]
Giamos, O
D. Giamos, O. Doucet, and P.-M. L´ eger, Continuous performance feedback: Investigating the effects of feedback content and feedback sources on performance, motivation to improve performance and task engagement, Journal of Organizational Behavior 16 Management 44, 194 (2024)
2024
-
[30]
D. J. Nicol and D. Macfarlane-Dick, Formative assessment and self-regulated learning: A model and seven principles of good feedback practice, Studies in higher education 31, 199 (2006)
2006
-
[31]
Sirnoorkar, A
A. Sirnoorkar, A. P. Jambuge, K. D. Rainey, A. Adamson, B. R. Wilcox, and J. T. Laverty, Theoretical approach for providing feedback for instructors through a standardized assessment for undergraduate physics, in Proceedings of the 17th International Conference of the Learning Sciences-ICLS 2023, pp. 1326-1329 (International Society of the Learning Scienc...
2023
-
[32]
J. T. Laverty, A. Sirnoorkar, A. P. Jambuge, K. D. Rainey, J. Weaver, A. Adamson, and B. R. Wilcox, A new paradigm for research-based assessment development, in Physics Education Research Conference Proceedings (2022)
2022
-
[33]
W. Allen, A. Shanker, and N. S. Rebello, Students’ perceptions to a large language model’s generated feedback and scores of argumentation essays, arXiv preprint arXiv:2508.14759 (2025)
Pith/arXiv arXiv 2025
-
[34]
Polverini and B
G. Polverini and B. Gregorcic, How understanding large language models can inform the use of chatgpt in physics education, European Journal of Physics 45, 025701 (2024)
2024
-
[35]
Polverini and B
G. Polverini and B. Gregorcic, Evaluating vision-capable chatbots in interpreting kinematics graphs: a comparative study of free and subscription-based models, in Frontiers in Education , Vol. 9 (Frontiers Media SA, 2024) p. 1452414
2024
-
[36]
B. Chen, Z. Zhang, N. Langren´ e, and S. Zhu, Unleashing the potential of prompt engineering for large language models, Patterns (2025)
2025
-
[37]
R. Hamed, A. Sirnoorkar, and N. S. Rebello, Dual-role dynamics in prompting: Elementary pre-service teachers’ ai prompting strategies for representational choices, arXiv preprint arXiv:2508.14760 (2025)
Pith/arXiv arXiv 2025
-
[38]
Kortemeyer, The boiling-frog problem of physics education, arXiv preprint arXiv:2508.08842 (2025)
G. Kortemeyer, The boiling-frog problem of physics education, arXiv preprint arXiv:2508.08842 (2025)
arXiv 2025
-
[39]
Kortemeyer, Could an artificial-intelligence agent pass an introductory physics course?, Physical Review Physics Education Research 19, 010132 (2023)
G. Kortemeyer, Could an artificial-intelligence agent pass an introductory physics course?, Physical Review Physics Education Research 19, 010132 (2023)
2023
-
[40]
Sirnoorkar, D
A. Sirnoorkar, D. Zollman, J. T. Laverty, A. J. Magana, N. S. Rebello, and L. A. Bryan, Student and ai responses to physics problems examined through the lenses of sensemaking and mechanistic reasoning, Computers and Education: Artificial Intelligence 7, 100318 (2024)
2024
-
[41]
Kortemeyer, J
G. Kortemeyer, J. N¨ ohl, and D. Onishchuk, Grading assistance for a handwritten thermodynamics exam using artificial intelligence: An exploratory study, Physical Review Physics Education Research 20, 020144 (2024)
2024
-
[42]
Kortemeyer, Performance of the pre-trained large language model gpt-4 on automated short answer grading, Discover Artificial Intelligence 4, 47 (2024)
G. Kortemeyer, Performance of the pre-trained large language model gpt-4 on automated short answer grading, Discover Artificial Intelligence 4, 47 (2024)
2024
-
[43]
Kortemeyer and J
G. Kortemeyer and J. N¨ ohl, Assessing confidence in ai-assisted grading of physics exams through psychometrics: An exploratory study, Physical Review Physics Education Research 21, 010136 (2025)
2025
-
[44]
G.-G. Lee, E. Latif, X. Wu, N. Liu, and X. Zhai, Applying large language models and chain-of-thought for automatic scoring, Computers and Education: Artificial Intelligence 6, 100213 (2024)
2024
-
[45]
Chen and T
Z. Chen and T. Wan, Grading explanations of problem-solving process and generating feedback using large language models at human-level accuracy, Physical Review Physics Education Research 21, 010126 (2025)
2025
-
[46]
Y. Wei, R. Zhang, J. Zhang, D. Qi, and W. Cui, Research on intelligent grading of physics problems based on large language models, Education Sciences 15, 116 (2025)
2025
-
[47]
Wan and Z
T. Wan and Z. Chen, Exploring generative ai assisted feedback writing for students’ written responses to a physics conceptual question with prompt engineering and few-shot learning, Physical Review Physics Education Research 20, 010152 (2024)
2024
-
[48]
X. Dai, Z. Wen, J. Jiang, H. Liu, and Y. Zhang, How students use ai feedback matters: Experimental evidence on physics achievement and autonomy, arXiv preprint arXiv:2505.08672 (2025)
Pith/arXiv arXiv 2025
-
[49]
Mills, A
E. Mills, A. Mizouri, and A. Peach, Prompting better feedback: A study of custom gpt for formative assessment in undergraduate physics, Education Sciences 15, 1058 (2025)
2025
-
[50]
Https://www.qualtrics.com/
-
[51]
R. W. Chabay and B. A. Sherwood, Matter and interactions (John Wiley & Sons, 2015)
2015
-
[52]
S. E. Toulmin, The uses of argument (Cambridge university press, 2003)
2003
-
[53]
K. L. McNeill and J. Krajcik, Inquiry and scientific explanations: Helping students use evidence and reasoning, Science as inquiry in the secondary setting 121, 34 (2008)
2008
-
[54]
J. W. Creswell, W. E. Hanson, V. L. Clark Plano, and A. Morales, Qualitative research designs: Selection and implementation, The counseling psychologist35, 236 (2007)
2007
-
[55]
Sawalha, I
G. Sawalha, I. Taj, and A. Shoufan, Analyzing student prompts and their effect on chatgpt’s performance, Cogent Education 11, 2397200 (2024)
2024
-
[56]
Zhang, Z
Z. Zhang, Z. Dong, Y. Shi, T. Price, N. Matsuda, and D. Xu, Students’ perceptions and preferences of generative artificial intelligence feedback for programming, in Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38 (2024) pp. 23250–23258
2024
-
[57]
S. Tassoti, Assessment of students use of generative artificial intelligence: Prompting strategies and prompt engineering in chemistry education, Journal of Chemical Education 101, 2475 (2024)
2024
-
[58]
Valeri, P
F. Valeri, P. Nilsson, and A.-M. Cederqvist, Exploring students’ experience of chatgpt in stem education, Computers and Education: Artificial Intelligence 8, 100360 (2025)
2025
-
[59]
Henderson, M
M. Henderson, M. Bearman, J. Chung, T. Fawns, S. Buckingham Shum, K. E. Matthews, and J. de Mello Heredia, Comparing generative ai and teacher feedback: student perceptions of usefulness and trustworthiness, Assessment & Evaluation in Higher Education , 1 (2025)
2025
-
[60]
S. Prompiengchai, C. Narreddy, and S. Joordens, A practical guide for supporting formative assessment and feedback using generative ai, arXiv preprint arXiv:2505.23405 (2025)
Pith/arXiv arXiv 2025
-
[61]
M. N. Dahlkemper, S. Z. Lahme, and P. Klein, How do physics students evaluate artificial intelligence responses on comprehension questions? a study on the perceived scientific accuracy and linguistic quality of chatgpt, Physical Review Physics Education Research 19, 010142 (2023)
2023
This paper was first reviewed by deepseek-v4-flash on August 4, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.