Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

From Self-Crafted to Engineered Prompts: Student Evaluations of AI-Generated Feedback in Introductory Physics

T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Across 1,235 introductory physics students, roughly 91% ranked AI feedback from structured prompts above feedback from their own self-crafted prompts, with the prompt combining prompt-engineering and effective-feedback principles most often

desk verdict Large but confounded: the study's fixed A→B→C presentation order means the 91.1% preference for structured-prompt feedback cannot be cleanly attributed to prompt engineering. read the letter →

arxiv 2508.09825 v1 pith:6VTGBYHQ submitted 2025-08-13 physics.ed-ph

classification physics.ed-ph
keywords AI-generatedfeedbackpromptengineeringeffectivephysicseducationstudentpreferencesintroductorygenerativeAIClaim-Evidence-Reasoning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether the way a student phrases a request to an AI assistant changes how useful the resulting feedback feels, and whether students can tell the difference. In a large introductory physics course, 1,235 students wrote a Claim-Evidence-Reasoning argument about a satellite's gravitational potential energy, then requested feedback three times: once with a prompt of their own devising, once with a template built from foundational prompt-engineering techniques, and once with that template extended by principles of effective feedback. Roughly 91% of students ranked one of the two structured-prompt outputs as most useful, with the combined prompt-engineering-plus-effective-feedback version favored by about 61.4%. The authors read this as evidence that students recognize and prefer better-prompted feedback, and they argue that introductory courses should teach prompt engineering and build AI feedback requests on the established feedback literature.

What carries the argument

The central mechanism is a three-condition prompt comparison embedded in one extra-credit activity. Each student solves a satellite gravitational-potential-energy problem (a plot of $U$ versus distance $r$), writes a Claim-Evidence-Reasoning argument, and submits that same argument to an AI of their choice three times: with a self-crafted prompt (Feedback A); with a provided template that assigns the AI an expert-physicist role, specifies the problem context, separates sections with delimiters, and requests relevant feedback (Feedback B); and with that template plus an explicit request that the feedback state the correct answer, appraise the student's answer (strengths and limitations), pinp

What would settle it

Inspect the students' self-crafted prompts: if a substantial fraction already contain engineered features (role assignments, delimiters, explicit requests for strengths, limitations, or action steps), the contrast behind the 91.1% figure collapses. A clean test: fix one AI model and one student argument, regenerate Feedback A, B, and C, anonymize the three outputs by stripping identifying formatting, and have students rank them; if preferences largely vanish, the effect is driven by feedback presentation rather than prompt structure.

Watch

Extended reading notes

Core claim

Across 1,235 responses, students ranked AI feedback produced by three prompt types: their own self-crafted prompt (A); a provided template embedding foundational prompt-engineering techniques such as role-based prompting, clarity, context specification, and delimiters (B); and that template extended with instructions to state the correct answer, evaluate the student's answer with strengths and limitations, identify gaps, stay within 200 words, and list three improvement strategies (C). Feedback C was ranked most useful by about 61.4%, B by 29.7%, and A by only 8.9%, so 91.1% preferred one of the structured prompts. The self-crafted-prompt output was the least preferred for 61.2% of students.

Load-bearing premise

The load-bearing premise is that students' self-crafted prompts were genuinely naive — devoid of prompt-engineering techniques and effective-feedback features — so that Feedback A differs from B and C only in prompt structure; if many students already wrote engineered prompts, the 91.1% preference for structured prompts partly evaporates.

Editorial extensions

If this is right

  • If the preference is real, basic prompt-engineering instruction in introductory physics would plausibly raise the quality of AI feedback students obtain, since most students' own prompts produced their least-preferred feedback.
  • Feedback designers should encode effective-feedback principles (desired performance, current performance, gap-closing steps) into AI prompts: the version that did so was the most frequent top choice.
  • The most popular format (C) also drew the largest share of last-place votes (about 30%), so a single best prompt will not serve every student.
  • Second-choice patterns differ sharply by first choice — 90% of B-first students chose A second, while 95.5% of C-first students chose B — so beyond the broad preference for structured prompts there is no uniform consensus hierarchy.
  • Self-crafted prompts yielding the least-preferred feedback supports the paper's call to integrate AI literacy and prompting practice into introductory courses.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Students may be responding to the feedback's visible format — explicit correct answer, bulleted action steps, 200-word limit — rather than to the prompt structure that produced it; an anonymized ranking of the three outputs with the prompts unidentified would separate prompt effects from presentation effects.
  • The 90% of B-first choosers who ranked A second suggests a distinct learner cluster that values detailed, mathematical confirmation of their work; segmenting students by preference pattern could motivate personalized prompt defaults.
  • Because the AI model was unrecorded and students chose their own platform, the headline percentages are plausibly model-dependent; a replication that fixes the model and logs versions would test the stability of the 91.1% figure.
  • Teaching students to name the feedback elements they want (correct answer, strengths, gaps, next steps) could reproduce C's advantage with far less prompt-engineering apparatus — a lightweight alternative to the paper's curriculum recommendation.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper reports a classroom study in which 1235 introductory physics students constructed a CER argument about the gravitational potential energy of a satellite and then solicited AI feedback using three prompt types in a fixed sequence: (A) their own self-crafted prompt, (B) a provided prompt-engineering template, and (C) the same template augmented with principles of effective feedback. Students ranked the three feedback statements by usefulness. The descriptive results show that 61.4% ranked C first, 29.7% ranked B first, and only 8.9% ranked A first, yielding a combined first-choice rate of 91.1% for the structured prompts (B and C). The authors interpret this as a strong student preference for feedback generated from structured prompts and discuss implications for integrating prompt engineering into introductory physics courses.

Significance. The topic is timely and practically important, and the study uses a large, ecologically valid classroom sample (1235 of 2044 students). The paper is transparent about its tabulation procedures and includes representative student justifications, and it explicitly acknowledges two threats to its assumptions in Section V. If the preference for structured prompts were established, the instructional implication would be meaningful. However, the fixed presentation order and the lack of inferential statistics currently limit the strength of the central claim. The core finding is better characterized as a description of students' rankings under a particular fixed sequence than as evidence that prompt structure alone drives preference.

major comments (3)
  1. [§III (Methods)] The three tasks are administered in a fixed order—students first craft their own prompt (A), then use the prompt-engineering template (B), then the combined template (C)—with no counterbalancing or between-subjects control. This perfectly confounds prompt type with presentation order. The combined 91.1% first-choice rate for B/C could reflect recency, practice, or accumulated context in the same chat session rather than prompt structure. The limitations section (§V) does not mention this threat. The claim that 'students preferred feedback generated using structured prompts' should be qualified as a preference observed under this fixed sequence, or the design must be changed.
  2. [§IV (Results)] All conclusions are based on raw percentages (61.4%, 29.7%, 8.9%, etc.) with no confidence intervals, standard errors, or significance tests. With N=1235, even small differences are likely statistically meaningful, but the current reporting does not allow the reader to assess sampling variability. Report multinomial confidence intervals or a chi-square test against a uniform/random ranking baseline, and, for the second-choice patterns in Figure 3, a formal test of conditional independence. Without this, the ranking differences remain descriptive only.
  3. [§V (Limitations)] The paper explicitly assumes that students' self-crafted prompts were devoid of prompt-engineering and effective-feedback features, making them structurally distinct from B and C. This assumption is load-bearing for the contrast between A and B/C: if a nontrivial fraction of students wrote prompts resembling B or C, the observed A–B/C gap cannot be attributed to prompt structure. The cited literature mitigates but does not eliminate the risk. Please provide at least a characterization of a subsample of the actual student prompts, coded for the features described in §II, to demonstrate that A is indeed distinct.
minor comments (4)
  1. [§I (Introduction)] Typos: 'compared compared' appears twice; please correct.
  2. [References] Reference [23] misspells 'Sherwood' as 'Shwerwood'; reference [26] is a bare URL and should be formatted per journal style; reference [6] and [27] contain 'V ol.' spacing errors.
  3. [§IV–V] The text states there was 'no common preference' in second-choice selections, but Figure 3 and the accompanying results show strong conditional patterns (90% of B-first students chose A second; 95.5% of C-first students chose B second). Please reconcile this apparent contradiction or clarify the intended meaning.
  4. [§III] Students were allowed to use any AI platform and the specific model was not recorded. This is acknowledged in §V, but it would strengthen the paper to state whether all three tasks for a given student were run in the same chat session or in fresh conversations, since session context can affect the perceived quality of later feedback.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper reports self-reported student rankings; no fitted prediction or derivation reduces to its own inputs.

full rationale

This is an empirical survey study, not a derivation. The central results—61.4% ranking Feedback C first, 29.7% ranking B first, 8.9% ranking A first, and 91.1% preferring either structured prompt—are direct tabulations of student rankings (Section IV), not predictions generated from a model fitted to those same rankings. There is no fitted parameter that is later renamed as a prediction, and no equation that defines one quantity in terms of another. The acknowledged assumption that students' self-crafted prompts lacked prompt-engineering and effective-feedback features (Section V) is a threat to construct validity, but it is not circular: the study does not derive the ranking from that assumption, and the assumption is stated as a limitation rather than used as an input to generate the outcome. The fixed presentation order (A→B→C) is a methodological confound, but again it does not make the result circular; it weakens causal attribution, which is a separate validity concern. The paper's self-citations (e.g., refs. [4], [6], [7]) are background literature citations and are not load-bearing for the preference results. Therefore, no circular step meeting the specified criteria can be exhibited.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper's central comparison depends on two unverified premises: that self-crafted prompts lacked the features being compared, and that AI model choice did not differ systematically across tasks. Both are acknowledged in the limitations section. No new entities or fitted parameters are introduced.

assumptions (3)
  • ad hoc to paper Students' self-crafted prompts (Feedback A) were devoid of prompt engineering and effective feedback features, making them structurally distinct from B and C.
    Stated as an underlying assumption in Section V; if false, the experimental contrast loses meaning.
  • domain assumption The AI platform/model used by each student was the same across the three feedback tasks.
    The authors note in Section V that model details were not recorded; the interpretation of prompt effects assumes model consistency within student.
  • domain assumption Student rankings of feedback usefulness are a meaningful proxy for feedback quality or learning benefit.
    The study measures perceived preferences; implications about learning rest on this unmeasured link.

how reviews work

0 comments
Cite this review

Pith. "Pith review of From Self-Crafted to Engineered Prompts: Student Evaluations of AI-Generated Feedback in Introductory Physics." pith.science (2026). https://pith.science/paper/6VTGBYHQ

@misc{pith2026250809825,
  author       = {Pith},
  title        = {Pith review of: From Self-Crafted to Engineered Prompts: Student Evaluations of AI-Generated Feedback in Introductory Physics},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6VTGBYHQ}},
  note         = {Machine review of arXiv:2508.09825}
}
read the original abstract

The abilities of Generative-Artificial Intelligence (AI) to produce real-time, sophisticated responses across diverse contexts has promised a huge potential in physics education, particularly in providing customized feedback. In this study, we investigate around 1200 introductory students' preferences about AI-feedback generated from three distinct prompt types: (a) self-crafted, (b) entailing foundational prompt-engineering techniques, and (c) entailing foundational prompt-engineering techniques along with principles of effective-feedback. The results highlight an overwhelming fraction of students preferring feedback generated using structured prompts, with those entailing combined features of prompt engineering and effective feedback to be favored most. However, the popular choice also elicited stronger preferences with students either liking or disliking the feedback. Students also ranked the feedback generated using their self-crafted prompts as the least preferred choice. Students' second preferences given their first choice and implications of the results such as the need to incorporate prompt engineering in introductory courses are discussed.

Figures

Figures reproduced from arXiv: 2508.09825 by the authors.

Figure 1
Figure 1. FIG. 1. Statement of the problem for which students constructed [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2. Plot highlighting students’ ranking of the three types of feed [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Feedback That Clicks: Introductory Physics Students' Valued Features in AI Feedback Generated From Self-Crafted and Engineered Prompts

    physics.ed-ph 2025-09 conditional novelty 5.0 of 10

    Introductory physics students rarely use prompt engineering on their own, but they rate AI feedback as most useful when the prompt explicitly requests evaluation, the correct answer, and improvement suggestions.

Reference graph

Works this paper leans on

36 extracted references · 32 canonical work pages · cited by 1 Pith paper

  1. [1]

    The feature of designing effective instructions to AI that yield outputs devoid of ambiguity and misinterpretations

    Providing instructions. The feature of designing effective instructions to AI that yield outputs devoid of ambiguity and misinterpretations

  2. [2]

    Explain the meaning of resistance in the context of electric circuits

    Being clear and precise. The effective instructions to AI should be further accompanied by clear and precise com- mands (for instance, “ Explain the meaning of resistance in the context of electric circuits.”) as general instructions (e.g., “ Explain the meaning of R in science” ) can yield contextually broad and ambiguous responses

  3. [3]

    Instructing AI to simulate specific roles (such as that of a domain expert) can yield task- specific and accurate outputs

    Role-based prompting. Instructing AI to simulate specific roles (such as that of a domain expert) can yield task- specific and accurate outputs

  4. [4]

    Use of symbols such as triple quotes (" " ") to separate different parts of a prompt, particularly when the prompt is complex and has multiple components

    Use of delimiters. Use of symbols such as triple quotes (" " ") to separate different parts of a prompt, particularly when the prompt is complex and has multiple components. This can ensure AI to better interpret and differentiate various input elements

  5. [5]

    Trying several times

    Zero-shot, One-shot, and Few-shot prompting. The tech- nique of providing one example (‘One-shot’), multiple ex- amples (‘Few-shot’), or no examples (‘Zero-shot’) when eliciting responses from AI platforms. When no examples are accompanied with carefully crafted prompts, AI gen- erates responses based on its pre-trained data. While there are additional fo...

  6. [6]

    Where should the learner be?

    Desired performance (“Where should the learner be? ”) highlighting expected or ideal response to the task. For instance, in the context of physics problem solving, this feature would correspond to the correct answer

  7. [7]

    Where is the learner now?

    Current performance (“Where is the learner now?”) high- lighting details of the learner’s current state or perfor- mance. This is often accompanied with affective features such as encouragements and positive reinforcements about the current performance that motivate learners to engage constructively with the feedback

  8. [8]

    How to get there

    Concrete approaches to fill the gaps (“How to get there”) between the expected and current performances. This in- cludes cognitive features such as task descriptions, specific areas for improvement, actionable suggestions, and plans for implementing the changes (wherever possible). In addition to these features, timeliness and personalization contribute t...

Show all 36 references
  1. [9]

    Sirnoorkar, A

    A. Sirnoorkar, A. P. Jambuge, K. D. Rainey, A. Adamson, B. R. Wilcox, and J. T. Laverty, Theoretical approach for pro- viding feedback for instructors through a standardized assess- ment for undergraduate physics, in Proceedings of the 17th In- ternational Conference of the Le...

  2. [10]

    Czajka, G

    D. Czajka, G. Reynders, C. Stanford, R. Cole, J. Lantz, and S. Ruder, A novel rubric format for providing feedback on pro- cess skills to stem undergraduate students, Journal of college science teaching 50, 48 (2021)

  3. [11]

    Kortemeyer, Could an artificial-intelligence agent pass an introductory physics course?, Physical Review Physics Educa- tion Research 19, 010132 (2023)

    G. Kortemeyer, Could an artificial-intelligence agent pass an introductory physics course?, Physical Review Physics Educa- tion Research 19, 010132 (2023)

  4. [12]

    Bralin, A

    A. Bralin, A. Sirnoorkar, Y . Zhang, and N. S. Rebello, Mapping the literature landscape of artificial intelligence and machine learning in physics education research, in Proceedings of the Physics Education Research Conference-PERC 2024, pp. 52- 59 (2024)

  5. [13]

    Polverini and B

    G. Polverini and B. Gregorcic, Performance of chatgpt on the test of understanding graphs in kinematics, Physical review physics education research 20, 010109 (2024)

  6. [14]

    Zollman, A

    D. Zollman, A. Sirnoorkar, and J. Laverty, Comparing ai and student responses on variations of questions through the lens of sensemaking and mechanistic reasoning, in Journal of Physics: Conference Series, V ol. 2693 (IOP Publishing, 2024) p. 012019

  7. [15]

    Sirnoorkar, D

    A. Sirnoorkar, D. Zollman, J. T. Laverty, A. J. Magana, N. S. Rebello, and L. A. Bryan, Student and ai responses to physics problems examined through the lenses of sensemaking and mechanistic reasoning, Computers and Education: Artificial Intelligence 7, 100318 (2024)

  8. [16]

    Kortemeyer, J

    G. Kortemeyer, J. Nöhl, and D. Onishchuk, Grading assistance for a handwritten thermodynamics exam using artificial intelli- gence: An exploratory study, Physical Review Physics Educa- tion Research 20, 020144 (2024)

  9. [17]

    Kortemeyer, Performance of the pre-trained large language model gpt-4 on automated short answer grading, Discover Ar- tificial Intelligence 4, 47 (2024)

    G. Kortemeyer, Performance of the pre-trained large language model gpt-4 on automated short answer grading, Discover Ar- tificial Intelligence 4, 47 (2024)

  10. [18]

    Kortemeyer and J

    G. Kortemeyer and J. Nöhl, Assessing confidence in ai-assisted grading of physics exams through psychometrics: An ex- ploratory study, Physical Review Physics Education Research 21, 010136 (2025)

  11. [19]

    G.-G. Lee, E. Latif, X. Wu, N. Liu, and X. Zhai, Applying large language models and chain-of-thought for automatic scoring, Computers and Education: Artificial Intelligence 6, 100213 (2024)

  12. [20]

    Chen and T

    Z. Chen and T. Wan, Grading explanations of problem-solving process and generating feedback using large language models at human-level accuracy, Physical Review Physics Education Research 21, 010126 (2025)

  13. [21]

    Wan and Z

    T. Wan and Z. Chen, Exploring generative ai assisted feedback writing for students’ written responses to a physics conceptual question with prompt engineering and few-shot learning, Phys- ical Review Physics Education Research 20, 010152 (2024)

  14. [22]

    Meyer, T

    A. Meyer, T. Bleckmann, and G. Friege, Automatic feedback on physics tasks using open-source generative artificial intelli- gence, International Journal of Science Education , 1 (2025)

  15. [23]

    B. Chen, Z. Zhang, N. Langrené, and S. Zhu, Unleashing the potential of prompt engineering for large language models, Pat- terns (2025)

  16. [24]

    Hattie and H

    J. Hattie and H. Timperley, The power of feedback, Review of educational research 77, 81 (2007)

  17. [25]

    D. R. Sadler, Formative assessment and the design of instruc- tional systems, Instructional science 18, 119 (1989)

  18. [26]

    Black and D

    P. Black and D. Wiliam, Developing the theory of formative assessment, Educational Assessment, Evaluation and Account- ability (formerly: Journal of personnel evaluation in education) 21, 5 (2009)

  19. [27]

    Noroozi, S

    O. Noroozi, S. K. Banihashem, N. Taghizadeh Kerman, M. Par- vaneh Akhteh Khaneh, M. Babaee, H. Ashrafi, and H. J. Bie- mans, Gender differences in students’ argumentative essay writing, peer review performance and uptake in online learn- ing environments, Interactive Learning ...

  20. [28]

    M. M. Patchan, C. D. Schunn, and R. J. Correnti, The nature of feedback: How peer feedback features affect students’ imple- mentation rate and quality of revisions., Journal of Educational Psychology 108, 1098 (2016)

  21. [29]

    Carless, D

    D. Carless, D. Salter, M. Yang, and J. Lam, Developing sustain- able feedback practices, Studies in higher education 36, 395 (2011)

  22. [30]

    Pardo, J

    A. Pardo, J. Jovanovic, S. Dawson, D. Gaševi ´c, and N. Mirri- ahi, Using learning analytics to scale the provision of person- alised feedback, British journal of educational technology 50, 128 (2019)

  23. [31]

    R. W. Chabay and B. A. Sherwood, Matter and interactions (John Wiley & Sons, 2015)

  24. [32]

    S. E. Toulmin, The uses of argument (Cambridge university press, 2003)

  25. [33]

    K. L. McNeill and J. Krajcik, Inquiry and scientific explana- tions: Helping students use evidence and reasoning, Science as inquiry in the secondary setting 121, 34 (2008)

  26. [34]

    Https://www.qualtrics.com/

  27. [35]

    Zhang, Z

    Z. Zhang, Z. Dong, Y . Shi, T. Price, N. Matsuda, and D. Xu, Students’ perceptions and preferences of generative artificial intelligence feedback for programming, in Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 38 (2024) pp. 23250–23258

  28. [36]

    Sawalha, I

    G. Sawalha, I. Taj, and A. Shoufan, Analyzing student prompts and their effect on chatgpt’s performance, Cogent Education 11, 2397200 (2024)

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.