Pith. sign in

REVIEW 4 major objections 6 minor 61 references

Students who prefer a hybrid chatbot style tend to hold more sophisticated views of physics knowledge, though the effect fades under stricter statistics.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Students who chose a combination chatbot (guided inquiry then answers) scored slightly higher on an epistemological-beliefs survey than answer-preference students, but the difference was not robust to multiple-testing correction.

T0 review reviewed 2026-08-03 challenge →

load-bearing objection Honest exploratory study of a new question, but the headline association rests on an unvalidated one-item preference measure and tiny effects; worth refereeing, not citing as established. the 4 major comments →

arxiv 2607.29385 v1 pith:ZHFFTEY7 submitted 2026-07-31 physics.ed-ph

Students' Epistemological Beliefs and their Chatbot Preferences in AI-mediated Physics Learning

classification physics.ed-ph
keywords epistemological beliefschatbot preferencesphysics educationEBAPSAI-mediated learningguided inquiryundergraduate physics
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether students' beliefs about the nature of physics knowledge and learning line up with how they want an AI chatbot to help them. In an online waves module with a built-in chatbot, students picked one of three preferred chatbot styles: direct answers, guided questions, or a combination that starts with guidance and gives answers only if they get stuck. The paper reports that, at the usual statistical threshold, students who chose the combination style scored higher on the EBAPS measure of epistemological beliefs than students who chose direct answers, especially on beliefs about evolving knowledge and about whether ability is fixed or can grow. But those differences shrank below significance once the analyses were adjusted for multiple comparisons, so the authors describe the finding as suggestive rather than conclusive. A sympathetic reader would take the paper as an early map of a new design question: how to tailor AI tutoring responses to students' epistemic orientations.

Core claim

The paper's central claim is that a student's preference for how a chatbot should respond—give the answer, guide with questions, or switch from guidance to answers when needed—carries a detectable, if modest, association with that student's epistemological beliefs in physics. Using the standardized EBAPS survey and a custom waves activity with an embedded chatbot, the authors found that the group preferring the combination style had higher total EBAPS scores and higher scores on the 'evolving knowledge' and 'source of ability to learn' axes than the group preferring direct answers (nominal p<0.05). No significant differences emerged for the other three axes, and none of the omnibus findings

What carries the argument

The central instrument is the Epistemological Beliefs Assessment for Physical Sciences (EBAPS), a 30-item questionnaire scored into five non-orthogonal axes measuring beliefs about knowledge structure, learning, applicability, evolving knowledge, and the source of learning ability. The central behavioral measure is a single pre-activity multiple-choice item in which students choose one of three chatbot-response styles: direct answers, guided Q&A, or a combination. The statistical machinery is a Kruskal-Wallis omnibus test across the three preference groups followed by Dunn's post-hoc comparisons with Bonferroni correction; the paper also computes Cronbach's alpha for each axis to assess inte

Load-bearing premise

The load-bearing premise is that the EBAPS subscales—especially the very unreliable 'evolving knowledge' axis (alpha=0.13)—actually measure stable epistemological traits, and that a single multiple-choice question captures students' genuine chatbot preferences.

What would settle it

A direct replication with a new sample that (1) measures chatbot preference both by self-report and by observed behavior in an actual chatbot session, and (2) uses a more internally consistent measure of evolving-knowledge beliefs, would settle the claim: if the combination-vs-answer difference on Axis 4 and 5 does not reappear at p<0.05 (or fails to survive correction again), the suggested association would be unsupported. A simpler check: compute the correlation between EBAPS scores and preference group after controlling for time spent on the activity or interest in the topic; if it vanishes

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Chatbot designers can reasonably prototype a 'combination' mode—guided inquiry first, direct answers on request—as a default that fits the stated preferences of the majority of students in this sample.
  • The Axis 5 finding, if replicated, suggests that chatbot feedback which frames ability as developable (strategy hints, reflection prompts) aligns with more productive epistemologies.
  • The absence of any difference between guided-inquiry and answer-preferring groups on total EBAPS scores indicates that a pure inquiry approach is not noticeably associated with more sophisticated beliefs in this population.
  • Future work should check whether sustained use of such chatbots shifts students' epistemological beliefs, or only selects on pre-existing ones.
  • Because the effects failed the Bonferroni-corrected threshold, any design decision should be treated as tentative until replicated.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A testable implication the authors do not draw: the preference measure may be partly a statement of ideal help-seeking style rather than actual behavior, so a chatbot that logs real requests for answers after hints would test whether stated preference matches enacted preference.
  • The particularly low reliability of the 'evolving knowledge' subscale (alpha=0.13) means the Axis 4 difference is the most fragile finding; a revised or supplementary measure could resolve whether the association is real or an artifact of measurement noise.
  • The lopsided preference distribution (52% combination, 40% direct answers, 8% guided only) suggests the association is driven largely by the contrast between the two large groups, making the tiny guided-only group a difficult comparison base.
  • A direct experimental extension—randomly assigning students to chatbot styles and measuring EBAPS afterward—would distinguish selection (students with certain beliefs choose certain styles) from causation (the style itself shapes beliefs).
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This paper reports an observational study in a large-enrollment introductory physics course (N=1048 matched responses) examining whether students' stated preferences for chatbot behavior (direct answers, guided inquiry, or a combination) are associated with their epistemological beliefs as measured by the EBAPS. The authors use Kruskal-Wallis tests at α=0.05 and report nominally significant omnibus differences for the total EBAPS score and for Axes 4 and 5, with Dunn post-hoc tests indicating that students choosing the 'Combination' option scored higher than those choosing 'Direct Answers'. They also report that these differences do not remain significant after a Bonferroni-adjusted omnibus threshold (α=0.0083). The Discussion and Conclusion frame the findings as suggestive rather than conclusive.

Significance. If the association were robust, the study would provide a useful empirical bridge between students' epistemic cognition and their preferences for AI-based instructional scaffolding, an area with little prior PER research. The study has notable strengths: a large sample, use of a standardized instrument (EBAPS), honest reporting of the Bonferroni-adjusted null result, and explicit acknowledgment of low subscale reliabilities and modest effect sizes. These features are valuable for a preliminary investigation. However, the central claim is weakened by three intertwined problems: the main grouping variable is a single unvalidated forced-choice item, the only robust-looking subscale finding (Axis 4) rests on a scale with Cronbach's α=0.13, and effect sizes are near zero. The paper is therefore more a proof-of-concept for a research question than a demonstration of an association.

major comments (4)
  1. [Section III / IV] The grouping variable is a single pre-activity multiple-choice item ('Direct Answers', 'Guided Q&A', 'Combination') with no reported validity evidence. The authors state that they 'focus only on the information pertaining to the pre-completion of the activity as they highlight students' natural preferences', but they also collected a retrospective preference item and presumably logged actual chatbot interactions. Without using those data to establish test-retest stability or behavioral concordance, the preference item may not measure a stable disposition, and every EBAPS comparison loses its intended interpretation. This is load-bearing because the entire study is a comparison of EBAPS scores across these preference groups. Please report the pre-post agreement and, if possible, whether the stated preference predicts actual answer-seeking behavior in the module.
  2. [Section IV, Table III] Effect sizes are reported as 0.004–0.007 with no confidence intervals. For N≈1048, even trivial differences can reach p<0.05. The paper also does not report group sizes or distributional summaries beyond means and SDs. Furthermore, the Dunn's post-hoc p-values (0.029, 0.038, 0.021) are ambiguously described as 'with Bonferroni correction' — it is unclear whether these are already family-wise adjusted. If they are unadjusted, they would not survive even a within-test correction, much less the six omnibus tests. The authors should clarify this and provide CIs, group ns, and a clearer labeling of the effect-size statistic (the values likely correspond to epsilon-squared from H, not η²).
  3. [Section IV, Table I and Fig. 1] Axis 4 ('Evolving knowledge') has Cronbach's α=0.13, making it effectively unreliable, yet it is one of the two axes with a nominally significant post-hoc pairwise difference. The Discussion's emphasis on this axis as evidence of sophisticated beliefs is disproportionate to the reliability of the measure. The authors acknowledge the low alpha in the limitations, but the abstract and Discussion foreground the Axis 4 result. I recommend either removing Axis 4 from the headline claims or, at minimum, clearly labeling it as exploratory and uninterpretable due to low internal consistency. Relatedly, Fig. 1 indicates that Item 19 overlaps Axes 1 and 3 and Item 28 overlaps Axes 1 and 4; the scoring in Eq. (1) would count these items in both axes, potentially inflating inter-axis correlations. This should be addressed explicitly.
  4. [Abstract / Discussion] The abstract states that students preferring the 'Combination' option 'demonstrated sophisticated epistemological beliefs' than those preferring answer-providing chatbots. Given that the Bonferroni-adjusted omnibus test is not significant, all unadjusted axes have tiny effect sizes, and the relevant subscales have low reliability, the word 'demonstrated' overstates the evidence. The Discussion more appropriately uses 'suggestive than conclusive', which should be reflected in the abstract. Additionally, the sentence is grammatically incomplete ('...exhibited sophisticated epistemological beliefs than...'). Please rephrase and explicitly note in the abstract that the differences do not survive multiple-comparison correction.
minor comments (6)
  1. [Section III] There are typographical errors: 'a a large-enrollment' and 'was was not part of their course assessment'. These should be corrected.
  2. [Eq. (1)] The summation notation is garbled: 'Σna i=1si' is not typeset properly. Use standard limits, e.g., \(\sum_{i=1}^{n_a} s_i\). Also clarify whether unassigned items (4 and 21) are included in the total score, and if so, how they are weighted.
  3. [Table III / text] The word 'Insignificant' is nonstandard in statistics; use 'Not significant' or 'ns'. Also specify which effect-size statistic is reported (e.g., ε², η², or Cramér's V) and how it was computed.
  4. [References] References [39] and [43] are bare URLs. Include access dates and the title of the webpage (e.g., 'EBAPS website' and 'EBAPS scoring ideas').
  5. [Figure 1] The figure caption lists items that are unassigned or overlapping but does not explain the plot's axis labels or whether error bars/standard deviations are shown. Please add that information, and consider adding the sample size used for the averages.
  6. [Section IV / Table II] The mean total scores are all around 68–70%, but the axis scores have different item counts and possible ceiling/floor effects. Consider presenting standardized scores or item-level averages for comparability, especially since Axis 3 and Axis 5 have means near 80%.

Circularity Check

0 steps flagged

No significant circularity: the study reports an observed association from external instruments and statistical tests, with no prediction derived from fitted inputs or load-bearing self-citation.

full rationale

The paper's central claim is an empirical association between students' self-reported chatbot preference (a single self-designed multiple-choice item) and their EBAPS scores (an external standardized survey). There is no derivation chain in which an output is defined in terms of an input, no parameter fitted to a subset of data and then 'predicted' on a closely related quantity, and no uniqueness theorem or ansatz imported from the authors' prior work to force a conclusion. The only equation in the paper, Eq. (1), is the EBAPS weighted-score normalization; it simply rescales item sums by the maximum possible score and does not embed the study's outcome. The self-citations in the references (e.g., Sirnoorkar et al. 2020, 2024, 2025) are background literature on epistemology and AI in physics education; they are not load-bearing for the association reported here. The paper's own limitations section explicitly notes the single-item preference measure, the narrow three options, low Cronbach's alpha for Axis 4, and modest effect sizes, and the authors describe the results as 'more suggestive than conclusive.' These are construct-validity and statistical-power concerns, not circularity. No step in the manuscript reduces, by definition or by self-citation, to its own inputs. Therefore the appropriate circularity score is 0.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

The analysis has no fitted free parameters and introduces no new entities. Its load-bearing assumptions are measurement and sampling assumptions: the EBAPS subscales are interpretable despite very low reliability on key axes, the single-item preference measure is valid, and the voluntary sample supports inference. These are domain assumptions, not derivation steps.

axioms (4)
  • domain assumption The EBAPS instrument provides valid and interpretable measures of physics epistemological beliefs in this sample.
    Central outcome measure; authors cite the instrument's design but do not revalidate it here. Cronbach's α for Axis 4 is 0.13 and Axis 3 is 0.33 (Table I), so the assumption is fragile for axis-level claims.
  • domain assumption Students' single pre-activity selection among the three chatbot-preference options reflects their genuine, stable preference.
    Preference was captured once with an ad-hoc item on page 2 of the module (Section III); no validity evidence is provided, and the retrospective preference was collected but not analyzed.
  • standard math The Kruskal-Wallis and Dunn tests' assumptions are met by the response data.
    Authors justify the nonparametric approach by non-normal score distributions (Section IV); standard test assumptions of independent groups and ordinal data are invoked.
  • domain assumption The voluntary extra-credit sample of 1048 overlapping responses is representative enough for inference about introductory students.
    Participation was optional and the analysis uses only overlapping responses from a course of ~1800; no missing-data or demographic analysis is reported, so selection effects cannot be ruled out.

reviewed 2026-08-03 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Students' Epistemological Beliefs and their Chatbot Preferences in AI-mediated Physics Learning." pith.science (2026). https://pith.science/paper/ZHFFTEY7

@misc{pith2026260729385,
  author       = {Pith},
  title        = {Pith review of: Students' Epistemological Beliefs and their Chatbot Preferences in AI-mediated Physics Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZHFFTEY7}},
  note         = {Machine review of arXiv:2607.29385}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Evolving technologies have traditionally influenced pedagogical practices and Generative AI is one such technology that promises to transform higher education. In this study, we investigate the association between introductory students' preferences for chatbot behavior and their epistemological beliefs surrounding physics. Our context involves a custom built online module on waves containing simulations integrated with a chatbot. While students' chatbot preferences were captured through three provided options (guided-inquiry, direct answer, and a combination of inquiry and answer), their epistemological beliefs were captured through the standardized Epistemological Beliefs Assessment for Physical Sciences (EBAPS) survey. Results highlight that students who preferred chatbots that initially engage them in guided-inquiry but provide answers when explicitly sought (`Combination'), demonstrated sophisticated epistemological beliefs than those who preferred answer-providing chatbots. Notably, we did not observe any association between the EBAPS' total scores among students who preferred guided-inquiry and those who preferred answer-oriented chatbots. Furthermore, the observed differences did not remain statistically significant after applying a Bonferroni-adjusted significance level. Implications of these results for the design and instructional use of chatbots in physics education are discussed.

Figures

Figures reproduced from arXiv: 2607.29385 by Amogh Sirnoorkar, Omkar Mamidpalliwar.

Figure 1
Figure 1. Figure 1: FIG. 1. Students’ average percentage scores on EBAPS items, [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

61 extracted references · 7 linked inside Pith

  1. [1]

    (10) Whether scientific knowledge is perceived as a collection of weakly connected pieces of facts, or as a coherent, conceptual, and highly structured unified whole

    Structure of scientific knowledge. (10) Whether scientific knowledge is perceived as a collection of weakly connected pieces of facts, or as a coherent, conceptual, and highly structured unified whole

  2. [2]

    Nature of knowing and learning. (8) The extent to which learning science is viewed as primarily absorbing information, as opposed to actively constructing one’s own understanding by engaging with the material, connecting it to prior experiences and intuitions

  3. [3]

    (4) Whether scientific knowledge is perceived as applicable only in formal settings or real-life situations as well

    Real-life applicability. (4) Whether scientific knowledge is perceived as applicable only in formal settings or real-life situations as well

  4. [4]

    Evolving knowledge. (3) The extent to which students balance between the extremes of absolutism (viewing scientific knowledge as fixed and unchanging) and extreme relativism (failing to distinguish between evidence-based reasoning and mere opinion)

  5. [5]

    (5) Whether students view success in science as determined by fixed innate ability versus through effort and study strategies

    Source of ability to learn. (5) Whether students view success in science as determined by fixed innate ability versus through effort and study strategies. 3 FIG. 1. Students’ average percentage scores on EBAPS items, grouped by axis. Items 4 and 21 are not assigned to any of the five axes. Item 19 overlaps between Axes 1 and 3, whereas Item 28 overlaps be...

  6. [6]

    Direct Answers (I want the chatbot to give me answers directly rather than guiding me to figure them out on my own.)

  7. [7]

    Guided Q&A (I want the chatbot to guide me to figure out the questions on my own through conceptual hints and guiding questions.)

  8. [8]

    Combination

    Combination (I want the chatbot to guide me at first so I can figure out the questions on my own, but provide direct answers if I get stuck for a while.) While 52% highlighted the “Combination” option, 39.6% preferred “Direct Answers”. Only 7.6% highlighted the “Guided Q&A” preference. On each of the following pages a specific topic was focused by present...

  9. [9]

    Structure of knowledge 2.4 1.4 0.45

  10. [10]

    Nature of knowing 2.7 1.3 0.42

  11. [11]

    Real-life applicability 3.2 1.1 0.33

  12. [12]

    Evolving knowledge 2.8 1.4 0.13

  13. [13]

    the assessment items were designed so that students were allowed to disagree with themselves within a subscale

    Source of ability to learn 3.2 1.0 0.64 four questions. On the last page students were asked to provide their student ID, rate the activity on a scale of 5, highlight their perceived learning, and their qualitative feedback about the activity. In addition, we also collected retrospective reflection about their content familiarity and chatbot preferences w...

  14. [14]

    Elby and D

    A. Elby and D. Hammer, On the substance of a sophisticated epistemology, Science education 85, 554 (2001)

  15. [15]

    Hammer, A

    D. Hammer, A. Elby, R. E. Scherr, and E. F. Redish, Resources, framing, and transfer, Transfer of learning from a modern multidisciplinary perspective 89 (2005)

  16. [16]

    Chen, Epistemic uncertainty and the support of productive struggle during scientific modeling for knowledge co-development, Journal of Research in Science Teaching 59, 383 (2022)

    Y.-C. Chen, Epistemic uncertainty and the support of productive struggle during scientific modeling for knowledge co-development, Journal of Research in Science Teaching 59, 383 (2022)

  17. [17]

    Sirnoorkar, A

    A. Sirnoorkar, A. Mazumdar, and A. Kumar, Towards a content-based epistemic measure in physics, Physical Review Physics Education Research 16, 010103 (2020)

  18. [18]

    B. K. Hofer and P. R. Pintrich, The development of epistemological theories: Beliefs about knowledge and knowing and their relation to learning, Review of educational research 67, 88 (1997)

  19. [19]

    J.-C. J. Jehng, S. D. Johnson, and R. C. Anderson, Schooling and students’ epistemological beliefs about learning, Contemporary educational psychology 18, 23 (1993)

  20. [20]

    Hammer, Epistemological beliefs in introductory physics, Cognition and instruction 12, 151 (1994)

    D. Hammer, Epistemological beliefs in introductory physics, Cognition and instruction 12, 151 (1994)

  21. [21]

    Schommer, Effects of beliefs about the nature of knowledge on comprehension., Journal of educational psychology 82, 498 (1990)

    M. Schommer, Effects of beliefs about the nature of knowledge on comprehension., Journal of educational psychology 82, 498 (1990)

  22. [22]

    Stathopoulou and S

    C. Stathopoulou and S. Vosniadou, Exploring the relationship between physics-related epistemological beliefs and physics understanding, Contemporary Educational Psychology 32, 255 (2007)

  23. [23]

    Valanides and C

    N. Valanides and C. Angeli, Effects of instruction on changes in epistemological beliefs, Contemporary Educational Psychology 30, 314 (2005)

  24. [24]

    S. D. Belet and M. Guven, Meta-cognitive strategy usage and epistemological beliefs of primary school teacher trainees., Educational Sciences: Theory and Practice 11, 51 (2011)

  25. [25]

    Bromme, S

    R. Bromme, S. Pieschl, and E. Stahl, Epistemological beliefs are standards for adaptive learning: A functional theory about epistemological beliefs and metacognition, Metacognition and learning 5, 7 (2010)

  26. [26]

    A. Aypay, The adaptation of the teaching-learning conceptions questionnaire and its relationships with epistemological beliefs., Educational Sciences: Theory and Practice 11, 21 (2011)

  27. [27]

    Trautwein and O

    U. Trautwein and O. L¨ udtke, Epistemological beliefs, school achievement, and college major: A large-scale longitudinal study on the impact of certainty beliefs, Contemporary educational psychology 32, 348 (2007)

  28. [28]

    C. J. De Brabander and J. S. Rozendaal, Epistemological beliefs, social status, and school preference: An exploration of relationships, Scandinavian Journal of Educational Research 51, 141 (2007)

  29. [29]

    Kortemeyer, Could an artificial-intelligence agent pass an introductory physics course?, Physical Review Physics Education Research 19, 010132 (2023)

    G. Kortemeyer, Could an artificial-intelligence agent pass an introductory physics course?, Physical Review Physics Education Research 19, 010132 (2023)

  30. [30]

    Sirnoorkar, D

    A. Sirnoorkar, D. Zollman, J. T. Laverty, A. J. Magana, N. S. Rebello, and L. A. Bryan, Student and ai responses to physics problems examined through the lenses of sensemaking and mechanistic reasoning, Computers and Education: Artificial Intelligence 7, 100318 (2024)

  31. [31]

    Bralin, A

    A. Bralin, A. Sirnoorkar, Y. Zhang, and N. S. Rebello, Mapping the literature landscape of artificial intelligence and machine learning in physics education research, in Proceedings of the Physics Education Research Conference-PERC (2024) pp. 52–59

  32. [32]

    C. G. West, Ai and the fci: Can chatgpt project an understanding of introductory physics?, arXiv preprint arXiv:2303.01067 (2023)

  33. [33]

    Polverini, J

    G. Polverini, J. Melin, E. ¨Onerud, and B. Gregorcic, Performance of chatgpt on tasks involving physics visual representations: The case of the brief electricity and magnetism assessment, Physical Review Physics Education Research 21, 010154 (2025)

  34. [34]

    J. Qiu, J. Shi, X. Juan, Z. Zhao, J. Geng, S. Liu, H. Wang, S. Wu, and M. Wang, Physics supernova: Ai agent matches elite gold medalists at ipho 2025, arXiv preprint arXiv:2509.01659 (2025)

  35. [35]

    Tschisgale, H

    P. Tschisgale, H. Maus, F. Kieser, B. Kroehs, S. Petersen, and P. Wulff, Evaluating gpt-and reasoning-based large language models on physics olympiad problems: Surpassing human performance and implications for educational assessment, Physical Review Physics Education Research 21, 020115 (2025)

  36. [36]

    Kortemeyer, M

    G. Kortemeyer, M. Babayeva, G. Polverini, R. Widenhorn, and B. Gregorcic, Multilingual performance of a multimodal artificial intelligence system on multisubject physics concept inventories, Physical Review Physics Education Research 21, 020101 (2025)

  37. [37]

    Polverini and B

    G. Polverini and B. Gregorcic, Multimodal large language models and physics visual tasks: comparative analysis of performance and costs, European Journal of Physics 46, 055708 (2025)

  38. [38]

    Ravˇ selj, D

    D. Ravˇ selj, D. Kerˇ ziˇ c, N. Tomaˇ zeviˇ c, L. Umek, N. Brezovar, N. A. Iahad, A. A. Abdulla, A. Akopyan, M. W. A. Segura, J. AlHumaid, et al., Higher education students’ perceptions of chatgpt: A global study of early reactions, PLoS One 20, e0315011 (2025)

  39. [39]

    Fageeh, The rise of chatbots in higher education: Exploring user profiles, motivations, and integration strategies, Social Sciences & Humanities Open 12, 101996 (2025)

    A. Fageeh, The rise of chatbots in higher education: Exploring user profiles, motivations, and integration strategies, Social Sciences & Humanities Open 12, 101996 (2025)

  40. [40]

    Kosmyna, E

    N. Kosmyna, E. Hauptmann, Y. T. Yuan, J. Situ, X.-H. Liao, A. V. Beresnitzky, I. Braunstein, and P. Maes, Your brain on chatgpt: Accumulation of cognitive debt when using an ai assistant for essay writing task, arXiv preprint arXiv:2506.08872 4 (2025)

  41. [41]

    Sirnoorkar and N

    A. Sirnoorkar and N. S. Rebello, Feedback that clicks: Introductory physics students’ valued features in ai feedback generated from self-crafted and engineered prompts, arXiv preprint arXiv:2509.08516 (2025)

  42. [42]

    Mills, A

    E. Mills, A. Mizouri, and A. Peach, Prompting better feedback: A study of custom gpt for formative assessment in undergraduate physics, Education Sciences 15, 1058 (2025)

  43. [43]

    Allen, A

    W. Allen, A. Shanker, and N. S. Rebello, Students’ perceptions to a large language model’s generated feedback and scores of argumentation essays, arXiv preprint arXiv:2508.14759 (2025)

  44. [44]

    Wan and Z

    T. Wan and Z. Chen, Exploring generative ai assisted feedback writing for students’ written responses to a 7 physics conceptual question with prompt engineering and few-shot learning, Physical Review Physics Education Research 20, 010152 (2024)

  45. [45]

    M. N. Dahlkemper, S. Z. Lahme, and P. Klein, How do physics students evaluate artificial intelligence responses on comprehension questions? a study on the perceived scientific accuracy and linguistic quality of chatgpt, Physical Review Physics Education Research 19, 010142 (2023)

  46. [46]

    Lademann, J

    J. Lademann, J. Henze, and S. Becker-Genschow, Augmenting learning environments using ai custom chatbots: Effects on learning performance, cognitive load, and affective variables, Physical Review Physics Education Research 21, 010147 (2025)

  47. [47]

    Hamed, A

    R. Hamed, A. Sirnoorkar, and N. S. Rebello, Dual-role dynamics in prompting: Elementary pre-service teachers’ ai prompting strategies for representational choices, arXiv preprint arXiv:2508.14760 (2025)

  48. [48]

    Jiang, X.-M

    Y. Jiang, X.-M. Feng, Y. Liu, Y. Wang, L. Xie, and L. Bao, Generative ai for feedback and collaborative knowledge construction in preservice physics teacher education, Physical Review Physics Education Research 22, 010116 (2026)

  49. [49]

    K¨ uchemann, S

    S. K¨ uchemann, S. Steinert, N. Revenga, M. Schweinberger, Y. Dinc, K. E. Avila, and J. Kuhn, Can chatgpt support prospective teachers in physics task development?, Physical Review Physics Education Research 19, 020128 (2023)

  50. [50]

    S. F. A. Hashmi and N. S. Rebello, Analyzing undergraduate problem-solving in physics through interaction with an ai chatbot, arXiv preprint arXiv:2508.14778 (2025)

  51. [51]

    Elby, Helping physics students learn how to learn, American Journal of Physics 69, S54 (2001)

    A. Elby, Helping physics students learn how to learn, American Journal of Physics 69, S54 (2001)

  52. [52]

    Https://physics.umd.edu/ elby/EBAPS/home.htm

  53. [53]

    R. W. Chabay and B. A. Sherwood, Matter and interactions (John Wiley & Sons, 2015)

  54. [54]

    Https://www.qualtrics.com/

  55. [55]

    D. L. Streiner, Starting at the beginning: an introduction to coefficient alpha and internal consistency, Journal of personality assessment 80, 99 (2003)

  56. [56]

    Https://physics.umd.edu/ elby/EBAPS/idea.htm

  57. [57]

    Br¨ andle, T

    M. Br¨ andle, T. Bahr, J. B. Arnold, and B. Zinn, A comparative study of epistemological beliefs and ai chatbot usage among early adopters and later users in higher education, in 2025 IEEE Global Engineering Education Conference (EDUCON) (IEEE, 2025) pp. 1–10

  58. [58]

    Urhahne, L

    D. Urhahne, L. Kehle, L. Dietrich, and K. Kremer, The role of epistemic beliefs in predicting chatgpt adoption and avoidance in higher education, Acta Psychologica 263, 106334 (2026)

  59. [59]

    S.-C. J. Sin, Epistemological beliefs as predictors of generative ai familiarity, perceived issues likelihood, and usage, Proceedings of the Association for Information Science and Technology 62, 1076 (2025)

  60. [60]

    Avcı and F

    ¨O. Avcı and F. Ogan-Bekiro˘ glu, Exploring the relationship between pre-service science teachers’epistemological beliefs and attitudes towards artificial intelligence, in ICERI2025 Proceedings(IATED,

  61. [61]

    Zhou and K

    X. Zhou and K. T. Chiu, The role of epistemic beliefs in predicting deep learning strategies in an ai-assisted english approach, Language Testing in Asia (2025)

This paper was first reviewed by deepseek-v4-flash on August 3, 2026.