Pith. sign in

REVIEW 3 major objections 5 minor 159 references

Multilingual Performance of a Multimodal Artificial Intelligence System on Multisubject Physics Concept Inventories

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read On screenshots of physics concept inventories, GPT-4o outscored average post-instruction students in every subject category except laboratory skills.

desk verdict A useful descriptive benchmark with a real image-interpretation finding; the student-outperformance claim is overstated because the human baselines are best-effort and mostly English. read the letter →

arxiv 2501.06143 v3 pith:H54CW57J submitted 2025-01-10 physics.ed-ph cs.AI

classification physics.ed-phcs.AI
keywords physicseducationresearchconceptinventoriesGPT-4omultimodallanguagemodelsmultilingualperformancevisualinterpretationundergraduateassessmenteducationalequity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to map where a current multimodal AI system stands on the kind of conceptual physics questions used to evaluate undergraduate instruction. It screenshots 54 validated concept inventories in up to 35 languages and finds that the AI's average score beats the average post-instruction undergraduate in every subject category except laboratory skills. The two sharpest limits are visual: performance falls from 81% on text-only items to 49% when interpreting a diagram or graph is required, and scores approach random guessing in several non-European languages. The authors argue instructors should therefore treat AI's inventory scores as a partial capability profile, not as evidence of conceptual mastery.

What carries the argument

The carrying mechanism is the multimodal screenshot protocol: 3,662 item images, each submitted three times, for 14,022 solutions, to the model with a structured JSON prompt, then scored against transcribed answer keys. Each item is coded by whether it is text-only, contains an unneeded image, or requires image interpretation, and each inventory is compared to best-effort published post-instruction student scores. This design is what turns raw model outputs into the per-language, per-subject accuracy tables and the visual-interpretation contrast.

What would settle it

Collect a matched, representative sample of post-instruction undergraduates for the same inventories, in the same languages and screenshot formats, with the same scoring rules; if the AI's per-inventory average falls below the student average in more than one subject category, the paper's central outperformance claim fails.

Watch

Extended reading notes

Core claim

The central claim is that GPT-4o, when given screenshots of items from 54 validated physics concept inventories in up to 35 languages, outperforms the average post-instruction undergraduate on most instruments and in every subject category except laboratory skills. In English, its average is 71.1%; thermodynamics (85.2%) and astronomy (80.4%) are strongest, while laboratory skills are weakest at 35.0%. On the multimodal axis, the model answers 81% of text-only items correctly, 79% of items with decorative images, and only 49% of items where reading the image is necessary. The authors also find language dependence, with English and most European languages performing best and near-random scores in Punjabi (20%) and Tamil (22%) on the most widely translated inventory, while items that are hard in English tend to be hard in other languages. The paper frames these as exploratory findings about capability boundaries, not evidence of conceptual understanding.

Load-bearing premise

The load-bearing premise is that the published post-instruction student scores gathered from the literature are representative, comparable benchmarks for undergraduates across languages and institutions; the paper itself notes they are best-effort, heterogeneous, and mostly from English-speaking samples, so the outperformance comparison stands only as firmly as those numbers do.

Editorial extensions

If this is right

  • If a multimodal model scores 81% on text-only concept items, many current conceptual homework and quiz questions can be completed by AI, weakening the validity of unproctored versions of such assessments.
  • If required-image items drop to 49% accuracy, graph- and diagram-based questions are a more resilient format for assessing students without AI assistance.
  • If English and European languages score far above South Asian languages, AI tutoring tools will be least reliable for the students who may most need them, creating an equity risk.
  • If the AI's incorrect answers are consistent across languages rather than random, its error profile can be studied separately from student misconceptions instead of being treated as a noisy student proxy.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension the authors do not run: feed the same items as text with image content transcribed into captions; if accuracy on the 49% 'required image' items rises toward text-only levels, the bottleneck is visual encoding rather than the physics content itself.
  • If the English prompt biases the comparison, re-running with native-language prompts on a subset of inventories would isolate language-of-prompt effects; the paper notes structured outputs were unreliable for non-English prompts, so this would need a different output format.
  • The 66% repeat-incorrect pattern suggests the model has systematic attractors in its wrong answers, and linking those attractors to specific distractors would give instructors a map of where AI and student misconceptions coincide or diverge.
  • Because the student post-instruction benchmarks are mostly English-speaking, the outperformance claim is strongest for English-language assessments; how the gap looks in non-English classrooms remains open.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper reports an empirical measurement study in which GPT-4o (Azure version 2024-08-06) was asked to solve items from 54 physics concept inventories sourced from PhysPort, using screenshots of the items rather than text-only inputs, across 35 languages. The authors report inventory-level and category-level performance, language differences, an item-level cross-lingual difficulty analysis, a comparison with published post-instruction undergraduate student scores, and an analysis of text-only versus image-required items. The headline findings are that GPT-4o performs worst on laboratory-skills items, performs better in English and European languages than in several non-European languages, shows largely language-independent item difficulty, outperforms average post-instruction students in most subject categories except laboratory skills, and performs markedly worse on items requiring visual interpretation (49% correct) than on text-only items (81% correct).

Significance. The study's empirical base is substantial and well documented: 3,662 image files, 4,674 items, 14,022 scored responses, and clear descriptions of the prompting protocol, JSON output schema, and data-processing decisions. The central descriptive result that required-image items are much harder for GPT-4o than text-only items (49% vs 81%) replicates and extends earlier work on vision-capable chatbots, and the cross-language difficulty correlation is an interesting and falsifiable observation. The promised data release on PhysPort is a further strength. However, the comparison with student benchmarks, which supports the strongest claim in the abstract and conclusion, is built on best-effort, mostly English, non-representative literature values and is not language-matched; this issue must be addressed before the categorical outperformance claim can be accepted as stated.

major comments (3)
  1. [Section IV.C, Figure 5, and abstract] The claim that GPT-4o 'outperforms average post-instruction undergraduate students in all subject categories except laboratory skills' is not supported by a language-matched comparison. The %Post values in Tables II–V are, as the authors state in Section II, 'collected best-effort and not necessarily representative,' and Section VII concedes that most human data come from English-speaking students taking English versions. Figure 5, however, plots GPT-4o scores pooled over all languages against these benchmarks. For example, FCI %Post values of 38, 56, and 66 are compared with AI scores in 32 languages, several of which (20% in Punjabi, 22% in Tamil) are far below any of those student values. Because several categories rest on one or two inventories (RELA has only RCI; LAB has only CDPA with a student baseline, while MUQ has none), plausible variation in the benchmark values or restricting the comparison to English-language AI scores could change category-level conclusions. The authors should either redo the comparison with English-only or otherwise language-matched AI scores, or substantially soften the categorical wording in the abstract and conclusion.
  2. [Section IV.D and Figure 6] The coding of items into 'text-only,' 'unneeded image,' and 'required image' is done manually, but no coding rubric, second coder, or inter-rater reliability check is reported. This coding is load-bearing for RQ4 and for the abstract's claim that the AI 'performs worse on items requiring visual interpretation of images,' since the 49% versus 81% gap is computed entirely from these labels. The qualitative direction of the result is likely robust given the size of the gap, but the per-category comparisons in Figure 6, several of which involve small numbers of inventories, need at least a transparent coding protocol or a reliability check to support the reported magnitudes.
  3. [Sections IV.B–IV.C and Table VI] No uncertainty quantification is provided for the reported percentages, even though each item contributes three non-independent stochastic responses from a probabilistic system. Inventory-level and category-level differences, such as the LAB average of 35% versus the THERM average of 85%, or the FCI Portuguese score of 74% versus the Punjabi score of 20%, are discussed as exact values without confidence intervals or tests. A simple bootstrap by items or by inventories would establish which category and language differences are robust. The paper is explicitly exploratory, so this is not a fatal flaw, but the categorical claims in the abstract would be better supported by such an analysis.
minor comments (5)
  1. [Section V, paragraph 2] The text contains a typo: 'particularlyimportantnt' should read 'particularly important.'
  2. [Tables II–V, column 'Vers.'] The version identifiers such as '2.0,' 'F06,' '5.5.7,' and 'vf' are not explained anywhere; readers cannot tell what distinguishes these versions or which exact version was used for each inventory.
  3. [Section IV.B] The statement that 'all inventories — except TUG-K2.6 — were presented to the AI in English' is misleading: the screenshots contained text in the nominal language, and only the system prompt and JSON schema were in English. The wording should be corrected to 'prompted in English.'
  4. [Appendix C, Table X] The random-incorrect-answer baseline is computed assuming five-option items with one correct answer, but the dataset includes inventories with four or other numbers of options. The theoretical probabilities are therefore only approximately applicable; the analysis should either restrict itself to five-option items or recompute the baselines by item type.
  5. [Figure 6] The figure reports percentages without the number of items or submissions in each image-category and subject-category cell; these sample sizes matter, especially for categories with few inventories such as LAB and RELA.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: GPT-4o scores are direct empirical measurements, student benchmarks are external literature values, and no fitted parameter is renamed as a prediction.

full rationale

This is an empirical measurement study with no fitted equations, tuned parameters, or derived quantities that reduce to their inputs. The reported performances are direct aggregates of model outputs obtained by submitting inventory screenshots to GPT-4o and scoring the returned answers against PhysPort solution keys. The cross-language difficulty trend (Section IV.B and Table IX) is an empirical correlation between independent response sets in different languages, not a construction. The student comparison (Section IV.C) uses %Post values gathered from the published literature, as explicitly stated in Section II: 'these are collected best-effort and not necessarily representative.' The paper's own limitations section concedes that 'most of the human data came from English-speaking students taking the English versions of the inventories,' which is a threat to the external validity of the outperformance comparison, but it is not circularity: the student benchmarks are not derived from the GPT-4o outputs. Prior work by the authors (e.g., [31], [36], [37]) is cited as corroborating evidence for visual-interpretation difficulties, not as an input that defines the present measurements. Therefore, the central claims are self-contained empirical findings, and any weaknesses concern data representativeness rather than circular reasoning.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No free parameters are fitted; all numerical outcomes are direct aggregates of model responses and external student scores. The study makes several domain assumptions about the validity of the source inventories, their translations, the faithfulness of screenshots, the neutrality of the English prompt, and the representativeness of published student scores. These assumptions are disclosed in the paper, but they are load-bearing for the corresponding claims. No new entities are postulated.

assumptions (5)
  • domain assumption Inventories rated at least 'bronze star' on PhysPort are valid research-based conceptual assessments.
    Used as the inclusion criterion in Section II; the paper relies on PhysPort's review process for validity without independent re-validation.
  • domain assumption Available translations preserve the conceptual content of the original inventories.
    Section II notes the quality of most non-English translations could not be evaluated, so lower scores could reflect translation errors rather than model language ability.
  • domain assumption Screenshots, including manual edits to close page breaks, faithfully represent the items as students see them.
    Section III.A describes screenshot preparation; the authors treat the images as equivalent to paper versions.
  • domain assumption Using an English prompt with non-English item images does not systematically bias the measured language differences.
    Section III.B and Section VII acknowledge the English prompt could have influenced performance and language switching, so the language comparisons rest on this assumption.
  • domain assumption Published %Post scores are suitable post-instruction undergraduate benchmarks for comparison.
    Section II collects these 'best-effort' scores from multiple sources; Section VII notes most human data came from English-speaking students, making cross-language AI-student comparisons approximate.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multilingual Performance of a Multimodal Artificial Intelligence System on Multisubject Physics Concept Inventories." pith.science (2026). https://pith.science/paper/H54CW57J

@misc{pith2026250106143,
  author       = {Pith},
  title        = {Pith review of: Multilingual Performance of a Multimodal Artificial Intelligence System on Multisubject Physics Concept Inventories},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/H54CW57J}},
  note         = {Machine review of arXiv:2501.06143}
}
read the original abstract

We investigate the multilingual and multimodal performance of a large language model-based artificial intelligence (AI) system, GPT-4o, using a diverse set of physics concept inventories spanning multiple languages and subject categories. The inventories, sourced from the PhysPort website, cover classical physics topics such as mechanics, electromagnetism, optics, and thermodynamics, as well as relativity, quantum mechanics, astronomy, mathematics, and laboratory skills. Unlike previous text-only studies, we uploaded the inventories as images to reflect what a student would see on paper, thereby assessing the system's multimodal functionality. Our results indicate variation in performance across subjects, with laboratory skills standing out as the weakest. We also observe differences across languages, with English and European languages showing the strongest performance. Notably, the relative difficulty of an inventory item is largely independent of the language of the survey. When comparing AI results to existing literature on student performance, we find that the AI system outperforms average post-instruction undergraduate students in all subject categories except laboratory skills. Furthermore, the AI performs worse on items requiring visual interpretation of images than on those that are purely text-based. While our exploratory findings show GPT-4o's potential usefulness in physics education, they highlight the critical need for instructors to foster students' ability to critically evaluate AI outputs, adapt curricula thoughtfully in response to AI advancements, and address equity concerns associated with AI integration.

Figures

Figures reproduced from arXiv: 2501.06143 by the authors.

Figure 1
Figure 1. FIG. 1. Examples of uploaded problem images: FCI, items 8-11, in Persian (left panel) and HTCE, items 16-19, in Chinese [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2. Distribution of scores achieved by GPT-4o on physics [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. FIG. 3. Sina plot [156] of GPT-4o’s performance on English [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (8 more)
Figure 5
Figure 5. Figure 5: FIG. 5. Sina plots [156] of GPT-4o (all languages) and stu [PITH_FULL_IMAGE:figures/full_fig_p022_5.png]
Figure 4
Figure 4. Figure 4: FIG. 4. Sina plots [156] of GPT-4o’s performance on invento [PITH_FULL_IMAGE:figures/full_fig_p022_4.png]
Figure 6
Figure 6. Figure 6: FIG. 6. Performance on inventory items categorized by their [PITH_FULL_IMAGE:figures/full_fig_p023_6.png]
Figure 9
Figure 9. Figure 9: Role (job description) and prompt (task descrip [PITH_FULL_IMAGE:figures/full_fig_p024_9.png]
Figure 7
Figure 7. Figure 7: FIG. 7. The API call used in this study. “EthelOmni” is a [PITH_FULL_IMAGE:figures/full_fig_p024_7.png]
Figure 8
Figure 8. Figure 8: FIG. 8. The role used for this study [PITH_FULL_IMAGE:figures/full_fig_p024_8.png]
Figure 10
Figure 10. Figure 10: FIG. 10. The data structure used for this study. [PITH_FULL_IMAGE:figures/full_fig_p025_10.png]
Figure 11
Figure 11. Figure 11: FIG. 11. Typical output; each problem is independently solved three times (three “problems”-blocks inside of “solutions”). [PITH_FULL_IMAGE:figures/full_fig_p026_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

159 extracted references · 72 canonical work pages

  1. [1]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, Attention is all you need, Advances in neural informa- tion processing systems 30 (2017)

  2. [2]

    T. H. Kung, M. Cheatham, A. Medinilla, ChatGPT, C. Sillos, L. De Leon, C. Elepano, M. Madriaga, R. Ag- gabao, G. Diaz-Candido, et al., Performance of chatgpt on usmle: Potential for ai-assisted medical education using large language models, medRxiv , 2022 (2022). 17

  3. [3]

    html (accessed January 2023)

    Samantha Murphy Kelly, ChatGPT passes exams from law and business schools, https://edition.cnn.com/ 2023/01/26/tech/chatgpt-passes-exams/index. html (accessed January 2023)

  4. [4]

    OpenAI, ChatGPT, https://chat.openai.com/ (ac- cessed April 2024)

  5. [5]

    A. M. Turing, Computing machinery and intelligence, Mind , 433 (1950)

  6. [6]

    C. R. Jones and B. K. Bergen, People cannot distinguish gpt-4 from a human in a turing test, arXiv preprint arXiv:2405.08007 (2024)

  7. [7]

    Achiam, S

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Alt- man, S. Anadkat, et al. , GPT-4 technical report, arXiv preprint arXiv:2303.08774 (2023)

  8. [8]

    OpenAI, ChatGPT, https://openai.com/research/ gpt-4 (accessed April 2024)

Show all 159 references
  1. [9]

    Kortemeyer, Could an artificial-intelligence agent pass an introductory physics course?, Phys

    G. Kortemeyer, Could an artificial-intelligence agent pass an introductory physics course?, Phys. Rev. Phys. Educ. Res. 19, 010132 (2023)

  2. [10]

    Polverini and B

    G. Polverini and B. Gregorcic, How understanding large language models can inform the use of chatgpt in physics education, European Journal of Physics 45, 025701 (2024)

  3. [11]

    Kortemeyer and W

    G. Kortemeyer and W. Bauer, Cheat sites and artificial intelligence usage in online introductory physics courses: What is the extent and what effect does it have on as- sessments?, Phys. Rev. Phys. Educ. Res. 20, 010145 (2024)

  4. [12]

    Yeadon and T

    W. Yeadon and T. Hardy, The impact of AI in physics education: a comprehensive review from GCSE to uni- versity levels, Physics Education 59, 025010 (2024)

  5. [13]

    K. A. Pimbblet and L. J. Morrell, Can ChatGPT pass a physics degree? making a case for reformation of as- sessment of undergraduate degrees, European Journal of Physics 46, 015702 (2024)

  6. [14]

    Sperling and J

    A. Sperling and J. Lincoln, Artificial intelligence and high school physics, The Physics Teacher62, 314 (2024)

  7. [15]

    K¨ uchemann, M

    S. K¨ uchemann, M. Rau, A. Schmidt, and J. Kuhn, Chatgpt’s quality: Reliability and validity of concept inventory items, Frontiers in Psychology 15, 1426209 (2024)

  8. [16]

    Bitzenbauer, Chatgpt in physics education: A pi- lot study on easy-to-implement activities, Contempo- rary Educational Technology 15, ep430 (2023)

    P. Bitzenbauer, Chatgpt in physics education: A pi- lot study on easy-to-implement activities, Contempo- rary Educational Technology 15, ep430 (2023)

  9. [17]

    K¨ uchemann, S

    S. K¨ uchemann, S. Steinert, N. Revenga, M. Schwein- berger, Y. Dinc, K. E. Avila, and J. Kuhn, Can Chat- GPT support prospective teachers in physics task de- velopment?, Phys. Rev. Phys. Educ. Res. 19, 020128 (2023)

  10. [18]

    Kortemeyer, Using artificial-intelligence tools to make LaTeX content accessible to blind readers, TUG- boat 44, 390 (2023)

    G. Kortemeyer, Using artificial-intelligence tools to make LaTeX content accessible to blind readers, TUG- boat 44, 390 (2023)

  11. [19]

    Wan and Z

    T. Wan and Z. Chen, Exploring generative AI assisted feedback writing for students’ written responses to a physics conceptual question with prompt engineering and few-shot learning, Phys. Rev. Phys. Educ. Res. 20, 010152 (2024)

  12. [20]

    Chen and T

    Z. Chen and T. Wan, Grading explanations of problem- solving process and generating feedback using large lan- guage models at human-level accuracy, Physical Review Physics Education Research 21, 010126 (2025)

  13. [21]

    R. K. Fussell, M. Flynn, A. Damle, M. F. Fox, and N. Holmes, Comparing large language models for su- pervised analysis of students’ lab notes, Physical Review Physics Education Research 21, 010128 (2025)

  14. [22]

    Gregorcic, G

    B. Gregorcic, G. Polverini, and A. Sarlah, Chatgpt as a tool for honing teachers’ socratic dialogue skills, Physics Education 59, 045005 (2024)

  15. [23]

    Crawford, M

    J. Crawford, M. Cowling, and K.-A. Allen, Leadership is needed for ethical chatgpt: character, assessment, and learning using artificial intelligence (ai), Journal of Uni- versity Teaching and Learning Practice 20 (2023)

  16. [24]

    M. A. R. Vasconcelos and R. P. Dos Santos, Enhanc- ing stem learning with chatgpt and bing chat as objects to think with: a case study, Eurasia Journal of Mathe- matics, Science and Technology Education 19, em2296 (2023)

  17. [25]

    M. N. Dahlkemper, S. Z. Lahme, and P. Klein, How do physics students evaluate artificial intelligence responses on comprehension questions? a study on the perceived scientific accuracy and linguistic quality of ChatGPT, Phys. Rev. Phys. Educ. Res. 19, 010142 (2023)

  18. [26]

    L. Ding, T. Li, S. Jiang, and A. Gapud, Students’ per- ceptions of using ChatGPT in a physics class as a virtual tutor, International Journal of Educational Technology in Higher Education 20, 63 (2023)

  19. [27]

    C. G. West, AI and the FCI: Can ChatGPT project an understanding of introductory physics? (2023), arXiv:2303.01067 [physics.ed-ph]

  20. [28]

    Wheeler and R

    S. Wheeler and R. E. Scherr, Chatgpt reflects stu- dent misconceptions in physics, in Proceedings of the Physics Education Research Conference (PERC) (2023) pp. 386–390

  21. [29]

    N. Cho, An investigation of using Spark generative AI in solving physics concept inventories in english and chi- nese: Performance and issues, Discover Artificial Intel- ligence 4, 1 (2024)

  22. [30]

    Aldazharova, G

    S. Aldazharova, G. Issayeva, S. Maxutov, and N. Balta, Assessing AI’s problem solving in physics: Analyz- ing reasoning, false positives and negatives through the force concept inventory, Contemporary Educational Technology 16, ep538 (2024)

  23. [31]

    Polverini and B

    G. Polverini and B. Gregorcic, Evaluating vision- capable chatbots in interpreting kinematics graphs: a comparative study of free and subscription-based mod- els, in Frontiers in Education , Vol. 9 (Frontiers Media SA, 2024) p. 1452414

  24. [32]

    OpenAI, Hello GPT-4o, https://openai.com/index/ hello-gpt-4o/ (accessed June 2024)

  25. [33]

    Hestenes, M

    D. Hestenes, M. Wells, and G. Swackhamer, Force Con- cept Inventory, The Physics Teacher 30, 141 (1992), https://doi.org/10.1119/1.2343497

  26. [34]

    R. J. Beichner, Testing student interpretation of kine- matics graphs, American journal of Physics 62, 750 (1994)

  27. [35]

    L. Ding, R. Chabay, B. Sherwood, and R. Beichner, Evaluating an electricity and magnetism assessment tool: Brief electricity and magnetism assessment, Phys. Rev. ST Phys. Educ. Res. 2, 010105 (2006)

  28. [36]

    Polverini and B

    G. Polverini and B. Gregorcic, Performance of chatgpt on the test of understanding graphs in kinematics, Phys. Rev. Phys. Educ. Res. 20, 010109 (2024)

  29. [37]

    Polverini, J

    G. Polverini, J. Melin, E. ¨Onerud, and B. Gregorcic, Performance of chatgpt on tasks involving physics vi- sual representations: the case of the brief electricity and magnetism assessment, arXiv preprint arXiv:2412.10019 18 (2024)

  30. [38]

    J. I. Smith and K. Tanner, The problem of revealing how students think: concept inventories and beyond, CBE—Life Sciences Education 9, 1 (2010)

  31. [39]

    Sands, M

    D. Sands, M. Parker, H. Hedgeland, S. Jordan, and R. Galloway, Using concept inventories to measure understanding, Higher Education Pedagogies 3, 173 (2018)

  32. [40]

    Henderson, Common concerns about the force con- cept inventory, The Physics Teacher 40, 542 (2002)

    C. Henderson, Common concerns about the force con- cept inventory, The Physics Teacher 40, 542 (2002)

  33. [41]

    OpenAI, Introducing GPT-o1, https://openai.com/ o1/ (accessed January 2025)

  34. [42]

    Geisler, Quality metrics for automated evaluation of exercises within student-LLM dialogues, Unpublished M.Sc

    T. Geisler, Quality metrics for automated evaluation of exercises within student-LLM dialogues, Unpublished M.Sc. thesis, ETH Zurich (2025)

  35. [43]

    Gregorcic and A.-M

    B. Gregorcic and A.-M. Pendrill, ChatGPT and the frustrated Socrates, Physics Education 58, 035021 (2023)

  36. [44]

    Madsen, S

    A. Madsen, S. B. McKagan, and E. C. Sayre, Best prac- tices for administering concept inventories, The Physics Teacher 55, 530 (2017)

  37. [45]

    R. R. Hake, Interactive-engagement versus traditional methods: A six-thousand-student survey of mechanics test data for introductory physics courses, American journal of Physics 66, 64 (1998)

  38. [46]

    J. T. Laverty and M. D. Caballero, Analysis of the most common concept inventories in physics: What are we assessing?, Physical Review Physics Education Research 14, 010123 (2018)

  39. [47]

    S. M. Stoen, M. A. McDaniel, R. F. Frey, K. M. Hynes, and M. J. Cahill, Force concept inventory: More than just conceptual understanding, Physical Review Physics Education Research 16, 010105 (2020)

  40. [48]

    D. T. Brookes and E. Etkina, Using conceptual metaphor and functional grammar to explore how lan- guage used in physics affects student learning, Physical Review Special Topics—Physics Education Research 3, 010105 (2007)

  41. [49]

    Euler, E

    E. Euler, E. R ˚ adahl, and B. Gregorcic, Embodiment in physics learning: A social-semiotic look, Physical Re- view Physics Education Research 15, 010134 (2019)

  42. [50]

    D. T. Brookes, The role of language in learning physics , Ph.D. thesis, Rutgers University (2006)

  43. [51]

    P. Wulff, Physics language and language use in physics—what do we know and how ai might enhance language-related research and instruction, European Journal of Physics 45, 023001 (2024)

  44. [52]

    OpenAI, GPT-4, https://openai.com/index/ gpt-4-research/ (accessed December 2024)

  45. [53]

    Nicholas and A

    G. Nicholas and A. Bhatia, Lost in translation: large language models in non-english content analysis, arXiv preprint arXiv:2306.07377 (2023)

  46. [54]

    Deepseek, https://www.deepseek.com/ (accessed De- cember 2024)

  47. [55]

    Alibaba Cloud, Qwen, https://qwen-ai.com/ (ac- cessed December 2024)

  48. [56]

    AI, Swiss AI Initiative, https://www.swiss-ai.org/ (retrieved January 2025)

    S. AI, Swiss AI Initiative, https://www.swiss-ai.org/ (retrieved January 2025)

  49. [57]

    com/research/papers/the-ai-language-gap.pdf (ac- cessed December 2024)

    Cohere for AI, The AI language gap, https://cohere. com/research/papers/the-ai-language-gap.pdf (ac- cessed December 2024)

  50. [58]

    S. Feng, W. Shi, Y. Wang, W. Ding, O. Ahia, S. S. Li, V. Balachandran, S. Sitaram, and Y. Tsvetkov, Teach- ing LLMs to abstain across languages via multilingual feedback, arXiv preprint arXiv:2406.15948 (2024)

  51. [59]

    K. T. Kotsis, Chatgpt as teacher assistant for physics teaching, EIKI Journal of Effective Teaching Methods 2, https://doi.org/10.59652/jetm.v2i4.283 (2024)

  52. [60]

    Tschisgale, P

    P. Tschisgale, P. Wulff, and M. Kubsch, Integrating ar- tificial intelligence-based methods into qualitative re- search in physics education research: A case for compu- tational grounded theory, Phys. Rev. Phys. Educ. Res. 19, 020123 (2023)

  53. [61]

    Tschisgale, P

    P. Tschisgale, P. Wulff, and M. Kubsch, Erratum: Inte- grating artificial intelligence-based methods into quali- tative research in physics education research: A case for computational grounded theory [phys. rev. phys. educ. res. 19, 020123 (2023)], Phys. Rev. Phys. Educ. Res....

  54. [62]

    T. O. B. Odden, H. Tyseng, J. T. Mjaaland, M. F. Kreutzer, and A. Malthe-Sørenssen, Using text em- beddings for deductive qualitative research at scale in physics education, Phys. Rev. Phys. Educ. Res. 20, 020151 (2024)

  55. [63]

    Bakhtin, L

    A. Bakhtin, L. van der Maaten, J. Johnson, L. Gustafson, and R. Girshick, Phyre: A new bench- mark for physical reasoning, Advances in Neural Infor- mation Processing Systems 32 (2019)

  56. [64]

    C. Xue, V. Pinto, C. Gamage, E. Nikonova, P. Zhang, and J. Renz, Phy-q: A benchmark for physical reason- ing, arXiv preprint arXiv:2108.13696 3 (2021)

  57. [65]

    Melnik, R

    A. Melnik, R. Schiewer, M. Lange, A. Muresanu, M. Saeidi, A. Garg, and H. Ritter, Benchmarks for physical reasoning ai, arXiv preprint arXiv:2312.10728 (2023)

  58. [66]

    Physbench: Physical reasoning benchmark, https:// physbench.com/, accessed: 2025-05-07

  59. [67]

    Kortemeyer and J

    G. Kortemeyer and J. N¨ ohl, Assessing confidence in ai- assisted grading of physics exams through psychomet- rics: An exploratory study, Physical Review Physics Ed- ucation Research 21, 010136 (2025)

  60. [68]

    R. Mok, F. Akhtar, L. Clare, C. Li, J. Ida, L. Ross, and M. Campanelli, Using large language models for grad- ing in education: an applied test for physics, Physics Education 60, 035006 (2025)

  61. [69]

    S. Guo, E. Latif, Y. Zhou, X. Huang, and X. Zhai, Us- ing generative AI and multi-agents to provide automatic feedback, arXiv preprint arXiv:2411.07407 (2024)

  62. [70]

    Krupp, J

    L. Krupp, J. Bley, I. Gobbi, A. Geng, S. M¨ uller, S. Suh, A. Moghiseh, A. C. Medina, V. Bartsch, A. Widera, et al. , Llm-generated tips rival expert-created tips in helping students answer quantum-computing questions, EPJ Quantum Technology 12, 33 (2025)

  63. [71]

    J. R. Aguilar-Mej ´ ıa, S. Tejeda, C. V. Ramirez-Lopez, and C. L. Garay-Rondero, Design and use of a chatbot for learning selected topics of physics, Transactions on Computer Systems and Networks , 175–188 (2022)

  64. [72]

    V. R. Lee, D. Pope, S. Miles, and R. C. Zarate, Cheat- ing in the age of generative ai: A high school survey study of cheating behaviors before and after the release of chatgpt, Computers and Education: Artificial Intel- ligence 7, 100253 (2024)

  65. [73]

    J. L. Docktor and J. P. Mestre, Synthesis of discipline- based education research in physics, Physical Review Special Topics-Physics Education Research 10, 020119 (2014)

  66. [74]

    D. E. Meltzer and V. K. Otero, A brief history of physics education in the united states, American Journal of 19 Physics 83, 447 (2015)

  67. [75]

    Bulathwela, M

    S. Bulathwela, M. P´ erez-Ortiz, C. Holloway, M. Cukurova, and J. Shawe-Taylor, Artificial in- telligence alone will not democratise education: on educational inequality, techno-solutionism and inclusive tools, Sustainability 16, 781 (2024)

  68. [76]

    Liang, D

    Y. Liang, D. Zou, H. Xie, and F. L. Wang, Exploring the potential of using chatgpt in physics education, Smart Learning Environments 10, 52 (2023)

  69. [77]

    Latif, R

    E. Latif, R. Parasuraman, and X. Zhai, Physicsassis- tant: An LLM-powered interactive learning robot for physics lab investigations, in 2024 33rd IEEE Inter- national Conference on Robot and Human Interactive Communication (ROMAN) (IEEE, 2024) pp. 864–871

  70. [78]

    Lieb and T

    A. Lieb and T. Goel, Student interaction with newtbot: An LLM-as-tutor chatbot for secondary physics educa- tion, in Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (2024) pp. 1–8

  71. [79]

    Kahaleh and V

    R. Kahaleh and V. Lopez, Evaluating large language models in high school physics education: addressing misconceptions and fostering conceptual understanding, Physics Education 60, 025013 (2025)

  72. [80]

    K. E. Avila, S. Steinert, S. Ruzika, J. Kuhn, and S. K¨ uchemann, Using ChatGPT for teaching physics, The Physics Teacher 62, 536 (2024)

  73. [81]

    Zhu, Z.-Y

    Y. Zhu, Z.-Y. Khoo, J. S. C. Low, and S. Bressan, A personalised learning tool for physics undergraduate students built on a large language model for symbolic regression, in 2024 IEEE Conference on Artificial Intel- ligence (CAI) (IEEE, 2024) pp. 38–43

  74. [82]

    American Association of Physics Teachers, Physport, https://www.physport.org (2017), [retrieved Novem- ber 2024]

  75. [83]

    S. B. McKagan, L. E. Strubbe, L. J. Barbato, B. A. Mason, A. M. Madsen, and E. C. Sayre, PhysPort use and growth: Supporting physics teaching with research- based resources since 2011, The Physics Teacher58, 465 (2020)

  76. [84]

    Von Korff, B

    J. Von Korff, B. Archibeque, K. A. Gomez, T. Heck- endorf, S. B. McKagan, E. C. Sayre, E. W. Schenk, C. Shepherd, and L. Sorell, Secondary analysis of teach- ing methods in introductory physics: A 50 k-student study, American Journal of physics 84, 969 (2016)

  77. [85]

    Hufnagel, Development of the astronomy diagnostic test, Astronomy Education Review 1, 47 (2002)

    B. Hufnagel, Development of the astronomy diagnostic test, Astronomy Education Review 1, 47 (2002)

  78. [86]

    Brogt, D

    E. Brogt, D. Sabers, E. E. Prather, G. L. Deming, B. Hufnagel, and T. F. Slater, Analysis of the astron- omy diagnostic test, Astronomy Education Review 6, 25 (2007)

  79. [87]

    N. O. Koca, N. alhuda Al Saqri, H. Al Hamrashdi, and N. Al Kindi, Evaluating the students’ learning on the electricity and magnetism using a conceptual survey bema, Physics Education 60, 015022 (2024)

  80. [88]

    S. J. Pollock, Comparing student learning with multiple research-based conceptual surveys: Csem and bema., in AIP Conference Proceedings, Vol. 1064 (American In- stitute of Physics, 2008) pp. 171–174

  81. [89]

    Epstein, The calculus concept inventory, National STEM Assessment, Washington, DC , 60 (2006)

    J. Epstein, The calculus concept inventory, National STEM Assessment, Washington, DC , 60 (2006)

  82. [90]

    Maciejewski, Flipping the calculus classroom: an evaluative study, Teaching Mathematics and its Appli- cations: An International Journal of the IMA 35, 187 (2016)

    W. Maciejewski, Flipping the calculus classroom: an evaluative study, Teaching Mathematics and its Appli- cations: An International Journal of the IMA 35, 187 (2016)

  83. [91]

    Day and D

    J. Day and D. Bonn, Development of the concise data processing assessment, Phys. Rev. ST Phys. Educ. Res. 7, 010114 (2011)

  84. [92]

    D. P. Maloney, T. L. O’Kuma, C. J. Hieggelke, and A. Van Heuvelen, Surveying students’ conceptual knowledge of electricity and magnetism, American Jour- nal of Physics 69, S12 (2001)

  85. [93]

    Tapping, G

    R. Tapping, G. Lepage, and N. Holmes, Visualizing pat- terns in csem responses to assess student conceptual un- derstanding, in 2018 Physics Education Research Con- ference (PERC) (2019) pp. 419–422

  86. [94]

    A. E. Lawson, The development and validation of a classroom test of formal reasoning., Journal of Research in Science Teaching (1978)

  87. [95]

    J. C. Moore and L. J. Rubbo, Scientific reasoning abilities of nonscience majors in physics-based courses, Physical Review Special Topics—Physics Education Re- search 8, 010106 (2012)

  88. [96]

    P. V. Engelhardt and R. J. Beichner, Students’ under- standing of direct current resistive electrical circuits, American journal of physics 72, 98 (2004)

  89. [97]

    Sangam and B

    D. Sangam and B. K. Jesiek, Conceptual understanding of resistive electric circuits among first-year engineering students, in 2012 ASEE Annual Conference & Exposi- tion (2012) pp. 25–339

  90. [98]

    Yeend, M

    R. Yeend, M. Loverude, and B. Gonzales, Student un- derstanding of density: a cross-age investigation, in Physics Education Research Conference (2001)

  91. [99]

    Zenger and P

    T. Zenger and P. Bitzenbauer, Exploring german sec- ondary school students’ conceptual knowledge of den- sity, Science Education International 33, 86 (2022)

  92. [100]

    L. Ding, R. Chabay, and B. Sherwood, How do stu- dents in an innovative principle-based mechanics course understand energy concepts?, Journal of research in sci- ence teaching 50, 722 (2013)

  93. [101]

    D. R. Sokoloff, Teaching electric circuit concepts us- ing microcomputer-based current/voltage probes, in Microcomputer–based labs: Educational research and standards (Springer, 1996) pp. 129–146

  94. [102]

    Kortemeyer, D

    G. Kortemeyer, D. Anderson, A. M. Desrochers, A. Hackbardt, K. Hoekstra, A. Holt, A. Iftekhar, T. Kabaker, N. Keller, Z. Korzecke, et al., Using a com- puter game to teach circuit concepts, European Journal of Physics 40, 055703 (2019)

  95. [103]

    M. W. McColgan, R. A. Finn, D. L. Broder, and G. E. Hassel, Assessing students’ conceptual knowledge of electricity and magnetism, Physical Review Physics Ed- ucation Research 13, 020121 (2017)

  96. [104]

    Singh and D

    C. Singh and D. Rosengrant, Multiple- choice test of energy and momentum con- cepts, American Journal of Physics 71, 607 (2003), https://pubs.aip.org/aapt/ajp/article- pdf/71/6/607/7531054/607 1 online.pdf

  97. [105]

    M. Sahin, The impact of problem-based learning on en- gineering students’ beliefs about physics and concep- tual understanding of energy and momentum, European Journal of Engineering Education 35, 519 (2010)

  98. [106]

    A. J. Mason, Learning goals and perceived irrelevance to major within life science majors in introductory physics, arXiv preprint arXiv:2012.09898 (2020)

  99. [107]

    Kortemeyer, Gender differences in the use of an on- line homework system in an introductory physics course, Phys

    G. Kortemeyer, Gender differences in the use of an on- line homework system in an introductory physics course, Phys. Rev. ST Phys. Educ. Res. 5, 010107 (2009). 20

  100. [108]

    J. Han, L. Bao, L. Chen, T. Cai, Y. Pi, S. Zhou, Y. Tu, and K. Koenig, Dividing the force concept inventory into two equivalent half-length tests, Phys. Rev. ST Phys. Educ. Res. 11, 010112 (2015)

  101. [109]

    R. K. Thornton and D. R. Sokoloff, Assessing student learning of newton’s laws: The force and motion con- ceptual evaluation and the evaluation of active learning laboratory and lecture curricula, American Journal of Physics 66, 338 (1998)

  102. [110]

    Cummings, J

    K. Cummings, J. Marx, R. Thornton, and D. Kuhl, Evaluating innovation in studio physics, American jour- nal of physics 67, S38 (1999)

  103. [111]

    S. T. Kalinowski and S. Willoughby, Development and validation of a scientific (formal) reasoning test for col- lege students, Journal of Research in Science Teaching 56, 1269 (2019)

  104. [112]

    Kaltakci-Gurel, A

    D. Kaltakci-Gurel, A. Eryilmaz, and L. C. McDermott, Development and application of a four-tier test to as- sess pre-service physics teachers’ misconceptions about geometrical optics, ReseaRch in science & Technological educaTion 35, 238 (2017)

  105. [113]

    Rosenblatt and A

    R. Rosenblatt and A. F. Heckler, Systematic study of student understanding of the relationships between the directions of force, velocity, and acceleration in one di- mension, Physical Review Special Topics - Physics Ed- ucation Research 7, 020112 (2011)

  106. [114]

    J. M. Keller, Part I. development of a concept inven- tory addressing students’ beliefs and reasoning difficul- ties regarding the greenhouse effect, part II. distribution of chlorine measured by the mars odyssey gamma ray spectrometer (The University of Arizona, 2006)

  107. [115]

    Tanahoung, M

    C. Tanahoung, M. D. Sharma, I. D. Johnston, R. Chita- ree, and C. Soankwan, Surveying sydney introductory physics students’ understandings of heat and temper- ature, in Australian Institute of Physics 17th National Congress, Brisbane, Paper No. WC0233 (2006)

  108. [116]

    Halloun, Evaluation of the impact of the new physics curriculum on the conceptual profiles of secondary stu- dents

    I. Halloun, Evaluation of the impact of the new physics curriculum on the conceptual profiles of secondary stu- dents. 1–25, https://www.halloun.net/wp-content/ uploads/2016/10/LU-Summative-Report-10-07.pdf (2007)

  109. [117]

    Ndihokubwayo, J

    K. Ndihokubwayo, J. Uwamahoro, I. Ndayambaje, and M. Ralph, Light phenomena conceptual assessment: an inventory tool for teachers, Physics Education 55, 035009 (2020)

  110. [118]

    R. S. Lindell and J. P. Olsen, Developing the lunar phases concept inventory, in Proceedings of the 2002 Physics Education Research Conference (New York: PERC Publishing, 2002)

  111. [119]

    E. M. Bardar, E. E. Prather, K. Brecher, and T. F. Slater, Development and validation of the light and spectroscopy concept inventory, Astronomy Education Review 5, 103 (2007)

  112. [120]

    C. S. Wallace, T. G. Chambers, and E. E. Prather, Item response theory evaluation of the light and spectroscopy concept inventory national data set, Physical Review Physics Education Research 14, 010149 (2018)

  113. [121]

    Hestenes and M

    D. Hestenes and M. Wells, A mechanics baseline test, The physics teacher 30, 159 (1992)

  114. [122]

    C. P. Mill´ an and S. Otranto, Thirty-six years of the forced concept inventory and the mechanics baseline test: is aristotle still playing hide and seek in our class- rooms?, Latin-American Journal of Physics Education 15, 9 (2021)

  115. [123]

    Antwi, R

    V. Antwi, R. Hanson, A. Sam, E. Savelsbergh, and H. Eijkelhof, The impact of interactive-engagement (ie) teaching on students understanding of concepts in me- chanics: The use of force concept inventory (fci) and mechanics baseline test (mbt), International Journal of Educatio...

  116. [124]

    K´ ad´ ar and P

    C. K´ ad´ ar and P. Tasn´ adi, The knowledge of hungarian students in the light of the mechanics baseline test, in Journal of Physics: Conference Series , Vol. 1286 (IOP Publishing, 2019) p. 012026

  117. [125]

    Li and C

    J. Li and C. Singh, Developing and validating a concep- tual survey to assess introductory physics students’ un- derstanding of magnetism, European Journal of Physics 38, 025702 (2016)

  118. [126]

    D. L. Deardorff, Introductory physics students’ treat- ment of measurement uncertainty (North Carolina State University, 2001)

  119. [127]

    Tongchai, M

    A. Tongchai, M. D. Sharma, I. D. Johnston, K. Aray- athanitkul, and C. Soankwan, Developing, evaluating and demonstrating the use of a conceptual survey in mechanical waves, International Journal of Science Ed- ucation 31, 2437 (2009)

  120. [128]

    P. H. Santoso, E. Istiyono, and H. Haryanto, Princi- pal component analysis and exploratory factor analysis of the mechanical waves conceptual survey, JP3I (Ju- rnal Pengukuran Psikologi dan Pendidikan Indonesia) 11, 209 (2022)

  121. [129]

    K. E. Williamson, S. D. Willoughby, and E. E. Prather, Development of the Newtonian gravity concept inven- tory, Astron. Educ. Rev. 12, 1 (2013)

  122. [130]

    P. V. Engelhardt, S. Robinson, E. P. Price, P. S. Smith, and F. Goldberg, Developing a conceptual assessment for a modular curriculum, in Physics Education Re- search Conference (2018)

  123. [131]

    White Brahmia, A

    S. White Brahmia, A. Olsho, T. I. Smith, A. Boudreaux, P. Eaton, and C. Zimmerman, Physics inventory of quantitative literacy: A tool for assessing mathemati- cal reasoning in introductory physics, Physical Review Physics Education Research 17, 020129 (2021)

  124. [132]

    H. R. Sadaghiani and S. J. Pollock, Quantum mechanics concept assessment: Development and validation study, Physical Review Special Topics-Physics Education Re- search 11, 010110 (2015)

  125. [133]

    McKagan, K

    S. McKagan, K. Perkins, and C. Wieman, Design and validation of the quantum mechanics conceptual survey, Physical Review Special Topics – Physics Education Re- search 6, 020121 (2010)

  126. [134]

    Marshman and C

    E. Marshman and C. Singh, Validation and administra- tion of a conceptual survey on the formalism and pos- tulates of quantum mechanics, Physical Review Physics Education Research 15, 020128 (2019)

  127. [135]

    E. M. Marshman, Improving the quantum mechanics content knowledge and pedagogical content knowledge of physics graduate students , Ph.D. thesis, University of Pittsburgh (2015)

  128. [136]

    Zhu and C

    G. Zhu and C. Singh, Surveying students’ understand- ing of quantum mechanics in one spatial dimension, American Journal of Physics 80, 252 (2012)

  129. [137]

    Cataloglu and R

    E. Cataloglu and R. Robinett, Testing the development of student conceptual and visualization understanding in quantum mechanics through the undergraduate ca- reer, American Journal of Physics 70, 238 (2002)

  130. [138]

    Wuttiprom, M

    S. Wuttiprom, M. D. Sharma, I. D. Johnston, R. Chita- ree, and C. Soankwan, Development and use of a con- 21 ceptual survey in introductory quantum physics, Inter- national Journal of Science Education 31, 631 (2009)

  131. [139]

    R. J. Allain, Investigating the relationship between stu- dent difficulties with the concept of electric potential and the concept of rate of change (North Carolina State Uni- versity, 2001)

  132. [140]

    Aslanides and C

    J. Aslanides and C. M. Savage, Relativity concept in- ventory: Development, analysis, and results, Physical Review Special Topics – Physics Education Research 9, 010118 (2013)

  133. [141]

    Nieminen, A

    P. Nieminen, A. Savinainen, and J. Viiri, Force con- cept inventory-based multiple-choice test for investigat- ing students’ representational consistency, Physical Re- view Special Topics – Physics Education Research 6, 020109 (2010)

  134. [142]

    Mashood and V

    K. Mashood and V. A. Singh, An inventory on ro- tational kinematics of a particle: unravelling miscon- ceptions and pitfalls in reasoning, European Journal of Physics 33, 1301 (2012)

  135. [143]

    Su´ arez, S

    M. Su´ arez, S. Pandiella, and J. Benegas, Tutorials+ PhET: a simple and efficient active-learning approach for the teaching of kinematics of circular motion in a technically-oriented high school, Physics Education 58, 035005 (2023)

  136. [144]

    L. G. Rimoldini and C. Singh, Student understanding of rotational and rolling motion concepts, Physical Review Special Topics – Physics Education Research 1, 010102 (2005)

  137. [145]

    Singh, Student understanding of symmetry and Gauss’s law of electricity, American journal of physics 74, 923 (2006)

    C. Singh, Student understanding of symmetry and Gauss’s law of electricity, American journal of physics 74, 923 (2006)

  138. [146]

    J. M. Bailey, B. Johnson, E. E. Prather, and T. F. Slater, Development and validation of the star proper- ties concept inventory, International Journal of Science Education 34, 2257 (2012)

  139. [147]

    Brown and C

    B. Brown and C. Singh, Development and validation of a conceptual survey instrument to evaluate stu- dents’ understanding of thermodynamics, Physical Re- view Physics Education Research 17, 010104 (2021)

  140. [148]

    Yeo and M

    S. Yeo and M. Zadnik, Introductory thermal concept evaluation: Assessing students’ understanding, The Physics Teacher 39, 496 (2001)

  141. [149]

    Wattanakasiwich, P

    P. Wattanakasiwich, P. Taleab, M. D. Sharma, and I. D. Johnston, Construction and implementation of a con- ceptual survey in thermodynamics, International Jour- nal of Innovation in Science and Mathematics Education 21 (2013)

  142. [150]

    S. J. Slater, The development and validation of the test of astronomy standards (TOAST)., Journal of Astron- omy & Earth Sciences Education 1, 1 (2014)

  143. [151]

    Klein, A

    P. Klein, A. Lichtenberger, S. K¨ uchemann, S. Becker, M. Kekule, J. Viiri, C. Baadte, A. Vaterlaus, and J. Kuhn, Visual attention while solving the test of un- derstanding graphs in kinematics: an eye-tracking anal- ysis, European Journal of Physics 41, 025701 (2020)

  144. [152]

    Barniol and G

    P. Barniol and G. Zavala, Test of understanding of vectors: A reliable multiple-choice vector concept test, Physical Review Special Topics-Physics Education Re- search 10, 010121 (2014)

  145. [153]

    OpenAI, How ChatGPT and our foundation models are developed, https://help.openai.com/en/articles/ 7842364-how-chatgpt-and-our-foundation-models-are-developed (accessed March 2025)

  146. [154]

    microsoft.com/en-us/products/ai-services (ac- cessed June 2024)

    Microsoft, Azure AI Services, https://azure. microsoft.com/en-us/products/ai-services (ac- cessed June 2024)

  147. [155]

    Danil´ ak, langdetect, https://pypi.org/project/ langdetect/ (accessed December 2024)

    M. Danil´ ak, langdetect, https://pypi.org/project/ langdetect/ (accessed December 2024)

  148. [156]

    Sidiropoulos, S

    N. Sidiropoulos, S. H. Sohi, T. L. Pedersen, B. T. Porse, O. Winther, N. Rapin, and F. O. Bagger, Sinaplot: an enhanced chart for simple and truthful representation of single observations over multiple classes, Journal of Computational and Graphical Statistics 27, 673 (2018)

  149. [157]

    Kieser, P

    F. Kieser, P. Wulff, J. Kuhn, and S. K¨ uchemann, Ed- ucational data augmentation in physics education re- search using ChatGPT, Phys. Rev. Phys. Educ. Res. 19, 020150 (2023)

  150. [158]

    com/index/introducing-o3-and-o4-mini/ (accessed May 2025)

    OpenAI, OpenAI o3 and o4-mini, https://openai. com/index/introducing-o3-and-o4-mini/ (accessed May 2025)

  151. [159]

    multilin- gual performance of a multimodal artificial intelligence system on multisubject physics concept inventories

    G. Kortemeyer, M. Babayeva, G. Polverini, R. Widen- horn, and B. Gregorcic, Data for the paper “multilin- gual performance of a multimodal artificial intelligence system on multisubject physics concept inventories”, https://www.physport.org/XXXXX (2025). 22 | | | | | | | | | |...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.