ChatGPT-4o scores 67% on BEMA when answers are coded by meaning, above the student average of 53.4%, but systematically fails on right-hand-rule and spatial-coordination items.
However, we can see that many errors made by the chatbot would be quite atypical for human students, based on our experience as physics instructors
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
physics.ed-ph 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
Performance of ChatGPT on tasks involving physics visual representations: the case of the Brief Electricity and Magnetism Assessment
ChatGPT-4o scores 67% on BEMA when answers are coded by meaning, above the student average of 53.4%, but systematically fails on right-hand-rule and spatial-coordination items.