ChatGPT 4.1 mini, Gemini 2.5 Flash, Claude 4.0 Sonnet, and DeepSeek R1 average 82–92% on AP Physics 1/2 free-response questions but systematically fail spatial, visual, and conceptual tasks.
Survey: Automatic generation of attack trees and attack graphs,
1 Pith paper cite this work, alongside 45 external citations. Polarity classification is still indexing.
1
Pith paper citing it
45
external citations · external index
fields
physics.ed-ph 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
How Well Do AI Systems Solve AP Physics? A Comparative Evaluation of Large Language Models on Algebra-Based Free Response Questions
ChatGPT 4.1 mini, Gemini 2.5 Flash, Claude 4.0 Sonnet, and DeepSeek R1 average 82–92% on AP Physics 1/2 free-response questions but systematically fail spatial, visual, and conceptual tasks.