An LLM agent with image-reading and answer-review tools scored 23.5/30 on IPhO 2025 theory problems, matching the median gold-medalist theory score.
Phybench: Holistic evaluation of physical perception and reasoning in large language models, 2025
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
baseline 1
citation-polarity summary
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1roles
baseline 1polarities
baseline 1representative citing papers
citing papers explorer
-
Physics Supernova: AI Agent Matches Elite Gold Medalists at IPhO 2025
An LLM agent with image-reading and answer-review tools scored 23.5/30 on IPhO 2025 theory problems, matching the median gold-medalist theory score.