On the MaScQA benchmark, Claude-3.5-Sonnet and GPT-4o achieve about 84 percent accuracy, while the best open-source models (Llama3-70b, Phi3-14b) reach about 56 and 43 percent.
MCQ tasks, while simpler, can be impacted by pattern exploitation where models rely on super- ficial cues rather than true conceptual understanding
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
physics.comp-ph 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Exploring the Expertise of Large Language Models in Materials Science and Metallurgical Engineering
On the MaScQA benchmark, Claude-3.5-Sonnet and GPT-4o achieve about 84 percent accuracy, while the best open-source models (Llama3-70b, Phi3-14b) reach about 56 and 43 percent.