A new 971-question chemistry benchmark shows that even the best large language models, given full context, still fail on many multi-step reasoning questions.
Graph of thoughts: Solving elaborate problems with large language models
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Evaluating Multi-Hop Reasoning in Large Language Models: A Chemistry-Centric Case Study
A new 971-question chemistry benchmark shows that even the best large language models, given full context, still fail on many multi-step reasoning questions.