On the GPQA Diamond benchmark, chain-of-thought prompting gives only small average accuracy gains for non-reasoning models, little or no gain for reasoning models, and large increases in response time and tokens.
Miller E (2024) Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Prompting Science Report 2: The Decreasing Value of Chain of Thought in Prompting
On the GPQA Diamond benchmark, chain-of-thought prompting gives only small average accuracy gains for non-reasoning models, little or no gain for reasoning models, and large increases in response time and tokens.