On the GPQA Diamond benchmark, chain-of-thought prompting gives only small average accuracy gains for non-reasoning models, little or no gain for reasoning models, and large increases in response time and tokens.
This is our strictest condition for tasks where there is no room for errors
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Prompting Science Report 2: The Decreasing Value of Chain of Thought in Prompting
On the GPQA Diamond benchmark, chain-of-thought prompting gives only small average accuracy gains for non-reasoning models, little or no gain for reasoning models, and large increases in response time and tokens.