On the GPQA Diamond benchmark, chain-of-thought prompting gives only small average accuracy gains for non-reasoning models, little or no gain for reasoning models, and large increases in response time and tokens.
easy" questions that the model would otherwise get right, harming performance on the
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Prompting Science Report 2: The Decreasing Value of Chain of Thought in Prompting
On the GPQA Diamond benchmark, chain-of-thought prompting gives only small average accuracy gains for non-reasoning models, little or no gain for reasoning models, and large increases in response time and tokens.