Most large reasoning models overthink easy questions, and Think-Bench provides a benchmark and metrics to quantify this inefficiency and the quality of their chain-of-thought.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models
Most large reasoning models overthink easy questions, and Think-Bench provides a benchmark and metrics to quantify this inefficiency and the quality of their chain-of-thought.