D-GPTScore, which averages GPT-4o's per-aspect ratings of concept-customized images, correlates with human preference at 0.78 Pearson on the new CC-AlignBench, beating prior metrics.
Do LLMs Really Think Step-by-step In Implicit Reasoning?
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
It has been well-known that Chain-of-Thought can remarkably enhance LLMs' performance on complex tasks. However, because it also introduces slower inference speeds and higher computational costs, many researches have attempted to use implicit CoT, which does not need LLMs to explicitly generate the intermediate steps. However, the invisible reasoning process leaves us a doubt that, can implicit CoT really be equal to explicit CoT? Therefore, in this study, we address this question through experiments. We probe the information of intermediate steps from the model's hidden states when it is either trained or prompted to perform implicit CoT. The results surprisingly indicate that when prompted, LLMs hardly think about intermediate steps, suggesting they may just rely on experience rather than strict step-by-step reasoning. But when trained, they indeed calculate intermediate steps. Moreover, in both situations, we find the effect of using implicit CoT is susceptible to the format of the problem, reaffirming the current deficiency of implicit CoT.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation
D-GPTScore, which averages GPT-4o's per-aspect ratings of concept-customized images, correlates with human preference at 0.78 Pearson on the new CC-AlignBench, beating prior metrics.