Pith. sign in

REVIEW 9 cited by

Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.21333 v4 pith:PCENNCEJ submitted 2024-10-27 cs.LG cs.AIcs.CLcs.CY

classification cs.LGcs.AIcs.CLcs.CY
keywords performancehumanstasksmodelsthinkingchain-of-thoughtcognitivedeliberation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Chain-of-thought (CoT) prompting has become a widely used strategy for improving large language and multimodal model performance. However, it is still an open question under which settings CoT systematically reduces performance. In this paper, we seek to identify the characteristics of tasks where CoT reduces performance by drawing inspiration from cognitive psychology, focusing on six representative tasks from the psychological literature where deliberation hurts performance in humans. In three of these tasks, state-of-the-art models exhibit significant performance drop-offs with CoT (up to 36.3\% absolute accuracy for OpenAI o1-preview compared to GPT-4o), while in others, CoT effects are mixed, with positive, neutral, and negative changes. While models and humans do not exhibit perfectly parallel cognitive processes, considering cases where thinking has negative consequences for humans helps identify settings where it negatively impacts models. By connecting the literature on human verbal thinking and deliberation with evaluations of CoT, we offer a perspective for understanding the impact of inference-time reasoning.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    An ensemble of three post-trained small specialist language models (knowledge, reasoning, coding) with calibrated confidence and abstention outperforms frontier reasoning models on high-precision missing-value predict...

  2. Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation?

    cs.AI 2025-08 conditional novelty 6.0 of 10

    Distilled-reasoning models actively depend on chain-of-thought for soft reasoning tasks, and a chain's causal influence can diverge from its explanatory faithfulness.

  3. The Other Mind: How Language Models Exhibit Human Temporal Cognition

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Larger LLMs develop a subjective 'present' around the current date, and their year similarity judgments follow a logarithmic Weber-Fechner compression, with supporting neural and representational evidence.

  4. Unveiling Confirmation Bias in Chain-of-Thought Reasoning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    LLMs exhibit confirmation bias in chain-of-thought: strong internal beliefs, approximated by direct answer probabilities, skew both reasoning generation and how the final answer is chosen, which helps explain why CoT ...

  5. Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    The monotonicity of token probabilities during initial decoding predicts chain-of-thought gains, enabling dynamic selection between CoT and direct answers.

  6. From Token to Action: State Machine Reasoning to Mitigate Overthinking in Information Retrieval

    cs.IR 2025-05 conditional novelty 6.0 of 10

    A state-machine framework that replaces token-level chain-of-thought with discrete query-refinement and reranking actions reduces token use by 74% while improving nDCG@10 on retrieval benchmarks.

  7. Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems

    cs.LG 2025-05 conditional novelty 6.0 of 10

    LLMs struggle to use passive observations for reverse engineering, but active intervention improves performance, largely through the process of generating queries rather than the data obtained.

  8. Data Shifts Hurt CoT: A Theoretical Study

    cs.LG 2025-06 reject novelty 5.0 of 10

    The authors derive a joint condition on distribution skew and poisoned reasoning labels that decides whether chain-of-thought training on k-parity succeeds, and they claim a paradox where maximal information leakage m...

  9. Knowing Before Saying: LLM Representations Encode Information About Chain-of-Thought Success Before Completion

    cs.CL 2025-05 conditional novelty 5.0 of 10

    LLM hidden states encode enough information to predict chain-of-thought success before any reasoning tokens are generated, outperforming a text-only classifier.

Pith tools