Pith. sign in

REVIEW 7 cited by

Break the Chain: Large Language Models Can be Shortcut Reasoners

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.06580 v1 pith:CVMYZLMZ submitted 2024-06-04 cs.CL cs.AI

classification cs.CLcs.AI
keywords reasoningstrategiesbreakchainshortcutscomplexlanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent advancements in Chain-of-Thought (CoT) reasoning utilize complex modules but are hampered by high token consumption, limited applicability, and challenges in reproducibility. This paper conducts a critical evaluation of CoT prompting, extending beyond arithmetic to include complex logical and commonsense reasoning tasks, areas where standard CoT methods fall short. We propose the integration of human-like heuristics and shortcuts into language models (LMs) through "break the chain" strategies. These strategies disrupt traditional CoT processes using controlled variables to assess their efficacy. Additionally, we develop innovative zero-shot prompting strategies that encourage the use of shortcuts, enabling LMs to quickly exploit reasoning clues and bypass detailed procedural steps. Our comprehensive experiments across various LMs, both commercial and open-source, reveal that LMs maintain effective performance with "break the chain" strategies. We also introduce ShortcutQA, a dataset specifically designed to evaluate reasoning through shortcuts, compiled from competitive tests optimized for heuristic reasoning tasks such as forward/backward reasoning and simplification. Our analysis confirms that ShortcutQA not only poses a robust challenge to LMs but also serves as an essential benchmark for enhancing reasoning efficiency in AI.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. NeoQA: Evidence-based Question Answering with Generated News Events

    cs.CL 2025-05 conditional novelty 7.0 of 10

    A benchmark built from fictional news timelines shows LLMs frequently answer with shortcuts instead of deflecting when evidence is insufficient.

  2. InteChar: A Unified Oracle Bone Character List for Ancient Chinese Language Modeling

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    InteChar is a proposed standard character list for digitizing oracle bone script, paired with a corpus that reportedly improves ancient Chinese language modeling.

  3. CoT-Valve: Length-Compressible Chain-of-Thought Tuning

    cs.AI 2025-02 conditional novelty 6.0 of 10

    A single LoRA task vector, scaled up or down at inference, controls chain-of-thought length in LLMs and compresses reasoning tokens with little accuracy loss.

  4. Not All Errors Are Created Equal: ASCoT Addresses Late-Stage Fragility in Efficient LLM Reasoning

    cs.CL 2025-08 reject novelty 5.0 of 10

    ASCoT claims later reasoning errors are more harmful than early ones and uses a position-weighted verifier to prune and correct CoT steps, but its key evidence is internally inconsistent.

  5. Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Reward shaping with a powered length penalty makes LLMs answer easy questions with far fewer tokens while preserving or slightly improving accuracy on hard math benchmarks.

  6. Thought Manipulation: External Thought Can Be Efficient for Large Reasoning Models

    cs.CL 2025-04 conditional novelty 5.0 of 10

    Injecting external chain-of-thought text into a reasoning model's thinking tokens makes it skip internal reasoning, cutting output tokens by 30-90% with moderate accuracy trade-offs.

  7. Generative AI Act II: Test Time Scaling Drives Cognition Engineering

    cs.CL 2025-04 conditional novelty 3.0 of 10

    Test-time scaling techniques such as long chain-of-thought, tree search, and self-correction define the paper's 'cognition engineering' paradigm, which it surveys, taxonomizes, and tutorials.

Pith tools