Pith. sign in

REVIEW 31 cited by

Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.15402 v3 pith:XE73OF3K submitted 2023-09-27 cs.CL cs.AI

classification cs.CLcs.AI
keywords futurereasoningresearchfrontiersintelligenceacademicsadvancedadvances
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Reasoning, a fundamental cognitive process integral to human intelligence, has garnered substantial interest within artificial intelligence. Notably, recent studies have revealed that chain-of-thought prompting significantly enhances LLM's reasoning capabilities, which attracts widespread attention from both academics and industry. In this paper, we systematically investigate relevant research, summarizing advanced methods through a meticulous taxonomy that offers novel perspectives. Moreover, we delve into the current frontiers and delineate the challenges and future directions, thereby shedding light on future research. Furthermore, we engage in a discussion about open questions. We hope this paper serves as an introduction for beginners and fosters future research. Resources have been made publicly available at https://github.com/zchuz/CoT-Reasoning-Survey

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 31 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data

    physics.chem-ph 2026-07 conditional novelty 6.0 of 10

    A two-stage AI pipeline — spectral hypothesis generation followed by mass-constrained molecular refinement — reconstructs organic structures from multimodal spectra, with 93.8% top-1 accuracy on simulated QM9 data and...

  2. LLM-based Question-Answer Framework for Sensor-driven HVAC System Interaction

    cs.AI 2025-07 conditional novelty 6.0 of 10

    JARVIS, an LLM-based HVAC question-answering framework with an Expert-LLM, a parameterized SQL builder, and bottom-up planning, outperforms a text-to-SQL baseline and its own ablations on a small expert-curated dataset.

  3. Intelligent Channel Allocation for IEEE 802.11be Multi-Link Operation: When MAB Meets LLM

    cs.NI 2025-06 conditional novelty 6.0 of 10

    BAI-MCTS and an LLM-initialized variant solve the WiFi 7 channel allocation problem as a multi-armed bandit, converging faster than prior bandit-MCTS baselines.

  4. Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test

    cs.AI 2025-05 conditional novelty 6.0 of 10

    With chain-of-thought prompting and text inputs, GPT-4o, Gemini-1.5 Pro, and Claude-3.5 Sonnet reach or exceed human-level set-shifting on the WCST, but not with visual inputs or direct answers.

  5. Skip-Thinking: Chunk-wise Chain-of-Thought Distillation Enable Smaller Language Models to Reason Better and Faster

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Chunk-wise training with loss-guided chunking and skip-thinking training improves small-model reasoning accuracy and speed over standard chain-of-thought distillation.

  6. DeepRec: Towards a Deep Dive Into the Item Space with Large Language Model Based Recommendation

    cs.IR 2025-05 conditional novelty 6.0 of 10

    An LLM trained by reinforcement learning to interact over multiple turns with a preference-aware recommender model outperforms both traditional and LLM-based baselines on sequential recommendation benchmarks.

  7. Systematic Evaluation of Machine-Generated Reasoning and PHQ-9 Labeling for Depression Detection Using Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A subtask decomposition of depression detection shows LLMs are biased by explicit depression keywords, and DPO fine-tuning on quality-filtered machine-generated rationales improves joint PHQ-9 labeling on the hardest samples.

  8. OPA-Pack: Object-Property-Aware Robotic Bin Packing

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A property-aware bin packing framework that uses a vision-language model to label object properties and a deep Q-network to separate incompatible items and reduce pressure on fragile objects.

  9. CHAINSFORMER: Numerical Reasoning on Knowledge Graphs from a Chain Perspective

    cs.AI 2025-04 conditional novelty 6.0 of 10

    ChainsFormer converts knowledge graph numerical reasoning into a chain-encoding problem, using hyperbolic filtering and attention weighting to improve prediction accuracy over prior graph-based methods.

  10. GCoT: Chain-of-Thought Prompt Learning for Graphs

    cs.CL 2025-02 conditional novelty 6.0 of 10

    GCoT improves few-shot graph classification by iteratively generating node-specific prompts from intermediate encoder states, mimicking chain-of-thought reasoning for text-free graphs.

  11. SensorChat: Answering Qualitative and Quantitative Questions during Long-Term Multimodal Sensor Interactions

    cs.AI 2025-02 conditional novelty 6.0 of 10

    A three-stage pipeline with LLM decomposition, pretrained embedding retrieval, and LLM assembly outperforms prior sensor QA systems on long-duration, high-frequency data, with caveats on evaluation leakage.

  12. MASTER: A Multi-Agent System with LLM Specialized MCTS

    cs.AI 2025-01 conditional novelty 6.0 of 10

    A multi-agent framework whose tree search is guided by LLM self-evaluation instead of simulations, reporting 76% on HotpotQA, 80% on WebShop, and 91% on MBPP.

  13. Navigating Chemical-Linguistic Sharing Space with Heterogeneous Molecular Encoding

    cs.CE 2024-12 conditional novelty 6.0 of 10

    A multi-view molecular encoder with fragment-based chain-of-thought improves text-to-molecule and molecule-to-text generation, supported by a new one-million-molecule conditional design dataset.

  14. Refining Answer Distributions for Improved Large Language Model Reasoning

    cs.CL 2024-12 conditional novelty 6.0 of 10

    RAD iteratively refines a distribution over answers by marginalizing over refinement samples, improving accuracy on six arithmetic benchmarks over self-consistency and hint-based prompting.

  15. IntentGPT: Few-shot Intent Discovery with Large Language Models

    cs.CL 2024-11 conditional novelty 6.0 of 10

    A training-free LLM prompting pipeline with semantic few-shot retrieval and feedback of discovered intents outperforms trained baselines on few-shot intent discovery benchmarks.

  16. Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning

    cs.CL 2025-07 conditional novelty 5.0 of 10

    CRFT selects critical internal representations via attention and saliency scores and fine-tunes only them, improving GSM8K accuracy over ReFT from 29.0% to 32.8% on LLaMA-2-7B.

  17. Attack Effect Model based Malicious Behavior Detection

    cs.CR 2025-06 conditional novelty 5.0 of 10

    FEAD derives security monitoring items from attack reports with an LLM, decomposes them across existing collectors, and applies locality-aware graph analysis, reporting an 8.23% higher F1-score than ThreaTrace with 5....

  18. Optimization Problem Solving Can Transition to Evolutionary Agentic Workflows

    math.OC 2025-05 conditional novelty 5.0 of 10

    An evolutionary loop of foundation-model agents could automate the full optimization pipeline, but the paper's evidence only covers two isolated components.

  19. CDW-CoT: Clustered Distance-Weighted Chain-of-Thoughts Reasoning

    cs.LG 2025-01 reject novelty 5.0 of 10

    CDW-CoT groups a reasoning dataset into clusters, learns a prompt distribution per cluster, and interpolates these distributions by embedding distance for each new query, reporting higher exact-match accuracy than thr...

  20. LLM-Powered User Simulator for Recommender System

    cs.IR 2024-12 reject novelty 5.0 of 10

    A simulator that combines LLM-generated like/dislike keywords, semantic similarity, and a sequential model to produce binary user feedback for recommender-system training.

  21. Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation

    cs.CL 2025-09 reject novelty 4.0 of 10

    Difficulty-aware compression of CoT traces plus SFT and DPO lets LLMs shorten reasoning on easy math problems, cutting tokens by up to 30% with mixed accuracy effects.

  22. CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks

    cs.CV 2025-07 reject novelty 4.0 of 10

    Adding ground-truth object localization, tracking, and relation annotations as context improves VideoLLM answers, but the models never generate these steps themselves in the experiments.

  23. DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A reward model that multiplies stepwise correctness and potential scores improves best-of-N verification accuracy for math reasoning.

  24. ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning

    cs.AI 2025-05 reject novelty 4.0 of 10

    ALAS combines role-specialized LLM agents, persistent state, and a local compensation protocol to produce disruption-tolerant schedules, reporting a 0.86% mean gap on a subset of Taillard instances and 19.09% on Demirkol-DMU.

  25. A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations

    cs.CR 2025-02 conditional novelty 4.0 of 10

    A survey of LVLM safety that adds a lifecycle taxonomy and new benchmark results showing Janus-Pro-7B has weaker safety than several open-source LVLMs.

  26. Empowering AIOps: Leveraging Large Language Models for IT Operations Management

    cs.SE 2025-01 conditional novelty 4.0 of 10

    Tool-using LLM agents can resolve many Kubernetes IT operations tasks; GPT-4o led advanced multi-tool tasks while Anthropic models led simple ones, and Mixtral 8x22B failed with hallucinations.

  27. Natural Language Fine-Tuning

    cs.CL 2024-12 reject novelty 4.0 of 10

    NLFT weights each token by how much its probability shifts under different natural language prompts and claims to beat supervised fine-tuning with 50 examples, but the paper's loss equation and reported gains are inte...

  28. OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous Driving

    cs.CV 2024-12 conditional novelty 4.0 of 10

    Adding a chain-of-thought reasoning step before predicting speed and curvature improves zero-shot trajectory planning of open multimodal LLMs on nuScenes, with code released.

  29. Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

    cs.CL 2024-12 conditional novelty 4.0 of 10

    LLM judges are often internally inconsistent across seed variations, with McDonald's omega reliability scores mostly below acceptable thresholds.

  30. Dspy-based Neural-Symbolic Pipeline to Enhance Spatial Reasoning in LLMs

    cs.AI 2024-11 conditional novelty 4.0 of 10

    A DSPy-orchestrated LLM plus Answer Set Programming pipeline reports 82% average accuracy on StepGame and 69% on SparQA, well above direct prompting baselines.

  31. Foundations of Large Language Models

    cs.CL 2025-01 unverdicted

    A textbook-style review of core LLM concepts, drawn from the authors' existing NLPBook, with no new experimental or theoretical results.

Pith tools