REVIEW 31 cited by
Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Reasoning, a fundamental cognitive process integral to human intelligence, has garnered substantial interest within artificial intelligence. Notably, recent studies have revealed that chain-of-thought prompting significantly enhances LLM's reasoning capabilities, which attracts widespread attention from both academics and industry. In this paper, we systematically investigate relevant research, summarizing advanced methods through a meticulous taxonomy that offers novel perspectives. Moreover, we delve into the current frontiers and delineate the challenges and future directions, thereby shedding light on future research. Furthermore, we engage in a discussion about open questions. We hope this paper serves as an introduction for beginners and fosters future research. Resources have been made publicly available at https://github.com/zchuz/CoT-Reasoning-Survey
Forward citations
Cited by 31 Pith papers
-
Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data
A two-stage AI pipeline — spectral hypothesis generation followed by mass-constrained molecular refinement — reconstructs organic structures from multimodal spectra, with 93.8% top-1 accuracy on simulated QM9 data and...
-
LLM-based Question-Answer Framework for Sensor-driven HVAC System Interaction
JARVIS, an LLM-based HVAC question-answering framework with an Expert-LLM, a parameterized SQL builder, and bottom-up planning, outperforms a text-to-SQL baseline and its own ablations on a small expert-curated dataset.
-
Intelligent Channel Allocation for IEEE 802.11be Multi-Link Operation: When MAB Meets LLM
BAI-MCTS and an LLM-initialized variant solve the WiFi 7 channel allocation problem as a multi-armed bandit, converging faster than prior bandit-MCTS baselines.
-
Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test
With chain-of-thought prompting and text inputs, GPT-4o, Gemini-1.5 Pro, and Claude-3.5 Sonnet reach or exceed human-level set-shifting on the WCST, but not with visual inputs or direct answers.
-
Skip-Thinking: Chunk-wise Chain-of-Thought Distillation Enable Smaller Language Models to Reason Better and Faster
Chunk-wise training with loss-guided chunking and skip-thinking training improves small-model reasoning accuracy and speed over standard chain-of-thought distillation.
-
DeepRec: Towards a Deep Dive Into the Item Space with Large Language Model Based Recommendation
An LLM trained by reinforcement learning to interact over multiple turns with a preference-aware recommender model outperforms both traditional and LLM-based baselines on sequential recommendation benchmarks.
-
Systematic Evaluation of Machine-Generated Reasoning and PHQ-9 Labeling for Depression Detection Using Large Language Models
A subtask decomposition of depression detection shows LLMs are biased by explicit depression keywords, and DPO fine-tuning on quality-filtered machine-generated rationales improves joint PHQ-9 labeling on the hardest samples.
-
OPA-Pack: Object-Property-Aware Robotic Bin Packing
A property-aware bin packing framework that uses a vision-language model to label object properties and a deep Q-network to separate incompatible items and reduce pressure on fragile objects.
-
CHAINSFORMER: Numerical Reasoning on Knowledge Graphs from a Chain Perspective
ChainsFormer converts knowledge graph numerical reasoning into a chain-encoding problem, using hyperbolic filtering and attention weighting to improve prediction accuracy over prior graph-based methods.
-
GCoT: Chain-of-Thought Prompt Learning for Graphs
GCoT improves few-shot graph classification by iteratively generating node-specific prompts from intermediate encoder states, mimicking chain-of-thought reasoning for text-free graphs.
-
SensorChat: Answering Qualitative and Quantitative Questions during Long-Term Multimodal Sensor Interactions
A three-stage pipeline with LLM decomposition, pretrained embedding retrieval, and LLM assembly outperforms prior sensor QA systems on long-duration, high-frequency data, with caveats on evaluation leakage.
-
MASTER: A Multi-Agent System with LLM Specialized MCTS
A multi-agent framework whose tree search is guided by LLM self-evaluation instead of simulations, reporting 76% on HotpotQA, 80% on WebShop, and 91% on MBPP.
-
Navigating Chemical-Linguistic Sharing Space with Heterogeneous Molecular Encoding
A multi-view molecular encoder with fragment-based chain-of-thought improves text-to-molecule and molecule-to-text generation, supported by a new one-million-molecule conditional design dataset.
-
Refining Answer Distributions for Improved Large Language Model Reasoning
RAD iteratively refines a distribution over answers by marginalizing over refinement samples, improving accuracy on six arithmetic benchmarks over self-consistency and hint-based prompting.
-
IntentGPT: Few-shot Intent Discovery with Large Language Models
A training-free LLM prompting pipeline with semantic few-shot retrieval and feedback of discovered intents outperforms trained baselines on few-shot intent discovery benchmarks.
-
Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning
CRFT selects critical internal representations via attention and saliency scores and fine-tunes only them, improving GSM8K accuracy over ReFT from 29.0% to 32.8% on LLaMA-2-7B.
-
Attack Effect Model based Malicious Behavior Detection
FEAD derives security monitoring items from attack reports with an LLM, decomposes them across existing collectors, and applies locality-aware graph analysis, reporting an 8.23% higher F1-score than ThreaTrace with 5....
-
Optimization Problem Solving Can Transition to Evolutionary Agentic Workflows
An evolutionary loop of foundation-model agents could automate the full optimization pipeline, but the paper's evidence only covers two isolated components.
-
CDW-CoT: Clustered Distance-Weighted Chain-of-Thoughts Reasoning
CDW-CoT groups a reasoning dataset into clusters, learns a prompt distribution per cluster, and interpolates these distributions by embedding distance for each new query, reporting higher exact-match accuracy than thr...
-
LLM-Powered User Simulator for Recommender System
A simulator that combines LLM-generated like/dislike keywords, semantic similarity, and a sequential model to produce binary user feedback for recommender-system training.
-
Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
Difficulty-aware compression of CoT traces plus SFT and DPO lets LLMs shorten reasoning on easy math problems, cutting tokens by up to 30% with mixed accuracy effects.
-
CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks
Adding ground-truth object localization, tracking, and relation annotations as context improves VideoLLM answers, but the models never generate these steps themselves in the experiments.
-
DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning
A reward model that multiplies stepwise correctness and potential scores improves best-of-N verification accuracy for math reasoning.
-
ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning
ALAS combines role-specialized LLM agents, persistent state, and a local compensation protocol to produce disruption-tolerant schedules, reporting a 0.86% mean gap on a subset of Taillard instances and 19.09% on Demirkol-DMU.
-
A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations
A survey of LVLM safety that adds a lifecycle taxonomy and new benchmark results showing Janus-Pro-7B has weaker safety than several open-source LVLMs.
-
Empowering AIOps: Leveraging Large Language Models for IT Operations Management
Tool-using LLM agents can resolve many Kubernetes IT operations tasks; GPT-4o led advanced multi-tool tasks while Anthropic models led simple ones, and Mixtral 8x22B failed with hallucinations.
-
Natural Language Fine-Tuning
NLFT weights each token by how much its probability shifts under different natural language prompts and claims to beat supervised fine-tuning with 50 examples, but the paper's loss equation and reported gains are inte...
-
OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous Driving
Adding a chain-of-thought reasoning step before predicting speed and curvature improves zero-shot trajectory planning of open multimodal LLMs on nuScenes, with code released.
-
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
LLM judges are often internally inconsistent across seed variations, with McDonald's omega reliability scores mostly below acceptable thresholds.
-
Dspy-based Neural-Symbolic Pipeline to Enhance Spatial Reasoning in LLMs
A DSPy-orchestrated LLM plus Answer Set Programming pipeline reports 82% average accuracy on StepGame and 69% on SparQA, well above direct prompting baselines.
-
Foundations of Large Language Models
A textbook-style review of core LLM concepts, drawn from the authors' existing NLPBook, with no new experimental or theoretical results.
Discussion (0). Continue with ORCID to comment.