REVIEW 18 cited by
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large language models (LLMs) have recently shown impressive performance on tasks involving reasoning, leading to a lively debate on whether these models possess reasoning capabilities similar to humans. However, despite these successes, the depth of LLMs' reasoning abilities remains uncertain. This uncertainty partly stems from the predominant focus on task performance, measured through shallow accuracy metrics, rather than a thorough investigation of the models' reasoning behavior. This paper seeks to address this gap by providing a comprehensive review of studies that go beyond task accuracy, offering deeper insights into the models' reasoning processes. Furthermore, we survey prevalent methodologies to evaluate the reasoning behavior of LLMs, emphasizing current trends and efforts towards more nuanced reasoning analyses. Our review suggests that LLMs tend to rely on surface-level patterns and correlations in their training data, rather than on sophisticated reasoning abilities. Additionally, we identify the need for further research that delineates the key differences between human and LLM-based reasoning. Through this survey, we aim to shed light on the complex reasoning processes within LLMs.
Forward citations
Cited by 18 Pith papers
-
HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation
HALO learns a vision-based navigation reward from human preference rankings on egocentric video, and an IQL policy using it beats several baselines in 10-trial real-world tests.
-
Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies
Prompt structure and LoRA adaptation both strongly affect F1 on NLI4CT clinical NLI, but the claimed consistent +8 to 12 point LoRA gains and over 97% validity are not supported by the reported per-configuration results.
-
Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?
LLMs perform much worse on causal questions built from post-cutoff news articles, suggesting their apparent causal skill is mostly memorization, and a general-knowledge prompt method only partly closes the gap.
-
Propositional Logic for Probing Generalization in Neural Networks
Standard neural architectures generalize to unseen variable and operator combinations, but systematically fail when negation is applied to an operator that was hidden during training.
-
MINERVA: Evaluating Complex Video Reasoning
MINERVA provides 1,515 multi-step video QA questions with human reasoning traces; frontier models score far below humans and fail mainly on temporal localization and perception.
-
Reasoning-as-Logic-Units: Scaling Test-Time Reasoning in Large Language Models Through Logic Unit Alignment
RaLU aligns natural-language reasoning with program logic by decomposing generated code into control-flow units and self-correcting each one, reporting modest accuracy gains on math and code benchmarks.
-
LLM-based Discriminative Reasoning for Knowledge Graph Question Answering
READS reframes KGQA as three constrained selection tasks and reports SOTA results on WebQSP and CWQ with a 7B LLM.
-
Evaluating the Robustness of Analogical Reasoning in Large Language Models
GPT models solve original analogy tasks but fail many simple variants that humans handle easily, showing their analogy performance is not robust.
-
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
LLM reasoning benchmark scores vary substantially across repeated runs under the same model, strategy, and task, so single-run evaluation can misrank systems.
-
LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue
LLMs can imitate mental-state annotation in team dialogue but systematically err on spatial reasoning and prosodic cues, per a six-dialogue CReST pilot.
-
Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models
LLMs show measurable interlocutor awareness: they identify same-family models well and adapt behavior when told who they are talking to, which helps cooperation but raises alignment and safety risks.
-
SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models
SCOPE estimates a model's position bias with nonsense prompts, puts correct answers in disliked slots, and spreads similar distractors apart to cap lucky guessing.
-
From Thinking to Output: Chain-of-Thought and Text Generation Characteristics in Reasoning Language Models
Reasoning models differ in reflection frequency and in how closely their thinking and outputs resemble GPT-o1, but these differences are descriptive and not tied to answer accuracy.
-
Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation
Prompt rewrites of LeetCode problems cause large accuracy swings in nine LLMs, but invalid negation test cases and inconsistent tables make the headline numbers unreliable.
-
Generative to Agentic AI: Survey, Conceptualization, and Challenges
Agentic AI is characterized over Generative AI by iterative reasoning, environment interaction, memory, and tool use, with autonomy as the defining difference.
-
Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI
A conceptual analysis arguing that OpenAI's o3 solves ARC-AGI by brute-force search over predefined operations, so its high score is not evidence of AGI, and proposes a new definition and benchmark for intelligence.
-
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
A structured survey of LLM safety evaluation that proposes a why/what/where/how taxonomy and catalogs metrics, datasets, benchmarks, evaluators, and frameworks.
-
Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models
A narrative review arguing that medical LLM evaluations should examine reasoning behaviour, not only accuracy, and proposing two conceptual transparency frameworks.
Discussion (0). Continue with ORCID to comment.