Pith. sign in

REVIEW 18 cited by

Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.01869 v2 pith:JMRF4NZX submitted 2024-04-02 cs.CL cs.AI

classification cs.CLcs.AI
keywords reasoningllmsmodelsaccuracybehaviorsurveyabilitiesbeyond
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large language models (LLMs) have recently shown impressive performance on tasks involving reasoning, leading to a lively debate on whether these models possess reasoning capabilities similar to humans. However, despite these successes, the depth of LLMs' reasoning abilities remains uncertain. This uncertainty partly stems from the predominant focus on task performance, measured through shallow accuracy metrics, rather than a thorough investigation of the models' reasoning behavior. This paper seeks to address this gap by providing a comprehensive review of studies that go beyond task accuracy, offering deeper insights into the models' reasoning processes. Furthermore, we survey prevalent methodologies to evaluate the reasoning behavior of LLMs, emphasizing current trends and efforts towards more nuanced reasoning analyses. Our review suggests that LLMs tend to rely on surface-level patterns and correlations in their training data, rather than on sophisticated reasoning abilities. Additionally, we identify the need for further research that delineates the key differences between human and LLM-based reasoning. Through this survey, we aim to shed light on the complex reasoning processes within LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation

    cs.RO 2025-08 reject novelty 6.0 of 10

    HALO learns a vision-based navigation reward from human preference rankings on egocentric video, and an IQL policy using it beats several baselines in 10-trial real-world tests.

  2. Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies

    cs.CL 2025-07 reject novelty 6.0 of 10

    Prompt structure and LoRA adaptation both strongly affect F1 on NLI4CT clinical NLI, but the claimed consistent +8 to 12 point LoRA gains and over 97% validity are not supported by the reported per-configuration results.

  3. Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?

    cs.AI 2025-06 conditional novelty 6.0 of 10

    LLMs perform much worse on causal questions built from post-cutoff news articles, suggesting their apparent causal skill is mostly memorization, and a general-knowledge prompt method only partly closes the gap.

  4. Propositional Logic for Probing Generalization in Neural Networks

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Standard neural architectures generalize to unseen variable and operator combinations, but systematically fail when negation is applied to an operator that was hidden during training.

  5. MINERVA: Evaluating Complex Video Reasoning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    MINERVA provides 1,515 multi-step video QA questions with human reasoning traces; frontier models score far below humans and fail mainly on temporal localization and perception.

  6. Reasoning-as-Logic-Units: Scaling Test-Time Reasoning in Large Language Models Through Logic Unit Alignment

    cs.AI 2025-02 conditional novelty 6.0 of 10

    RaLU aligns natural-language reasoning with program logic by decomposing generated code into control-flow units and self-correcting each one, reporting modest accuracy gains on math and code benchmarks.

  7. LLM-based Discriminative Reasoning for Knowledge Graph Question Answering

    cs.CL 2024-12 conditional novelty 6.0 of 10

    READS reframes KGQA as three constrained selection tasks and reports SOTA results on WebQSP and CWQ with a 7B LLM.

  8. Evaluating the Robustness of Analogical Reasoning in Large Language Models

    cs.CL 2024-11 conditional novelty 6.0 of 10

    GPT models solve original analogy tasks but fail many simple variants that humans handle easily, showing their analogy performance is not robust.

  9. ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning

    cs.AI 2025-12 reject novelty 5.0 of 10

    LLM reasoning benchmark scores vary substantially across repeated runs under the same model, strategy, and task, so single-run evaluation can misrank systems.

  10. LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue

    cs.CL 2025-09 reject novelty 5.0 of 10

    LLMs can imitate mental-state annotation in team dialogue but systematically err on spatial reasoning and prosodic cues, per a six-dialogue CReST pilot.

  11. Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    LLMs show measurable interlocutor awareness: they identify same-family models well and adapt behavior when told who they are talking to, which helps cooperation but raises alignment and safety risks.

  12. SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models

    cs.CL 2025-07 reject novelty 4.0 of 10

    SCOPE estimates a model's position bias with nonsense prompts, puts correct answers in disliked slots, and spreads similar distractors apart to cap lucky guessing.

  13. From Thinking to Output: Chain-of-Thought and Text Generation Characteristics in Reasoning Language Models

    cs.CL 2025-06 reject novelty 4.0 of 10

    Reasoning models differ in reflection frequency and in how closely their thinking and outputs resemble GPT-o1, but these differences are descriptive and not tied to answer accuracy.

  14. Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation

    cs.CL 2025-06 reject novelty 4.0 of 10

    Prompt rewrites of LeetCode problems cause large accuracy swings in nine LLMs, but invalid negation test cases and inconsistent tables make the headline numbers unreliable.

  15. Generative to Agentic AI: Survey, Conceptualization, and Challenges

    cs.AI 2025-04 conditional novelty 4.0 of 10

    Agentic AI is characterized over Generative AI by iterative reasoning, environment interaction, memory, and tool use, with autonomy as the defining difference.

  16. Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI

    cs.AI 2025-01 conditional novelty 4.0 of 10

    A conceptual analysis arguing that OpenAI's o3 solves ARC-AGI by brute-force search over predefined operations, so its high score is not evidence of AGI, and proposes a new definition and benchmark for intelligence.

  17. The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

    cs.CL 2025-06 conditional novelty 3.0 of 10

    A structured survey of LLM safety evaluation that proposes a why/what/where/how taxonomy and catalogs metrics, datasets, benchmarks, evaluators, and frameworks.

  18. Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models

    cs.CL 2024-12 conditional novelty 3.0 of 10

    A narrative review arguing that medical LLM evaluations should examine reasoning behaviour, not only accuracy, and proposing two conceptual transparency frameworks.

Pith tools