Pith. sign in

REVIEW 8 cited by

Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.17169 v3 pith:CL7EVHI4 submitted 2024-06-24 cs.CL cs.AI

classification cs.CLcs.AI
keywords reasoningllmslogicalmulti-logievalmulti-stepabilityevaluatinginference
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

As Large Language Models (LLMs) continue to exhibit remarkable performance in natural language understanding tasks, there is a crucial need to measure their ability for human-like multi-step logical reasoning. Existing logical reasoning evaluation benchmarks often focus primarily on simplistic single-step or multi-step reasoning with a limited set of inference rules. Furthermore, the lack of datasets for evaluating non-monotonic reasoning represents a crucial gap since it aligns more closely with human-like reasoning. To address these limitations, we propose Multi-LogiEval, a comprehensive evaluation dataset encompassing multi-step logical reasoning with various inference rules and depths. Multi-LogiEval covers three logic types--propositional, first-order, and non-monotonic--consisting of more than 30 inference rules and more than 60 of their combinations with various depths. Leveraging this dataset, we conduct evaluations on a range of LLMs including GPT-4, ChatGPT, Gemini-Pro, Yi, Orca, and Mistral, employing a zero-shot chain-of-thought. Experimental results show that there is a significant drop in the performance of LLMs as the reasoning steps/depth increases (average accuracy of ~68% at depth-1 to ~43% at depth-5). We further conduct a thorough investigation of reasoning chains generated by LLMs which reveals several important findings. We believe that Multi-LogiEval facilitates future research for evaluating and enhancing the logical reasoning ability of LLMs. Data is available at https://github.com/Mihir3009/Multi-LogiEval.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models

    cs.AI 2025-05 conditional novelty 7.0 of 10

    DeepMath-Creative is a new 179-problem benchmark measuring LLM mathematical creativity through constructive proof and counterexample tasks; the best model, O3 Mini, reaches only about 70% accuracy on basic undergradua...

  2. Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse

    cs.CL 2026-08 conditional novelty 6.0 of 10

    CUE-Bench provides 51,823 Chinese discourse instances annotated with a nine-way Affective Stance defined by explicit-implicit polarity, plus pragmatic intent and fine-grained emotion labels.

  3. The Effect of Perceived Race and Gender on Police Language Use: Experimental Evidence from VR Simulations

    cs.CY 2026-08 conditional novelty 6.0 of 10

    In randomized VR police encounters, officers' language to Black male characters was less deferential by about 0.1 points per exchange on a 0-10 scale, while White and biracial or multiracial female officers showed mor...

  4. Throttling Web Agents Using Reasoning Gates

    cs.AI 2025-09 conditional novelty 6.0 of 10

    Rebus-based reasoning gates, puzzles built from random word/domain clue sets, impose token costs on LM web agents that are up to 9.2x the generator's cost.

  5. LogiPlan: A Structured Benchmark for Logical Planning and Relational Reasoning in LLMs

    cs.AI 2025-06 conditional novelty 6.0 of 10

    LogiPlan introduces a three-task benchmark with controllable graph complexity, showing that although reasoning models excel at plan generation, all models degrade sharply on cycle detection and deep comparison questions.

  6. Computational Reasoning of Large Language Models

    cs.CL 2025-04 conditional novelty 6.0 of 10

    TMBench measures LLM computational reasoning by having models simulate m-tag systems step by step, and its pass rates correlate with AIME2024, MATH500, GPQA Diamond, and MMLU Pro scores across 12 leading models.

  7. ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning

    cs.AI 2025-05 reject novelty 4.0 of 10

    ALAS combines role-specialized LLM agents, persistent state, and a local compensation protocol to produce disruption-tolerant schedules, reporting a 0.86% mean gap on a subset of Taillard instances and 19.09% on Demirkol-DMU.

  8. Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey

    cs.AI 2025-04 conditional novelty 4.0 of 10

    The paper surveys existing work on LLM meta-thinking and argues that multi-agent reinforcement learning is a promising missing ingredient for building self-correcting language models.

Pith tools