REVIEW 8 cited by
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recently developed large language models (LLMs) have been shown to perform remarkably well on a wide range of language understanding tasks. But, can they really "reason" over the natural language? This question has been receiving significant research attention and many reasoning skills such as commonsense, numerical, and qualitative have been studied. However, the crucial skill pertaining to 'logical reasoning' has remained underexplored. Existing work investigating this reasoning ability of LLMs has focused only on a couple of inference rules (such as modus ponens and modus tollens) of propositional and first-order logic. Addressing the above limitation, we comprehensively evaluate the logical reasoning ability of LLMs on 25 different reasoning patterns spanning over propositional, first-order, and non-monotonic logics. To enable systematic evaluation, we introduce LogicBench, a natural language question-answering dataset focusing on the use of a single inference rule. We conduct detailed analysis with a range of LLMs such as GPT-4, ChatGPT, Gemini, Llama-2, and Mistral using chain-of-thought prompting. Experimental results show that existing LLMs do not fare well on LogicBench; especially, they struggle with instances involving complex reasoning and negations. Furthermore, they sometimes overlook contextual information necessary for reasoning to arrive at the correct conclusion. We believe that our work and findings facilitate future research for evaluating and enhancing the logical reasoning ability of LLMs. Data and code are available at https://github.com/Mihir3009/LogicBench.
Forward citations
Cited by 8 Pith papers
-
Author-in-the-Loop Response Generation and Evaluation: Integrating Author Expertise and Intent in Responses to Peer Review
Introduces Re3Align dataset of review-response-revision triplets, REspGen author-in-the-loop generation framework, and REspEval multi-metric suite for controllable peer-review response generation.
-
Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models
A three-statistic Borda consensus over sparse-autoencoder features produces interpretable activation steering, but usable quality-preserving shifts are rare and highly localized.
-
PuzzleClone: A DSL-Powered Framework for Synthesizing Verifiable Data
A DSL plus SMT solver generates and validates 83,657 logic puzzles, and fine-tuning on them improves a 7B model's scores on several reasoning benchmarks.
-
HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation
HALO learns a vision-based navigation reward from human preference rankings on egocentric video, and an IQL policy using it beats several baselines in 10-trial real-world tests.
-
UniCode: Augmenting Evaluation for Code Reasoning
UniCode's LLM-generated coding benchmark drops top-model pass@1 to 70.3% and indicates current LLMs rely on memorized seed logic instead of generalizing to new algorithmic problems.
-
Boosting Self-Efficacy and Performance of Large Language Models via Verbal Efficacy Stimulations
Emotionally styled verbal prompts (encouraging, provocative, critical) modestly improve zero-shot LLM accuracy on many tasks, with the best style varying by model and task zone.
-
Evaluation of LLMs for mathematical problem solving
A three-model, three-dataset LLM math evaluation using a multi-dimensional reasoning rubric, undermined by contradictory accuracy tables.
-
BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
A purported impossibility theorem for LLM services reduces to the paper's own assumption that reasoning and authenticity necessarily consume extra inference budget.
Discussion (0). Continue with ORCID to comment.