Pith. sign in

REVIEW 7 cited by

JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.14851 v2 pith:75UWU7L3 submitted 2025-01-24 cs.CL cs.AIcs.LGcs.LO

classification cs.CLcs.AIcs.LGcs.LO
keywords reasoningdeductivejustlogicllmsmodelshumanknowledgeprior
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Logical reasoning is a critical component of Large Language Models (LLMs), and substantial research efforts in recent years have aimed to enhance their deductive reasoning capabilities. However, existing deductive reasoning benchmarks, which are crucial for evaluating and advancing LLMs, are inadequate due to their lack of task complexity, presence of prior knowledge as a confounder, and superficial error analysis. To address these deficiencies, we introduce JustLogic, a synthetically generated deductive reasoning benchmark designed for rigorous evaluation of LLMs. JustLogic is (i) highly complex, capable of generating a diverse range of linguistic patterns, vocabulary, and argument structures; (ii) prior knowledge independent, eliminating the advantage of models possessing prior knowledge and ensuring that only deductive reasoning is used to answer questions; and (iii) capable of in-depth error analysis on the heterogeneous effects of reasoning depth and argument form on model accuracy. Our experimental results on JustLogic reveal that (i) state-of-the-art (SOTA) reasoning LLMs perform on par or better than the human average but significantly worse than the human ceiling, and (ii) SOTA non-reasoning models still underperform the human average. All code and data are available at https://github.com/michaelchen-lab/JustLogic

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Abductive Corroboration of Probabilistic AI Models for Forensic Synthetic Media Detection

    cs.CR 2026-07 conditional novelty 6.0 of 10

    Multi-detector corroboration reduces FP/TP from ~0.22 to 0.02 (two models) or 0 (three models) while first measuring OpenAI SynthID production rollout and detector complementarity.

  2. ReportLogic: Evaluating Logical Quality in Deep Research Reports

    cs.CL 2026-01 conditional novelty 6.0 of 10

    An auditability-based benchmark with three logic layers and eight dimensions shows a distilled judge agrees with human experts ~74-75% versus ~62-74% for frontier LLM judges.

  3. Deductive Logic in Language Models: Horizontal vs Vertical Reasoning

    cs.AI 2025-10 conditional novelty 6.0 of 10

    A 2-layer, single-head attention-only transformer learns to perform multi-step logical deduction through induction-head circuits for rule completion, chaining, and final decision.

  4. Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL

    cs.AI 2025-09 conditional novelty 6.0 of 10

    DRER rewards CoT trajectories that increase the model's likelihood of the correct answer, plus a length penalty, and the new LogicTree benchmark reportedly lifts a 7B model's average accuracy from 0.13 to 0.60.

  5. PuzzleClone: A DSL-Powered Framework for Synthesizing Verifiable Data

    cs.AI 2025-08 conditional novelty 6.0 of 10

    A DSL plus SMT solver generates and validates 83,657 logic puzzles, and fine-tuning on them improves a 7B model's scores on several reasoning benchmarks.

  6. Logical Reasoning with Outcome Reward Models for Test-Time Scaling

    cs.CL 2025-08 conditional novelty 5.0 of 10

    Outcome reward models trained on multi-sample chain-of-thought plus deliberately flawed 'echo' rationales improve Best-of-N test-time verification for deductive reasoning.

  7. Improving Large Language Models with Concept-Aware Fine-Tuning

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Adding lightweight multi-token auxiliary heads with a weighted future-token loss improves supervised fine-tuning of Llama-3-8B-Instruct across five diverse tasks.

Pith tools