Pith. sign in

REVIEW 13 cited by

FOLIO: Natural Language Reasoning with First-Order Logic

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.00840 v3 pith:QCBSLD4S submitted 2022-09-02 cs.CL

classification cs.CL
keywords languagefolioreasoningmodelsnaturalnl-folannotationscomplex
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have achieved remarkable performance on a variety of natural language understanding tasks. However, existing benchmarks are inadequate in measuring the complex logical reasoning capabilities of a model. We present FOLIO, a human-annotated, logically complex and diverse dataset for reasoning in natural language (NL), equipped with first-order logic (FOL) annotations. FOLIO consists of 1,430 examples (unique conclusions), each paired with one of 487 sets of premises used to deductively reason for the validity of each conclusion. The logical correctness of the premises and conclusions is ensured by their FOL annotations, which are automatically verified by an FOL inference engine. In addition to the main NL reasoning task, NL-FOL pairs in FOLIO constitute a new NL-FOL translation dataset. Our experiments on FOLIO systematically evaluate the FOL reasoning ability of supervised fine-tuning on medium-sized language models. For both NL reasoning and NL-FOL translation, we benchmark multiple state-of-the-art language models. Our results show that a subset of FOLIO presents a challenge for one of the most capable {Large Language Model (LLM)} publicly available, GPT-4.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Steering LLM Thinking with Budget Guidance

    cs.CL 2025-06 conditional novelty 7.0 of 10

    A test-time budget guidance module softly biases token generation to keep LLM reasoning traces within a target token budget, improving accuracy under tight budgets and cutting tokens versus hard cutoffs.

  2. Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles

    cs.CL 2025-05 conditional novelty 7.0 of 10

    Enigmata's synthetic puzzles with verifiable rewards lift a 32B model to 32.8% on ARC-AGI, above o3-mini-high and o1, and give small apparent gains on math and STEM when added to Seed1.5-Thinking.

  3. Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL

    cs.AI 2025-09 conditional novelty 6.0 of 10

    DRER rewards CoT trajectories that increase the model's likelihood of the correct answer, plus a length penalty, and the new LogicTree benchmark reportedly lifts a 7B model's average accuracy from 0.13 to 0.60.

  4. Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    The monotonicity of token probabilities during initial decoding predicts chain-of-thought gains, enabling dynamic selection between CoT and direct answers.

  5. MIH-TCCT: Mitigating Inconsistent Hallucinations in LLMs via Event-Driven Text-Code Cyclic Training

    cs.AI 2025-02 conditional novelty 6.0 of 10

    MIH-TCCT reduces inconsistent hallucinations by cyclically training LLMs to translate event-based text into structured code and back, without task-specific fine-tuning.

  6. DAIS: Dependency-Aware Intermediate QA Supervision for Complex Reasoning

    cs.CL 2026-07 conditional novelty 5.5 of 10

    Dependency-conditioned intermediate QA supervision improves fine-tuned LLM final-answer accuracy over flat chain-of-thought and answer-only baselines on four reasoning benchmarks.

  7. Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation

    cs.AI 2026-07 conditional novelty 5.0 of 10

    Nested exception-chain eligibility breaks frontier LLMs in unstable ways; an SMT execution layer makes outcomes deterministic given authored rules.

  8. Mitigating Spurious Correlations in LLMs via Causality-Aware Post-Training

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Fine-tuning a 3B LLM on randomly symbolized reasoning questions reduces spurious-correlation failures and improves OOD accuracy on CLadder and PrOntoQA.

  9. KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation

    cs.CL 2025-05 conditional novelty 5.0 of 10

    KORGym introduces a 51-game, text and visual, multi-turn benchmark with a normalized scoring scheme, and uses it to compare 19 LLMs and 8 VLMs on six reasoning dimensions.

  10. Evaluating the Meta- and Object-Level Reasoning of Large Language Models for Question Answering

    cs.CL 2025-02 conditional novelty 5.0 of 10

    LLMs frequently generate rational plans for multi-step questions but often fail to produce answers, and a new Franklin dataset is especially hard for them.

  11. A Comparative Study of Neurosymbolic AI Approaches to Interpretable Logical Reasoning

    cs.AI 2025-08 unverdicted novelty 4.0 of 10

    A comparison of two neurosymbolic designs concludes that the hybrid design, pairing an LLM with a separate symbolic solver, is the more promising path to general logical reasoning.

  12. Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A survey that classifies AI agent evaluation benchmarks along environment and capability axes, and proposes five traits that distinguish agents from chatbots.

  13. Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression

    cs.LG 2025-05 conditional novelty 4.0 of 10

    ACBench tests compressed LLMs on agentic tasks and finds 4-bit quantization keeps tool use and workflow generation strong while hurting real-world application performance.

Pith tools