Pith. sign in

REVIEW 2 cited by

GLoRE: Evaluating Logical Reasoning of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.09107 v2 pith:OQRT2IZ3 submitted 2023-10-13 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelslanguagereasoninglargelogicalgloreperformancedatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have shown significant general language understanding abilities. However, there has been a scarcity of attempts to assess the logical reasoning capacities of these LLMs, an essential facet of natural language understanding. To encourage further investigation in this area, we introduce GLoRE, a General Logical Reasoning Evaluation platform that not only consolidates diverse datasets but also standardizes them into a unified format suitable for evaluating large language models across zero-shot and few-shot scenarios. Our experimental results show that compared to the performance of humans and supervised fine-tuning models, the logical reasoning capabilities of large reasoning models, such as OpenAI's o1 mini, DeepSeek R1 and QwQ-32B, have seen remarkable improvements, with QwQ-32B achieving the highest benchmark performance to date. GLoRE is designed as a living project that continuously integrates new datasets and models, facilitating robust and comparative assessments of model performance in both commercial and Huggingface communities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Think Deep, Speak Once: Relit, A Recursive Latent Implicit Transformer Framework

    cs.AI 2026-08 reject novelty 5.0 of 10

    ReLIT reaches 98.6% on ProofWriter and 97.6% on RuleTaker by adding a recursive latent block to a frozen TinyLlama backbone.

  2. AI Benchmarks and Datasets for LLM Evaluation

    cs.DC 2024-12 conditional novelty 1.0 of 10

    A catalog of 41 existing AI benchmarks and datasets tagged with EU Trustworthy AI categories, with no new benchmarks or experimental results.

Pith tools