Pith. sign in

REVIEW 2 cited by

Navigating Semantic Relations: Challenges for Language Models in Abstract Common-Sense Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.14086 v1 pith:HPFGBO55 submitted 2025-02-19 cs.CL cs.AI

classification cs.CLcs.AI
keywords promptingreasoningrelationsabstractcommon-sensemodelsllmsmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have achieved remarkable performance in generating human-like text and solving reasoning tasks of moderate complexity, such as question-answering and mathematical problem-solving. However, their capabilities in tasks requiring deeper cognitive skills, such as common-sense understanding and abstract reasoning, remain under-explored. In this paper, we systematically evaluate abstract common-sense reasoning in LLMs using the ConceptNet knowledge graph. We propose two prompting approaches: instruct prompting, where models predict plausible semantic relationships based on provided definitions, and few-shot prompting, where models identify relations using examples as guidance. Our experiments with the gpt-4o-mini model show that in instruct prompting, consistent performance is obtained when ranking multiple relations but with substantial decline when the model is restricted to predicting only one relation. In few-shot prompting, the model's accuracy improves significantly when selecting from five relations rather than the full set, although with notable bias toward certain relations. These results suggest significant gaps still, even in commercially used LLMs' abstract common-sense reasoning abilities, compared to human-level understanding. However, the findings also highlight the promise of careful prompt engineering, based on selective retrieval, for obtaining better performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning

    q-bio.NC 2025-08 unverdicted novelty 6.0 of 10

    Only the largest tested LLMs (about 70 billion parameters) match human accuracy on an abstract reasoning task, and the internal geometry of their best layers correlates moderately with human frontal EEG activity.

  2. Relational Schemata in BERT Are Inducible, Not Emergent: A Study of Performance vs. Competence in Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Relational schemata in BERT are not emergent from pretraining: high classification accuracy coexists with unstructured embeddings, and only fine-tuning induces clustering by relation type.

Pith tools