Pith. sign in

REVIEW 1 cited by

Simple Linguistic Inferences of Large Language Models (LLMs): Blind Spots and Blinds

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.14785 v2 pith:TVSJ2QJQ submitted 2023-05-24 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmsmodelscertainentailmentslanguageblindsembeddingentailment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We evaluate LLMs' language understanding capacities on simple inference tasks that most humans find trivial. Specifically, we target (i) grammatically-specified entailments, (ii) premises with evidential adverbs of uncertainty, and (iii) monotonicity entailments. We design evaluation sets for these tasks and conduct experiments in both zero-shot and chain-of-thought setups, and with multiple prompts and LLMs. The models exhibit moderate to low performance on these evaluation sets. Subsequent experiments show that embedding the premise in syntactic constructions that should preserve the entailment relations (presupposition triggers) or change them (non-factives), further confuses the models, causing them to either under-predict or over-predict certain entailment labels regardless of the true relation, and often disregarding the nature of the embedding context. Overall these results suggest that, despite LLMs' celebrated language understanding capacity, even the strongest models have blindspots with respect to certain types of entailments, and certain information-packaging structures act as ``blinds'' overshadowing the semantics of the embedded premise.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Position: Scaling LLM Agents Requires Asymptotic Analysis with LLM Primitives

    cs.CL 2025-02 conditional novelty 4.0 of 10

    The paper proposes that asymptotic analysis with LLM primitives, treating one forward pass as the cost unit, is the right framework for scaling multi-agent LLM systems.

Pith tools