Pith. sign in

REVIEW 2 cited by

"Is the Pope Catholic?" Applying Chain-of-Thought Reasoning to Understanding Conversational Implicatures

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.13826 v1 pith:YO6PC4AC submitted 2023-05-23 cs.CL

classification cs.CL
keywords humanimplicaturesaveragechain-of-thoughtconversationalperformancereasoningalthough
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Conversational implicatures are pragmatic inferences that require listeners to deduce the intended meaning conveyed by a speaker from their explicit utterances. Although such inferential reasoning is fundamental to human communication, recent research indicates that large language models struggle to comprehend these implicatures as effectively as the average human. This paper demonstrates that by incorporating Grice's Four Maxims into the model through chain-of-thought prompting, we can significantly enhance its performance, surpassing even the average human performance on this task.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pragmatic Theories Enhance Understanding of Implied Meanings in LLMs

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Zero-shot prompts that summarize Gricean pragmatics or Relevance Theory improve LLM accuracy on PRAGMEGA implied-meaning questions by up to 9.6 percentage points; just naming the theory helps larger models.

  2. They want to pretend not to understand: The Limits of Current LLMs in Interpreting Implicit Content of Political Discourse

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Even the best tested model, GPT4o-mini, falls more than 20 accuracy points short of the estimated human ceiling on a multiple-choice test and is fully correct in only about one-fourth of open-ended explanations.

Pith tools