Pith. sign in

REVIEW 1 cited by

"A good pun is its own reword": Can Large Language Models Understand Puns?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.13599 v2 pith:74S2EOQF submitted 2024-04-21 cs.CL

"A good pun is its own reword": Can Large Language Models Understand Puns?

classification cs.CL
keywords punsllmsmetricsunderstandingevaluationgenerationhumorlanguage
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Puns play a vital role in academic research due to their distinct structure and clear definition, which aid in the comprehensive analysis of linguistic humor. However, the understanding of puns in large language models (LLMs) has not been thoroughly examined, limiting their use in creative writing and humor creation. In this paper, we leverage three popular tasks, i.e., pun recognition, explanation and generation to systematically evaluate the capabilities of LLMs in pun understanding. In addition to adopting the automated evaluation metrics from prior research, we introduce new evaluation methods and metrics that are better suited to the in-context learning paradigm of LLMs. These new metrics offer a more rigorous assessment of an LLM's ability to understand puns and align more closely with human cognition than previous metrics. Our findings reveal the "lazy pun generation" pattern and identify the primary challenges LLMs encounter in understanding puns.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?

    cs.CL 2026-07 conditional novelty 6.5

    Δacc between low-frequency and novel xiehouyu is ~23.6% for Chinese frontier LLMs vs ~5.1% for English-centric models and ~2.9% for humans, while LLM-created xiehouyu rate below human creations.