Pith. sign in

REVIEW 2 cited by

Are LLMs Good Cryptic Crossword Solvers?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.12094 v2 pith:LJYU464Y submitted 2024-03-15 cs.AI cs.CLcs.LG

classification cs.AIcs.CLcs.LG
keywords llmscrypticlanguagemodelspuzzlestaskabilitiesability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Cryptic crosswords are puzzles that rely not only on general knowledge but also on the solver's ability to manipulate language on different levels and deal with various types of wordplay. Previous research suggests that solving such puzzles is a challenge even for modern NLP models. However, the abilities of large language models (LLMs) have not yet been tested on this task. In this paper, we establish the benchmark results for three popular LLMs -- LLaMA2, Mistral, and ChatGPT -- showing that their performance on this task is still far from that of humans.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Reasoning-Based Approach to Cryptic Crossword Clue Solving

    cs.CL 2025-06 conditional novelty 6.0 of 10

    An open pipeline of answer generation, wordplay proposal, and Python proof verification sets a new state of the art on the Cryptonite cryptic crossword benchmark.

  2. What Makes Cryptic Crosswords Challenging for LLMs?

    cs.CL 2024-12 conditional novelty 6.0 of 10

    LLMs solve cryptic crossword clues with at most 11.4% accuracy, far below human experts, and fail mainly at definition extraction and wordplay type identification.

Pith tools