REVIEW 2 cited by
Are LLMs Good Cryptic Crossword Solvers?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Cryptic crosswords are puzzles that rely not only on general knowledge but also on the solver's ability to manipulate language on different levels and deal with various types of wordplay. Previous research suggests that solving such puzzles is a challenge even for modern NLP models. However, the abilities of large language models (LLMs) have not yet been tested on this task. In this paper, we establish the benchmark results for three popular LLMs -- LLaMA2, Mistral, and ChatGPT -- showing that their performance on this task is still far from that of humans.
Forward citations
Cited by 2 Pith papers
-
A Reasoning-Based Approach to Cryptic Crossword Clue Solving
An open pipeline of answer generation, wordplay proposal, and Python proof verification sets a new state of the art on the Cryptonite cryptic crossword benchmark.
-
What Makes Cryptic Crosswords Challenging for LLMs?
LLMs solve cryptic crossword clues with at most 11.4% accuracy, far below human experts, and fail mainly at definition extraction and wordplay type identification.
Discussion (0). Continue with ORCID to comment.