Pith. sign in

REVIEW 2 cited by

Language Models are Crossword Solvers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.09043 v3 pith:MWB57U3F submitted 2024-06-13 cs.CL cs.AI

classification cs.CLcs.AI
keywords crosswordlanguagedemonstratellmsmodelscrosswordssolvingtackle
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Crosswords are a form of word puzzle that require a solver to demonstrate a high degree of proficiency in natural language understanding, wordplay, reasoning, and world knowledge, along with adherence to character and length constraints. In this paper we tackle the challenge of solving crosswords with large language models (LLMs). We demonstrate that the current generation of language models shows significant competence at deciphering cryptic crossword clues and outperforms previously reported state-of-the-art (SoTA) results by a factor of 2-3 in relevant benchmarks. We also develop a search algorithm that builds off this performance to tackle the problem of solving full crossword grids with out-of-the-box LLMs for the very first time, achieving an accuracy of 93% on New York Times crossword puzzles. Additionally, we demonstrate that LLMs generalize well and are capable of supporting answers with sound rationale.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Reasoning-Based Approach to Cryptic Crossword Clue Solving

    cs.CL 2025-06 conditional novelty 6.0 of 10

    An open pipeline of answer generation, wordplay proposal, and Python proof verification sets a new state of the art on the Cryptonite cryptic crossword benchmark.

  2. What Makes Cryptic Crosswords Challenging for LLMs?

    cs.CL 2024-12 conditional novelty 6.0 of 10

    LLMs solve cryptic crossword clues with at most 11.4% accuracy, far below human experts, and fail mainly at definition extraction and wordplay type identification.

Pith tools