Pith. sign in

REVIEW 2 cited by

Large Language Models Lack Understanding of Character Composition of Words

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.11357 v3 pith:O55JTCS5 submitted 2024-05-18 cs.CL

classification cs.CL
keywords languagellmstaskswordscharactercompositionlargemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have demonstrated remarkable performances on a wide range of natural language tasks. Yet, LLMs' successes have been largely restricted to tasks concerning words, sentences, or documents, and it remains questionable how much they understand the minimal units of text, namely characters. In this paper, we examine contemporary LLMs regarding their ability to understand character composition of words, and show that most of them fail to reliably carry out even the simple tasks that can be handled by humans with perfection. We analyze their behaviors with comparison to token level performances, and discuss the potential directions for future research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ChaI-TeA: A Benchmark for Evaluating Autocompletion of Interactions with LLM-based Chatbots

    cs.CL 2024-12 conditional novelty 6.0 of 10

    This paper defines the task of autocompleting user turns in LLM chatbot conversations and releases ChaI-TeA, a benchmark with datasets and a saved-effort metric showing that better ranking of generated completions is ...

  2. Why Do Large Language Models (LLMs) Struggle to Count Letters?

    cs.CL 2024-12 conditional novelty 5.0 of 10

    LLMs' letter-counting errors are driven mainly by repeated letters in a word, not by how common the word is.

Pith tools