Pith. sign in

REVIEW 2 cited by

How to Compute the Probability of a Word

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.14561 v2 pith:4OTQOL2J submitted 2024-06-20 cs.CL

classification cs.CL
keywords computinglanguageprobabilitiesprobabilitymodelsvalueswordaccurately
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Language models (LMs) estimate a probability distribution over strings in a natural language; these distributions are crucial for computing perplexity and surprisal in linguistics research. While we are usually concerned with measuring these values for words, most LMs operate over subwords. Despite seemingly straightforward, accurately computing probabilities over one unit given probabilities over the other requires care. Indeed, we show here that many recent linguistic studies have been incorrectly computing these values. This paper derives the correct methods for computing word probabilities, highlighting issues when relying on language models that use beginning-of-word (bow)-marking tokenisers, e.g., the GPT family. Empirically, we show that correcting the widespread bug in probability computations affects measured outcomes in sentence comprehension and lexical optimisation analyses.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Language Models over Tokens to Language Models over Characters

    cs.CL 2024-12 accept novelty 8.0 of 10

    A principled method computes and samples from the character-level distribution induced by any token-level language model, using a new covering enumeration with beam approximations.

  2. A Grounded Typology of Word Classes

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A new image-based measure of word meaning shows that across 30 languages, word classes form a consistent groundedness cline, with nouns most grounded and function words least, but still nonzero.

Pith tools