Pith. sign in

REVIEW 1 cited by

Revisiting the Uniform Information Density Hypothesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.11635 v1 pith:AWLQSZZF submitted 2021-09-23 cs.CL

classification cs.CL
keywords languagehypothesisacceptabilityinformationdensitylinguisticuniformityacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The uniform information density (UID) hypothesis posits a preference among language users for utterances structured such that information is distributed uniformly across a signal. While its implications on language production have been well explored, the hypothesis potentially makes predictions about language comprehension and linguistic acceptability as well. Further, it is unclear how uniformity in a linguistic signal -- or lack thereof -- should be measured, and over which linguistic unit, e.g., the sentence or language level, this uniformity should hold. Here we investigate these facets of the UID hypothesis using reading time and acceptability data. While our reading time results are generally consistent with previous work, they are also consistent with a weakly super-linear effect of surprisal, which would be compatible with UID's predictions. For acceptability judgments, we find clearer evidence that non-uniformity in information density is predictive of lower acceptability. We then explore multiple operationalizations of UID, motivated by different interpretations of the original hypothesis, and analyze the scope over which the pressure towards uniformity is exerted. The explanatory power of a subset of the proposed operationalizations suggests that the strongest trend may be a regression towards a mean surprisal across the language, rather than the phrase, sentence, or document -- a finding that supports a typical interpretation of UID, namely that it is the byproduct of language users maximizing the use of a (hypothetical) communication channel.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PDD-RRG: Posterior Diagnostic Decision for Study-level Radiology Report Generation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A decision-level fusion method aggregates multiple draft radiology reports via Bayesian posterior scoring and validation-tuned thresholds, improving CheXbert F1 scores on MIMIC-CXR across three base models.

Pith tools