Pith. sign in

REVIEW 5 major objections 4 minor 1 cited by

Entropy and type-token ratio in gigaword corpora

T0 review · 5 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read For large natural-language texts, Shannon word entropy and type-token ratio are linked by a single relation.

desk verdict Useful empirical relation between entropy and TTR, but the analytic validation is weaker than claimed because the fits use free prefactors that deviate from the predicted values by 1–2 orders of magnitude, and Eq. (16) has a sign error. read the letter →

arxiv 2411.10227 v3 pith:RB7KGYJK submitted 2024-11-15 cs.CL cs.IRphysics.soc-ph

classification cs.CLcs.IRphysics.soc-ph
keywords Shannonentropytype-tokenratioZipflawHeapslexicaldiversitygigawordcorpusstatisticallanguageuniversals
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that the two standard measures of lexical diversity, Shannon word entropy and type-token ratio, are not independent for sufficiently long texts but are tied by an explicit asymptotic relation. The relation is derived from Zipf's law for word frequencies and Heaps' law for vocabulary growth, and it is checked against six corpora of more than a billion tokens each, spanning three languages and three registers. A careful reader would care because the result turns two superficially different metrics into two projections of the same underlying statistical regularity of language, with consequences for how entropy and diversity are estimated in large text data.

What carries the argument

The argument is carried by two empirical laws of quantitative linguistics: Zipf's rank-frequency law $f = k/r^a$ with $a\simeq 1$ over the kernel vocabulary, and Heaps' law $V = \alpha L^\beta$ for vocabulary growth. The asymptotic expansion of the harmonic sum gives $H \sim \frac{1}{2}\ln V + \ln \ln V$, and substituting Heaps law yields both the entropy-vs-length formula and, through $TTR = V/L$, the final $H$–$TTR$ relation. The load-bearing numerical fact is that ranks in the second Zipf regime contribute roughly 20% of the entropy, so treating the whole distribution as a single $a=1$ power law is a good approximation.

What would settle it

Compute $H$ and $TTR$ on a corpus with a Zipf exponent in the kernel regime clearly different from one (for example a highly isolating language) and test whether the data obey $H = \frac{\beta}{2(1-\beta)}\ln(1/TTR) + \ln\ln(1/TTR)$ with no fitted multiplicative parameter; a systematic failure of this parameter-free prediction would falsify the universality of the relation.

Watch

Extended reading notes

Core claim

The central claim is that for $TTR \ll 1$, the Shannon word entropy of a natural-language text follows $H \sim \frac{\beta}{2(1-\beta)} \ln(1/TTR) + \ln \ln(1/TTR)$, where $\beta$ is the Heaps exponent. This expression results from combining the large-vocabulary entropy of a Zipf distribution with exponent one, $H \sim \frac{1}{2}\ln V + \ln \ln V$, with the Heaps-law observation that $TTR \sim L^{\beta-1}$. The authors verify the functional form on six gigaword corpora, showing that the same curve describes books, web media, and tweets in English, Spanish, and Turkish, with per-corpus fitting parameters absorbing the residual differences. They also show that the second, steeper regime of the Zipf distribution contributes only about 20% of the total entropy, which justifies the single-exponent derivation.

Load-bearing premise

Everything rests on the assumption that a single Zipf law with exponent exactly one captures enough of the word-frequency distribution that the higher-rank tail contributes only about a fifth of the entropy, and that this tail contribution stays roughly constant; if the tail dominates or varies with text length, the derived prefactors and the whole asymptotic relation change.

Editorial extensions

If this is right

  • In the asymptotic regime $TTR \ll 1$, the entropy and the type-token ratio carry the same information: the Heaps exponent and one of the two quantities determine the other.
  • Very large corpora can be cross-validated: estimates of $H$ from word-frequency tables should agree with TTR-based estimates through Eq. (15) once corpus-specific constants are calibrated.
  • Any corpus in any language that obeys Zipf and Heaps laws is predicted to fall on the same universal $H$–$TTR$ curve, with only the parameters $\beta$ and the additive constant varying.
  • The over-all (length-dependent) definition of TTR is essential; a mean-segmental TTR would produce a positive instead of negative correlation with entropy, as the paper notes.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: The derivation uses only Zipf and Heaps laws, so the same $H$–$TTR$ relation should appear in other heavy-tailed systems with type counts and sublinear type growth, such as ecological abundance data or genomic k-mer counts; this is a testable cross-domain prediction the authors list only as future work.
  • Editorial inference: The fitted slopes are well below the asymptotic prefactor, implying that real corpora of $10^9$ tokens have not yet reached the clean limit; practical entropy estimates from TTR should therefore be calibrated per corpus rather than used with the theoretical prefactor.
  • Editorial inference: The relation can serve as a distributional diagnostic for text generators: if a language model's long outputs trace a different $H$–$TTR$ curve than human corpora, that signals a systematic over- or under-repetition that perplexity alone would miss.
  • Editorial inference: A natural next test is to vary tokenization (subword vs. word) — the relation should still hold with modified exponents, since subword units also obey Zipf-like and Heaps-like laws in many languages.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. The paper examines the relation between Shannon word entropy H and the type-token ratio TTR in six gigaword corpora (English, Spanish, Turkish; books, web, and Twitter). The authors report a negative empirical correlation between H and TTR and derive an asymptotic analytical expression, H ~ β/(2(1−β)) ln(1/TTR) + ln ln(1/TTR) (Eq. 15), from Zipf's law with exponent a=1 and Heaps' law V=αL^β. They validate this expression by fitting semi-empirical forms with free prefactors and offsets (Eqs. 13 and 16) and report high correlation coefficients for most corpora. The paper argues that, for sufficiently large texts, lexical diversity measures such as TTR and entropy are functionally dependent, contrary to naive expectation.

Significance. If the analytical relation were quantitatively confirmed, the result would be of interest to quantitative linguistics and NLP: it would connect two widely used diversity metrics and imply that they are not independent for large texts. The paper's strengths include the use of very large corpora (over 10^9 tokens each) across three morphologically distinct languages and multiple registers, a reproducible analysis with publicly available code, a clear derivation from Zipf and Heaps laws, and a robust empirical demonstration of a negative H–TTR correlation. The main weakness is that the quantitative validation relies on free prefactors that deviate strongly from the predicted values, and the printed validation equation contains a sign error; as a result, the paper does not currently establish the claimed agreement with the asymptotic expression.

major comments (5)
  1. [Section V, Eq. (16)] The printed Eq. (16) contains a sign error: β/(2(β−1)) equals −β/(2(1−β)), so the leading term in Eq. (16) has the opposite sign to Eq. (15). Since the fitted p3 values in Table V are positive, the fitted function would decrease as ln(1/TTR) increases, contradicting the data in Fig. 6 unless the equation is a typo. This makes the validation as printed internally inconsistent and needs correction before the claimed agreement can be assessed.
  2. [Section III.C, Table III and Section V, Table V] The fits use free prefactors p1 and p3 that multiply the entire predicted expression. The predicted coefficients in Eqs. (12) and (15) are β/2 and β/(2(1−β)), respectively, corresponding to p1=1 and p3=1. Table III reports p1=0.03–0.27 and Table V reports p3=0.04–2.48, with most values one to two orders of magnitude from unity. With free prefactors, the fits test only the functional form and not the predicted coefficient, so the statement in Section V that the data agree with Eq. (15) is not supported by the evidence presented.
  3. [Section III.B and Appendix B, Table VI] The assumption that the second Zipf regime has a negligible contribution to the entropy is not supported by Table VI: R_H ranges from 0.14 to 0.24 and R'_H from 0.12 to 0.21, meaning the second regime contributes roughly 12–24% of the total entropy. This is not negligible, and if this fraction varies across corpora it changes the prefactor of the asymptotic relation derived from a pure a=1 Zipf law. The fitted p1 values far from unity are consistent with this concern and suggest that the two-regime structure materially affects the prefactor.
  4. [Section III.C and Appendix C] The range of L over which the fits are performed is selected post hoc using the condition 0.0025 < σ(H) < 0.025, with different ad hoc ranges for the two Turkish corpora (e.g., for TRCC100 a fixed vocabulary interval is chosen). Because the same data are used to select the fitting range and to compute the reported goodness-of-fit values, the high ρ² values in Tables III and V are not an out-of-sample confirmation. The authors should provide a pre-registered or otherwise justified range selection, or at least show that the fitted parameters are stable under reasonable variations of the range.
  5. [Table V, TwTR row] The TwTR corpus has ρ²=0.60 and p3=2.48, indicating a poor fit and a prefactor far from the predicted value. The paper acknowledges the poor statistics for TwTR but still concludes an 'overall good agreement' across the six corpora. With one of six corpora failing the quantitative test, the universality claim in the abstract and conclusion is overstated and should be qualified.
minor comments (4)
  1. [Eqs. (15) and (16)] The notation 'ln ln TTR^{-1}' is ambiguous; it should be written as ln(ln(1/TTR)) to avoid the possible misreading (ln ln TTR)^{-1}.
  2. [Section V, Fig. 5] The text says H_max corresponds to L=1.8×10^9 for all corpora, but SPGC and TRCC100 have total lengths larger than this; please clarify how H_max is computed for those corpora.
  3. [Section III.C, Table III] The fitted p1 values are reported without discussion; given that they deviate strongly from the predicted value of 1, the authors should either explain the deviation or temper the claim of 'excellent agreement'.
  4. [Introduction] The sentence 'This observation would suggest that type-token ratio and entropy are distinct and uncorrelated diversity measures' could be misread as a claim that they are uncorrelated in general; consider rephrasing to 'potentially uncorrelated'.

Circularity Check

2 steps flagged · score 6.0 of 10

The empirical validation of the central H–TTR relation is a semi-empirical fit with free prefactors p1/p3; the predicted coefficient of Eq. (15) is not independently tested.

  1. fitted input called prediction [Section III.C, Eq. (13), Table III]
    "where p1 and p2 are fitting parameters, independent of L, for the six datasets, in order to align the asymptotic approximation to the empirical data. In other words, Eq. (13) is a semi-empirical approach that retains the same functional behaviour of H(L) for all corpora with p1 and p2 being specific for each dataset."

    Eq. (12) predicts H ~ (beta/2) ln L + ln ln L with unit prefactor on the bracket. Fitting H_fit = p1[bracket] + p2 lets the free parameter p1 absorb any mismatch of the predicted coefficient, so the reported rho^2 cannot confirm the quantitative content of Eq. (12). Table III gives p1 = 0.03-0.27, i.e. far from the predicted value 1 for five of six corpora. The 'excellent agreement' is thus an agreement of a rescaled logarithmic template, not a confirmation of the derived coefficient; the fit itself supplies the match.

  2. fitted input called prediction [Section V, Eq. (16), Table V]
    "In order to validate Eq. (15), we fit Hfit(TTR) = p3[beta/(2(beta-1)) ln TTR^{-1} + ln ln TTR^{-1}] + p4, p3 and p4 being fitting parameters."

    Eq. (15) predicts H ~ beta/(2(1-beta)) ln(1/TTR) + ln ln(1/TTR) with no free prefactor. Validation via Eq. (16) multiplies the entire predicted logarithmic combination by a free parameter p3 and adds p4, so the fit can only test the functional shape, not the predicted coefficient beta/(2(1-beta)). Table V reports p3 = 0.04, 0.61, 0.24, 0.49, 0.43, and 2.48, not close to 1. Moreover, beta/(2(beta-1)) = -beta/(2(1-beta)), opposite in sign to Eq. (15); with the positive p3 values reported, the fitted leading term has the wrong sign relative to the data. The claimed agreement with Eq. (15) is therefore produced by the free fitting parameters rather than by an independent confirmation of the theoretical expression.

full rationale

The mathematical derivation of Eq. (15) from Zipf's law (a=1), Heaps' law, and the asymptotic entropy expression is not circular: it is explicit algebra from stated assumptions, supplemented by an empirical check that the high-rank tail contributes roughly 20% of the entropy (Appendix B). There is no load-bearing self-citation chain; the paper's references to prior work are ordinary citations. However, the central empirical validation is semi-empirical by the paper's own admission: the predicted functional forms are fitted to the same corpora with free prefactors p1 and p3 and offsets p2 and p4, while the Heaps exponents beta are also fitted per corpus. Because the fitted prefactors deviate from unity by factors up to roughly 30, the reported 'agreement' substantially reflects the flexibility of the fitting function rather than an independent test of the predicted coefficients. The relation in Eq. (15) is not empty, but its claimed empirical confirmation reduces in part to a fit with the predicted shape, which is why the circularity score is 6 rather than 0.

Assumptions & free parameters 6 free parameters · 6 assumptions · 0 invented entities

The paper introduces no new entities. Its free-parameter burden is substantial: Heaps exponents and two additional prefactors plus offsets are fitted per corpus, so the displayed agreement is semi-empirical. The main domain assumptions are the dominance of the a=1 Zipf regime and the applicability of Heaps' law to random fragments.

free parameters (6)
  • beta (Heaps exponent per corpus) = 0.53-0.61 across corpora (Fig. 1)
    Fitted from V versus L data for each corpus and used directly in Eqs. (12) and (15).
  • p1 (prefactor in Eq. 13) = SPGC 0.04, SPA 0.06, TRCC100 0.03, TwEN 0.05, TwES 0.05, TwTR 0.27
    Per-corpus fitted prefactor in the entropy-vs-length fit; values far from the theoretical value 1.
  • p2 (offset in Eq. 13) = SPGC 6.83, SPA 6.85, TRCC100 9.32, TwEN 6.91, TwES 6.91, TwTR 7.69
    Per-corpus fitted offset in the entropy-vs-length fit.
  • p3 (prefactor in Eq. 16) = SPGC 0.04, SPA 0.61, TRCC100 0.24, TwEN 0.49, TwES 0.43, TwTR 2.48
    Per-corpus fitted prefactor in the entropy-vs-TTR fit.
  • p4 (offset in Eq. 16) = SPGC 6.93, SPA 5.61, TRCC100 8.94, TwEN 5.95, TwES 6.13, TwTR 3.03
    Per-corpus fitted offset in the entropy-vs-TTR fit.
  • Two-regime Zipf parameters (a1, a2, r_c) = a1 ~ 0.91-1.19, a2 ~ 1.79-1.90, r_c ~ 4.9e3-5.8e4 (Table VI)
    Fitted to each corpus to justify the a=1 approximation by showing the second regime contributes only about 20% of entropy.
assumptions (6)
  • domain assumption Word frequencies follow Zipf's law with exponent a=1 over the ranks that dominate the entropy.
    Used in Section III.B to derive Eq. (9); the two-regime correction is treated as a small contribution in Appendix B.
  • domain assumption Vocabulary growth follows Heaps' law V = alpha L^beta with a single beta for each corpus over the sampled range.
    Used in Section III.B and Section IV to express entropy and TTR as functions of L.
  • domain assumption Random fragments of a given length are representative samples of the corpus bag-of-words distribution.
    The entropy and vocabulary measurements in Appendix A average over 25 random fragments and characterize them by their mean.
  • standard math Euler-Maclaurin and Stieltjes-constant asymptotic expansions are valid at vocabularies of order 10^6.
    Used in Eqs. (6)-(9) to obtain H ~ 1/2 ln V + ln ln V.
  • domain assumption The second Zipf regime contributes a small enough fraction of entropy to be dropped from the asymptotic formula.
    Appendix B reports R_H ~ 0.14-0.24, which the authors interpret as negligible; the fitted prefactors p1 do not fully support this assumption.
  • domain assumption Plug-in entropy estimates are accurate enough, as checked by the NSB estimator.
    Section III.C and Appendix D compare plug-in and NSB estimates and report no significant deviations.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Entropy and type-token ratio in gigaword corpora." pith.science (2026). https://pith.science/paper/RB7KGYJK

@misc{pith2026241110227,
  author       = {Pith},
  title        = {Pith review of: Entropy and type-token ratio in gigaword corpora},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RB7KGYJK}},
  note         = {Machine review of arXiv:2411.10227}
}
read the original abstract

There are different ways of measuring diversity in complex systems. In particular, in language, lexical diversity is characterized in terms of the type-token ratio and the word entropy. We here investigate both diversity metrics in six massive linguistic datasets in English, Spanish, and Turkish, consisting of books, news articles, and tweets. These gigaword corpora correspond to languages with distinct morphological features and differ in registers and genres, thus constituting a varied testbed for a quantitative approach to lexical diversity. We unveil an empirical functional relation between entropy and type-token ratio of texts of a given corpus and language, which is a consequence of the statistical laws observed in natural language. Further, in the limit of large text lengths we find an analytical expression for this relation relying on both Zipf and Heaps laws that agrees with our empirical findings.

Figures

Figures reproduced from arXiv: 2411.10227 by the authors.

Figure 1
Figure 1. FIG. 1. Frequency [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2. Measured entropy [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. FIG. 3. Incremental partial entropy, [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: FIG. 4. Range selection for the fits of the SPGC in Fig. [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 6
Figure 6. Figure 6: FIG. 6 [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 5
Figure 5. Figure 5: FIG. 5. Deviations from maximum entropy, [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 7
Figure 7. Figure 7: FIG. 7. (a) Typical probability distribution of the measure [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]
Figure 8
Figure 8. Figure 8: FIG. 8. Range selection for the fits in Fig. [PITH_FULL_IMAGE:figures/full_fig_p011_8.png]
Figure 9
Figure 9. Figure 9: FIG. 9. Measured entropy [PITH_FULL_IMAGE:figures/full_fig_p012_9.png]
Figure 10
Figure 10. Figure 10: FIG. 10 [PITH_FULL_IMAGE:figures/full_fig_p012_10.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generative Exaggeration in LLM Social Agents: Consistency, Bias, and Toxicity

    cs.HC 2025-07 conditional novelty 5.0 of 10

    When LLMs are given more context about a real social media user, they become more ideologically consistent but also more extreme, toxic, and stereotyped than the user actually is.

Reference graph

Works this paper leans on

101 extracted references · 80 canonical work pages · cited by 1 Pith paper

  1. [1]

    Johnson, Studies in language behavior: A program of research, Psychological Monographs56, 1 (1944)

    W. Johnson, Studies in language behavior: A program of research, Psychological Monographs56, 1 (1944)

  2. [2]

    M. C. Templin,Certain Language Skills in Children: Their Development and Interrelationships, Vol. 26 (Uni- versity of Minnesota Press, 1957)

  3. [3]

    Hardie and T

    A. Hardie and T. McEnery, Statistics, inEncyclopedia of Language and Linguistics, edited by K. Brown (Elsevier, Amsterdam, 2006) 2nd ed

  4. [4]

    Kettunen, Can Type-Token Ratio be Used to Show Morphological Complexity of Languages?, Journal of Quantitative Linguistics21, 223 (2014)

    K. Kettunen, Can Type-Token Ratio be Used to Show Morphological Complexity of Languages?, Journal of Quantitative Linguistics21, 223 (2014)

  5. [5]

    Bermel, Corpora and quantitative data in Slavic lan- guages, Russian Linguistics39, 275 (2015)

    N. Bermel, Corpora and quantitative data in Slavic lan- guages, Russian Linguistics39, 275 (2015)

  6. [6]

    Litvinova, P

    T. Litvinova, P. Seredin, O. Litvinova, and O. Zagorovskaya, Differences in type-token ratio and part-of-speech frequencies in male and female Rus- sian written texts, inProceedings of the Workshop on Stylistic Variation, edited by J. Brooke, T. Solorio, and M. Koppel (Association for Computational Linguistics, Copenhagen, Denmark, 2017) pp. 69–73

  7. [7]

    C. W. Hess, K. M. Sefton, and R. G. Landry, Sample size and type-token ratios for oral language of preschool children, Journal of Speech, Language, and Hearing Re- search29, 129 (1986)

  8. [8]

    T. C. Manschreck, B. A. Maher, T. M. Hoover, and D. Ames, The type—token ratio in schizophrenic disor- ders: clinical and research value, Psychological Medicine 14, 151–157 (1984)

Show all 101 references
  1. [9]

    Jost, Entropy and diversity, Oikos113, 363 (2006)

    L. Jost, Entropy and diversity, Oikos113, 363 (2006)

  2. [10]

    Tuomisto, A consistent terminology for quantifying species diversity? yes, it does exist, Oecologia164, 853 (2010)

    H. Tuomisto, A consistent terminology for quantifying species diversity? yes, it does exist, Oecologia164, 853 (2010)

  3. [11]

    Mazzarisi, A

    O. Mazzarisi, A. de Azevedo-Lopes, J. J. Arenzon, and F. Corberi, Maximal Diversity and Zipf’s Law, Phys. Rev. Lett.127, 128301 (2021)

  4. [12]

    Bentz, D

    C. Bentz, D. Alikaniotis, M. Cysouw, and R. Ferrer-i Cancho, The Entropy of Words—Learnability and Ex- pressivity across More than 1000 Languages, Entropy 19, 275 (2017)

  5. [13]

    M. A. Montemurro and D. H. Zanette, Universal en- tropy of word ordering across linguistic families, PloS one6, e19875 (2011)

  6. [14]

    C. E. Shannon, A mathematical theory of communica- tion, Bell System Technical Journal27, 379 (1948)

  7. [15]

    T. M. Cover and J. A. Thomas,Elements of information theory(John Wiley & Sons, 1999)

  8. [16]

    K. Liu, R. Ye, L. Zhongzhu, and R. Ye, Entropy-based discrimination between translated Chinese and original Chinese using data mining techniques, PloS one17, e0265633 (2022)

  9. [17]

    Friedrich, M

    R. Friedrich, M. Luzzatto, and E. Ash, Entropy in legal language, inNLLP 2020 Natural Legal Language Pro- cessing Workshop 2020. Proceedings of the Natural Le- gal Language Processing Workshop 2020 co-located with the 26th ACM SIGKDD International Conference on Knowledge Disco...

  10. [18]

    Gerlach, F

    M. Gerlach, F. Font-Clos, and E. G. Altmann, Similar- ity of Symbol Frequency Distributions with Heavy Tails, Phys. Rev. X6, 021009 (2016)

  11. [19]

    Gerlach, H

    M. Gerlach, H. Shi, and L. A. N. Amaral, A universal information theoretic approach to the identification of stopwords, Nature Machine Intelligence1, 606 (2019)

  12. [20]

    Mohseni, C

    M. Mohseni, C. Redies, and V. Gast, Comparative Anal- ysis of Preference in Contemporary and Earlier Texts Using Entropy Measures, Entropy25, 486 (2023)

  13. [21]

    I. I. Eliazar and I. M. Sokolov, Diversity of Poissonian populations, Phys. Rev. E81, 011122 (2010)

  14. [22]

    J. R. Romero-Arias, G. Ram ´ ırez-Santiago, J. X. Velasco-Hern´ andez, L. Ohm, and M. Hern´ andez- Rosales, Model for breast cancer diversity and spatial heterogeneity, Phys. Rev. E98, 032401 (2018)

  15. [23]

    G. K. Zipf, The psychology of language, inEncyclopedia of psychology(Philosophical Library, 1946) pp. 332–341

  16. [24]

    Herdan,Type-token mathematics: A textbook of mathematical linguistics(Moulton & Co, 1960)

    G. Herdan,Type-token mathematics: A textbook of mathematical linguistics(Moulton & Co, 1960)

  17. [25]

    Heaps,Information retrieval: Computational and theoretical aspects(Academic Press, 1978)

    H. Heaps,Information retrieval: Computational and theoretical aspects(Academic Press, 1978)

  18. [26]

    E. G. Altmann and M. Gerlach, Statistical Laws in Linguistics, inCreativity and Universality in Language, edited by M. Degli Esposti, E. G. Altmann, and F. Pa- chet (Springer International Publishing, Cham, 2016) pp. 7–26

  19. [27]

    Stanisz, S

    T. Stanisz, S. Dro˙ zd˙ z, and J. Kwapie´ n, Complex sys- tems approach to natural language, Physics Reports 1053, 1 (2024)

  20. [28]

    Arnon, S

    I. Arnon, S. Kirby, J. A. Allen, C. Garrigue, E. L. Car- roll, and E. C. Garland, Whale song shows language-like statistical structure, Science387, 649 (2025)

  21. [29]

    Mitzenmacher, A brief history of generative mod- els for power law and lognormal distributions, Internet Mathematics1, 226 (2004)

    M. Mitzenmacher, A brief history of generative mod- els for power law and lognormal distributions, Internet Mathematics1, 226 (2004)

  22. [30]

    M. ´A. Serrano, A. Flammini, and F. Menczer, Modeling statistical properties of written text, PloS one4, e5372 (2009)

  23. [31]

    Gerlach and E

    M. Gerlach and E. G. Altmann, Stochastic Model for the Vocabulary Growth in Natural Languages, Phys. Rev. X3, 021006 (2013)

  24. [32]

    Malvern, B

    D. Malvern, B. Richards, N. Chipere, and P. Dur´ an, Lexical diversity and language development(Springer, 2004)

  25. [33]

    Shi and L

    Y. Shi and L. Lei, Lexical Richness and Text Length: An Entropy-based Perspective, Journal of Quantitative Linguistics29, 62 (2022)

  26. [34]

    Gregori-Signes and B

    C. Gregori-Signes and B. Clavel-Arroitia, Analysing Lexical Density and Lexical Diversity in University Stu- dents’ Written Discourse, Procedia - Social and Behav- ioral Sciences198, 546 (2015)

  27. [35]

    Koplenig, S

    A. Koplenig, S. Wolfer, and C. M¨ uller-Spitzer, Studying Lexical Dynamics and Language Change via General- ized Entropies: The Problem of Sample Size, Entropy 21(2019)

  28. [36]

    A. E. Magurran, Measuring biological diversity, Current Biology31, R1174 (2021)

  29. [37]

    Comrie,Language universals and linguistic typology: Syntax and morphology(University of Chicago press, 1989)

    B. Comrie,Language universals and linguistic typology: Syntax and morphology(University of Chicago press, 1989)

  30. [38]

    J. R. Taylor,The Oxford handbook of the word(OUP Oxford, 2015)

  31. [39]

    Bird, Nltk: the natural language toolkit, inProceed- ings of the COLING/ACL 2006 Interactive Presenta- tion Sessions(2006) pp

    S. Bird, Nltk: the natural language toolkit, inProceed- ings of the COLING/ACL 2006 Interactive Presenta- tion Sessions(2006) pp. 69–72. 14

  32. [40]

    Gerlach and F

    M. Gerlach and F. Font-Clos, A Standardized Project Gutenberg Corpus for Statistical Analysis of Natural Language and Quantitative Linguistics, Entropy22, 126 (2020)

  33. [41]

    GitHub repository with codes and access to corpora, https://github.com/pablorosillo/entropy ttr gigaword/

  34. [42]

    Hart, Project Gutenberg (1971), https://www.gutenberg.org

    M. Hart, Project Gutenberg (1971), https://www.gutenberg.org

  35. [43]

    Dunn,Natural language processing for corpus linguis- tics(Cambridge University Press, 2022)

    J. Dunn,Natural language processing for corpus linguis- tics(Cambridge University Press, 2022)

  36. [44]

    Langer, M

    L. Langer, M. Burghardt, R. Borgards, K. B¨ ohning- Gaese, R. Seppelt, and C. Wirth, The rise and fall of biodiversity in literature: A comprehensive quantifica- tion of historical changes in the use of vernacular labels for biological taxa in Western creative literature, Peop...

  37. [45]

    Q. He, S. Cheng, Z. Li, R. Xie, and Y. Xiao, Can Pre- trained Language Models Interpret Similes as Smart as Human? (2022), arXiv:2203.08452 [cs.CL]

  38. [46]

    Shade and E

    B. Shade and E. G. Altmann, Quantifying the Dissimi- larity of Texts, Information14, 271 (2023)

  39. [47]

    Corral and I

    A. Corral and I. Serra, The Brevity Law as a Scaling Law, and a Possible Origin of Zipf’s Law for Word Fre- quencies, Entropy22, 224 (2020)

  40. [48]

    Davies, Corpus del Espa˜ nol: Web/Dialects (2016), https://www.corpusdelespanol.org/web-dial/

    M. Davies, Corpus del Espa˜ nol: Web/Dialects (2016), https://www.corpusdelespanol.org/web-dial/

  41. [49]

    Biber, M

    D. Biber, M. Davies, J. K. Jones, and N. Tracy-Ventura, Spoken and written register variation in Spanish: A multi-dimensional analysis, Corpora1, 1 (2006)

  42. [50]

    Daidone, Preterite and Imperfect in Spanish Instruc- tor Oral Input and Spanish Language Corpora, Hispania 102, 45 (2019)

    D. Daidone, Preterite and Imperfect in Spanish Instruc- tor Oral Input and Spanish Language Corpora, Hispania 102, 45 (2019)

  43. [51]

    Common Crawl, Common Crawl Overview (2024), https://commoncrawl.org/overview

  44. [52]

    Conneau, K

    A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzm´ an, E. Grave, M. Ott, L. Zettle- moyer, and V. Stoyanov, Unsupervised Cross-lingual Representation Learning at Scale, inProceedings of the 58th Annual Meeting of the Association for Compu- tational Linguist...

  45. [53]

    Wenzek, M.-A

    G. Wenzek, M.-A. Lachaux, A. Conneau, V. Chaud- hary, F. Guzm´ an, A. Joulin, and E. Grave, CCNet: Ex- tracting High Quality Monolingual Datasets from Web Crawl Data, inProceedings of the Twelfth Language Re- sources and Evaluation Conference, edited by N. Calzo- lari, F. B´ e...

  46. [54]

    For our corpora, we will maintain references to Twit- ter and not to X throughout the manuscript, as Twitter was the name of the social network while the posts com- pilation was performed

  47. [55]

    Eisenstein, B

    J. Eisenstein, B. O’Connor, N. A. Smith, and E. P. Xing, Diffusion of lexical change in social media, PloS one9, e113114 (2014)

  48. [56]

    Gon¸ calves and D

    B. Gon¸ calves and D. S´ anchez, Crowdsourcing dialect characterization through Twitter, PloS one9, e112074 (2014)

  49. [57]

    Gon¸ calves, L

    B. Gon¸ calves, L. Loureiro-Porto, J. J. Ramasco, and D. S´ anchez, Mapping the Americanization of English in space and time, PloS one13, e0197741 (2018)

  50. [58]

    J. L. Abitbol, M. Karsai, J.-P. Magu´ e, J.-P. Chevrot, and E. Fleury, Socioeconomic dependencies of linguis- tic patterns in twitter: A multivariate analysis, in Proceedings of the 2018 World Wide Web Conference, WWW ’18 (International World Wide Web Conferences Steering Comm...

  51. [59]

    Grieve, C

    J. Grieve, C. Montgomery, A. Nini, A. Murakami, and D. Guo, Mapping lexical dialect variation in British En- glish using Twitter, Frontiers in Artificial Intelligence2, 11 (2019)

  52. [60]

    Alshaabi, J

    T. Alshaabi, J. L. Adams, M. V. Arnold, J. R. Minot, D. R. Dewhurst, A. J. Reagan, C. M. Danforth, and P. S. Dodds, Storywrangler: A massive exploratorium for sociolinguistic, cultural, socioeconomic, and political timelines using Twitter, Science Advances7, eabe6534 (2021)

  53. [61]

    T. Louf, D. S´ anchez, and J. J. Ramasco, Capturing the diversity of multilingual societies, Physical Review Re- search3, 043146 (2021)

  54. [62]

    E. S. Tellez, D. Moctezuma, S. Miranda, M. Graff, and G. Ruiz, Regionalized models for Spanish language vari- ations based on Twitter, Language Resources and Eval- uation57, 1697 (2023)

  55. [63]

    T. Louf, J. J. Ramasco, D. S´ anchez, and M. Kar- sai, When Dialects Collide: How Socioeconomic Mix- ing Affects Language Use (2023), arXiv:2307.10016 [physics.soc-ph]

  56. [64]

    Dunn, Syntactic variation across the grammar: mod- elling a complex adaptive system, Frontiers in Complex Systems1, 1273741 (2023)

    J. Dunn, Syntactic variation across the grammar: mod- elling a complex adaptive system, Frontiers in Complex Systems1, 1273741 (2023)

  57. [65]

    Ghosh, T

    R. Ghosh, T. Surachawala, and K. Lerman, Entropy- based Classification of ’Retweeting’ Activity on Twitter (2011), arXiv:1106.0346 [cs.SI]

  58. [66]

    Paryani, A

    J. Paryani, A. K. TK, and K. George, Entropy-based model for estimating veracity of topics from tweets, inComputational Collective Intelligence: 9th Inter- national Conference, ICCCI 2017, Nicosia, Cyprus, September 27-29, 2017, Proceedings, Part II 9 (Springer, 2017) pp. 417–427

  59. [67]

    Kanavos, G

    A. Kanavos, G. Vonitsanos, A. Mohasseb, and P. My- lonas, An Entropy-based Evaluation for Sentiment Analysis of Stock Market Prices using Twitter Data, in2020 15th International Workshop on Semantic and Social Media Adaptation and Personalization (SMA (2020) pp. 1–7

  60. [68]

    Andrew Hayden and Jason Riesa, Com- pact Language Detector 2 (CLD2) (2025), https://github.com/CLD2Owners/cld2

  61. [69]

    van Leijenhorst and T

    D. van Leijenhorst and T. van der Weide, A formal derivation of Heaps’ Law, Information Sciences170, 263 (2005)

  62. [70]

    M. A. Serrano, A. Flammini, and F. Menczer, Modeling Statistical Properties of Written Text, PloS one4, 1 (2009)

  63. [71]

    L¨ u, Z.-K

    L. L¨ u, Z.-K. Zhang, and T. Zhou, Zipf’s law leads to heaps’ law: Analyzing their relation in finite-size sys- tems, PloS one5, 1 (2010)

  64. [72]

    I. I. Eliazar and M. H. Cohen, Power-law connections: From Zipf to Heaps and beyond, Annals of Physics332, 56 (2012)

  65. [73]

    Font-Clos, G

    F. Font-Clos, G. Boleda, and ´A. Corral, A scaling law 15 beyond Zipf’s law and its relation to Heaps’ law, New Journal of Physics15, 093033 (2013)

  66. [74]

    Loreto, V

    V. Loreto, V. D. Servedio, S. H. Strogatz, and F. Tria, Dynamics on expanding spaces: modeling the emer- gence of novelties, inCreativity and Universality in Lan- guage(Springer, 2016) pp. 59–83

  67. [75]

    A. M. Petersen, J. N. Tenenbaum, S. Havlin, H. E. Stan- ley, and M. Perc, Languages cool as they expand: Allo- metric scaling and the decreasing need for new words, Scientific Reports2, 943 (2012)

  68. [76]

    Yamamoto, S

    T. Yamamoto, S. Yamada, and T. Mizuguchi, Negative correlation of word rank sequence in written texts, The European Physical Journal B94, 1 (2021)

  69. [77]

    R. F. i Cancho and R. V. Sol´ e, Two Regimes in the Fre- quency of Words and the Origins of Complex Lexicons: Zipf’s Law Revisited, Journal of Quantitative Linguis- tics8, 165 (2001)

  70. [78]

    V. V. Bochkarev, E. Y. Lerner, and A. V. Shevlyakova, Deviations in the Zipf and Heaps laws in natural lan- guages, inJournal of Physics: Conference Series, Vol. 490 (IOP Publishing, 2014) p. 012009

  71. [79]

    T. M. Apostol, An Elementary View of Euler’s Sum- mation Formula, The American Mathematical Monthly 106, 409 (1999)

  72. [80]

    Liang and J

    J. Liang and J. Todd, The stieltjes constants, J. Res. Nat. Bur. Standards Sect. B76, 161 (1972)

  73. [81]

    M. W. Coffey, Series representations for the Stieltjes constants, Rocky Mountain Journal of Mathematics44, 443 (2014)

  74. [82]

    D. S. Jones,Elementary Information Theory(Oxford University Press, Oxford, 1979)

  75. [83]

    Corominas-Murtra and R

    B. Corominas-Murtra and R. V. Sol´ e, Universality of Zipf’s law, Phys. Rev. E82, 011102 (2010)

  76. [84]

    Visser, Zipf’s law, power laws and maximum entropy, New Journal of Physics15, 043021 (2013)

    M. Visser, Zipf’s law, power laws and maximum entropy, New Journal of Physics15, 043021 (2013)

  77. [85]

    K. W. Vugrin, L. P. Swiler, R. M. Roberts, N. J. Stucky- Mack, and S. P. Sullivan, Confidence region estima- tion techniques for nonlinear regression in groundwater flow: Three case studies, Water Resources Research43 (2007)

  78. [86]

    G. J. Sz´ ekely, M. L. Rizzo, and N. K. Bakirov, Measur- ing and testing dependence by correlation of distances, The Annals of Statistics35, 2769 (2007)

  79. [87]

    Ghosh, Distance Correlation in Python (2014), https://gist.github.com/satra/aa3d19a12b74e9ab7941

    S. Ghosh, Distance Correlation in Python (2014), https://gist.github.com/satra/aa3d19a12b74e9ab7941

  80. [88]

    Nemenman, F

    I. Nemenman, F. Shafee, and W. Bialek, Entropy and inference, revisited, Advances in neural information pro- cessing systems14(2001)

  81. [89]

    Arora, C

    A. Arora, C. Meister, and R. Cotterell, Estimating the Entropy of Linguistic Distributions, inProceedings of the 60th Annual Meeting of the Association for Com- putational Linguistics (Volume 2: Short Papers), edited by S. Muresan, P. Nakov, and A. Villavicencio (Asso- ciation...

  82. [90]

    De Gregorio, D

    J. De Gregorio, D. S´ anchez, and R. Toral, Entropy Esti- mators for Markovian Sequences: A Comparative Anal- ysis, Entropy26, 79 (2024)

  83. [91]

    Bentz, T

    C. Bentz, T. Ruzsics, A. Koplenig, and T. Samardˇ zi´ c, A Comparison Between Morphological Complexity Mea- sures: Typological Data vs. Language Corpora, in Proceedings of the Workshop on Computational Lin- guistics for Linguistic Complexity (CL4LC), edited by D. Brunato, F. D...

  84. [92]

    Ebeling and T

    W. Ebeling and T. P¨ oschel, Entropy and long-range cor- relations in literary english, Europhysics Letters26, 241 (1994)

  85. [93]

    E. G. Altmann, G. Cristadoro, and M. Degli Esposti, On the origin of long-range correlations in texts, Proceed- ings of the National Academy of Sciences109, 11582 (2012)

  86. [94]

    M. A. Montemurro and P. A. Pury, Long-range fractal correlations in literary corpora, Fractals10, 451 (2002)

  87. [95]

    Dro˙ zd˙ z, P

    S. Dro˙ zd˙ z, P. O´ swi¸ ecimka, A. Kulig, J. Kwapie´ n, K. Bazarnik, I. Grabska-Gradzi´ nska, J. Rybicki, and M. Stanuszek, Quantifying origin and character of long- range correlations in narrative texts, Information Sci- ences331, 32 (2016)

  88. [96]

    Tanaka-Ishii and A

    K. Tanaka-Ishii and A. Bunde, Long-range memory in literary texts: On the universal clustering of the rare words, PloS one11, e0164658 (2016)

  89. [97]

    S´ anchez, L

    D. S´ anchez, L. Zunino, J. De Gregorio, R. Toral, and C. Mirasso, Ordinal analysis of lexical patterns, Chaos: An Interdisciplinary Journal of Nonlinear Science33 (2023)

  90. [98]

    Kulig, J

    A. Kulig, J. Kwapie´ n, T. Stanisz, and S. Dro˙ zd˙ z, In nar- rative texts punctuation marks obey the same statistics as words, Information Sciences375, 98 (2017)

  91. [99]

    Camacho and R

    J. Camacho and R. V. Sol´ e, Scaling in ecological size spectra, Europhysics Letters55, 774 (2001)

  92. [100]

    Sol´ e and J

    R. Sol´ e and J. Bascompte,Self-Organization in Com- plex Ecosystems. (MPB-42), Monographs in Population Biology (Princeton University Press, 2006)

  93. [101]

    E. G. Altmann, L. Dias, and M. Gerlach, Generalized entropies and the similarity of texts, Journal of Statis- tical Mechanics: Theory and Experiment2017, 014002 (2017)

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.