{"id":"e1450391-d37d-4e0c-a205-12098ba798bc","arxiv_id":"2411.10227","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Word entropy and type-token ratio in billion-token corpora are linked by an asymptotic formula built from Zipf and Heaps laws, confirmed across English, Spanish and Turkish texts.","lead":"This paper measures word entropy and type-token ratio in six billion-word corpora and finds the two diversity metrics are linked by a common formula. The link comes from Zipf's and Heaps' laws, giving a quantitative bridge between vocabulary richness and predictability in large texts.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. (15) is not validated by the fits: free prefactors p1/p3 deviate from unity by 1-2 orders of magnitude, and Eq. (16) has the opposite sign from Eq. (15).","rationale":"The reader's weakest_assumption concerns the 20% entropy contribution from the high-rank tail. I think there is a more direct problem: even granting the a=1 approximation, the fits in Section V do not test the predicted prefactor. Equation (15) has no free prefactor, but Eq. (16) introduces p3 (and p4), and the reported p3 values in Table V are far from 1. Similarly, Eq. (13) introduces p1, with values in Table III far from 1. The high correlation coefficients only show that a logarithmic shape fits the data; they do not confirm the beta-dependent coefficient that is the novel quantitative prediction. The sign error in Eq. (16) makes the printed validation internally inconsistent with Eq. (15). I agree with the reader that the manuscript should not be accepted as is, and this concern is central rather than peripheral: it affects the main quantitative claim. A revised version that fixes the sign and re-fits with p3 fixed to unity, or that explicitly reframes the result as an empirical scaling law with fitted language-dependent prefactors, could be viable; but the current manuscript overstates the agreement.","tokens_in":18542,"tokens_out":9717,"duration_ms":97369,"concrete_test":"Re-run the H vs TTR analysis with p3 fixed to 1 and only the intercept p4 free, using the correct positive coefficient beta/(2(1-beta)) from Eq. (15) over the same L range; compute the residual sum of squares or a reduced chi-square and compare with the free-p3 fit. If the fixed-coefficient model is rejected or shows systematic residuals, the reported agreement is an artifact of the free prefactor. Also inspect the repository in Ref. [41] to verify whether Eq. (16) was implemented with beta/(2(beta-1)) or beta/(2(1-beta)); if the printed sign is the implemented one, the fit is internally inconsistent.","verdict_should_be":"REJECT","load_bearing_attack":"The central quantitative claim is Eq. (15), which predicts a fixed coefficient beta/(2(1-beta)) multiplying ln(1/TTR). The validation in Section V fits Eq. (16), which multiplies the whole predicted functional form by a free parameter p3 and adds p4. Since p3 is free, the fit tests only the shape, not the predicted coefficient. Table V reports p3 values of 0.04, 0.61, 0.24, 0.49, 0.43, and 2.48, systematically far from the predicted value of 1. The same issue appears in the H(L) fits: p1 in Table III is 0.03-0.06 for five corpora, also far from 1. A free prefactor that deviates by one to two orders of magnitude means the empirical slopes are not consistent with the asymptotic expression. Moreover, Eq. (16) is written with beta/(2(beta-1)) = -beta/(2(1-beta)), giving the leading term the opposite sign from Eq. (15). With the positive p3 values reported, the fitted function would decrease as ln(1/TTR) increases, contradicting the data in Fig. 6 unless the printed equation is a typo. Thus the paper's claim of quantitative agreement with Eq. (15) is not supported by the evidence presented.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper examines the relation between Shannon word entropy H and the type-token ratio TTR in six gigaword corpora (English, Spanish, Turkish; books, web, and Twitter). The authors report a negative empirical correlation between H and TTR and derive an asymptotic analytical expression, H ~ β/(2(1−β)) ln(1/TTR) + ln ln(1/TTR) (Eq. 15), from Zipf's law with exponent a=1 and Heaps' law V=αL^β. They validate this expression by fitting semi-empirical forms with free prefactors and offsets (Eqs. 13 and 16) and report high correlation coefficients for most corpora. The paper argues that, for sufficiently large texts, lexical diversity measures such as TTR and entropy are functionally dependent, contrary to naive expectation.","tokens_in":18853,"tokens_out":3312,"duration_ms":32733,"significance":"If the analytical relation were quantitatively confirmed, the result would be of interest to quantitative linguistics and NLP: it would connect two widely used diversity metrics and imply that they are not independent for large texts. The paper's strengths include the use of very large corpora (over 10^9 tokens each) across three morphologically distinct languages and multiple registers, a reproducible analysis with publicly available code, a clear derivation from Zipf and Heaps laws, and a robust empirical demonstration of a negative H–TTR correlation. The main weakness is that the quantitative validation relies on free prefactors that deviate strongly from the predicted values, and the printed validation equation contains a sign error; as a result, the paper does not currently establish the claimed agreement with the asymptotic expression.","major_comments":[{"comment":"The printed Eq. (16) contains a sign error: β/(2(β−1)) equals −β/(2(1−β)), so the leading term in Eq. (16) has the opposite sign to Eq. (15). Since the fitted p3 values in Table V are positive, the fitted function would decrease as ln(1/TTR) increases, contradicting the data in Fig. 6 unless the equation is a typo. This makes the validation as printed internally inconsistent and needs correction before the claimed agreement can be assessed.","section":"Section V, Eq. (16)"},{"comment":"The fits use free prefactors p1 and p3 that multiply the entire predicted expression. The predicted coefficients in Eqs. (12) and (15) are β/2 and β/(2(1−β)), respectively, corresponding to p1=1 and p3=1. Table III reports p1=0.03–0.27 and Table V reports p3=0.04–2.48, with most values one to two orders of magnitude from unity. With free prefactors, the fits test only the functional form and not the predicted coefficient, so the statement in Section V that the data agree with Eq. (15) is not supported by the evidence presented.","section":"Section III.C, Table III and Section V, Table V"},{"comment":"The assumption that the second Zipf regime has a negligible contribution to the entropy is not supported by Table VI: R_H ranges from 0.14 to 0.24 and R'_H from 0.12 to 0.21, meaning the second regime contributes roughly 12–24% of the total entropy. This is not negligible, and if this fraction varies across corpora it changes the prefactor of the asymptotic relation derived from a pure a=1 Zipf law. The fitted p1 values far from unity are consistent with this concern and suggest that the two-regime structure materially affects the prefactor.","section":"Section III.B and Appendix B, Table VI"},{"comment":"The range of L over which the fits are performed is selected post hoc using the condition 0.0025 < σ(H) < 0.025, with different ad hoc ranges for the two Turkish corpora (e.g., for TRCC100 a fixed vocabulary interval is chosen). Because the same data are used to select the fitting range and to compute the reported goodness-of-fit values, the high ρ² values in Tables III and V are not an out-of-sample confirmation. The authors should provide a pre-registered or otherwise justified range selection, or at least show that the fitted parameters are stable under reasonable variations of the range.","section":"Section III.C and Appendix C"},{"comment":"The TwTR corpus has ρ²=0.60 and p3=2.48, indicating a poor fit and a prefactor far from the predicted value. The paper acknowledges the poor statistics for TwTR but still concludes an 'overall good agreement' across the six corpora. With one of six corpora failing the quantitative test, the universality claim in the abstract and conclusion is overstated and should be qualified.","section":"Table V, TwTR row"}],"minor_comments":[{"comment":"The notation 'ln ln TTR^{-1}' is ambiguous; it should be written as ln(ln(1/TTR)) to avoid the possible misreading (ln ln TTR)^{-1}.","section":"Eqs. (15) and (16)"},{"comment":"The text says H_max corresponds to L=1.8×10^9 for all corpora, but SPGC and TRCC100 have total lengths larger than this; please clarify how H_max is computed for those corpora.","section":"Section V, Fig. 5"},{"comment":"The fitted p1 values are reported without discussion; given that they deviate strongly from the predicted value of 1, the authors should either explain the deviation or temper the claim of 'excellent agreement'.","section":"Section III.C, Table III"},{"comment":"The sentence 'This observation would suggest that type-token ratio and entropy are distinct and uncorrelated diversity measures' could be misread as a claim that they are uncorrelated in general; consider rephrasing to 'potentially uncorrelated'.","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":"The central empirical finding—a robust negative H–TTR correlation—is likely sound and valuable, but the analytical validation as presented is not quantitative because of the free prefactors and the sign error in Eq. (16). The paper could be made publishable if the authors (i) correct the sign error, (ii) test the predicted coefficients directly (e.g., by fixing p1=p3=1 or by reporting the actual slope and its confidence interval), (iii) address the tail-contribution concern with a quantitative bound, and (iv) qualify the claims for TwTR. Given the scope of these changes, major revision is appropriate. The paper fits the journal's profile; the issue is not scope but rigor."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nWhat you should know: this paper derives an analytic relation between Shannon word entropy and type-token ratio from Zipf and Heaps laws, and shows a consistent empirical H–TTR curve across six gigaword corpora. The empirical curve is real and worth knowing about. The analytic validation, however, is shaky: the fits to Eq. (15) use free prefactors p1 and p3 that are far from the predicted value of 1, and Eq. (16) as printed has the opposite sign from Eq. (15).\n\nWhat is actually new: combining the known asymptotic entropy of a Zipf distribution (H ~ 1/2 ln V + ln ln V) with Heaps' law to express H as a function of TTR is a natural but new step, and testing it on English, Spanish, and Turkish corpora from books, web, and Twitter is a solid empirical contribution. The authors provide code and data, check robustness with two entropy estimators, and are transparent about limitations like the overall versus mean-segmental TTR, unigram entropy, and the problematic Turkish Twitter corpus. Credit where due: the near-collapse of the H–TTR curves after shifting by H_max is an interesting object for quantitative linguistics.\n\nSoft spots, in proportion:\n\n1. The sign inconsistency between Eq. (15) and Eq. (16) is real. With beta in (0,1), beta/(2(beta-1)) is negative, so the fitted form in Eq. (16) would decrease with ln(1/TTR), contradicting both Eq. (15) and the data in Fig. 6. Almost certainly a typo, but it undermines the reader's trust.\n\n2. The free prefactors p1 and p3 make the validation much weaker than the text suggests. p1 is 0.03–0.06 for five corpora, not near 1; p3 ranges from 0.04 to 2.48. The fits test the shape of the logarithmic dependence, not the predicted coefficient. The paper calls the fits semi-empirical, which helps, but the abstract says the expression \"agrees with our empirical findings.\" That overstates what the evidence shows.\n\n3. The sigma(H)-based range selection is post hoc, though the authors explain it. The Turkish corpora need special handling, which is acknowledged.\n\n4. The derivation itself is fine as a leading-order asymptotic argument; the tail regime contributes roughly 20% of the entropy, so the a=1 approximation is acceptable for a first pass.\n\nWho is this for? Quantitative linguists and anyone using lexical diversity metrics who needs to know that H and TTR are not independent for large texts. It deserves a serious referee, but the authors should fix the sign error and either correct the coefficient prediction or reframe the contribution as an empirical scaling law with corpus-dependent prefactors. I would not cite the quantitative claim until that is resolved, but I would cite the empirical H–TTR curve.\n\nRecommendation: send to peer review with major revision. The empirical finding is solid enough to warrant referee time, but the analytic validation needs an honest rewrite.","headline":"Useful empirical relation between entropy and TTR, but the analytic validation is weaker than claimed because the fits use free prefactors that deviate from the predicted values by 1–2 orders of magnitude, and Eq. (16) has a sign error.","tokens_in":19433,"tokens_out":3138,"would_cite":true,"duration_ms":33035,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"For large natural-language texts, Shannon word entropy and type-token ratio are linked by a single relation.","keywords":["Shannon entropy","type-token ratio","Zipf law","Heaps law","lexical diversity","gigaword corpus","statistical language universals"],"falsifier":"Compute $H$ and $TTR$ on a corpus with a Zipf exponent in the kernel regime clearly different from one (for example a highly isolating language) and test whether the data obey $H = \\frac{\\beta}{2(1-\\beta)}\\ln(1/TTR) + \\ln\\ln(1/TTR)$ with no fitted multiplicative parameter; a systematic failure of this parameter-free prediction would falsify the universality of the relation.","tokens_in":18308,"feed_emoji":"📊","tokens_out":8458,"duration_ms":77096,"temperature":0.7,"pith_summary":"The paper tries to establish that the two standard measures of lexical diversity, Shannon word entropy and type-token ratio, are not independent for sufficiently long texts but are tied by an explicit asymptotic relation. The relation is derived from Zipf's law for word frequencies and Heaps' law for vocabulary growth, and it is checked against six corpora of more than a billion tokens each, spanning three languages and three registers. A careful reader would care because the result turns two superficially different metrics into two projections of the same underlying statistical regularity of language, with consequences for how entropy and diversity are estimated in large text data.","feed_headline":"One formula ties word entropy to type-token ratio in huge corpora","feed_subtitle":"Six billion-token corpora in English, Spanish, and Turkish confirm an asymptotic H–TTR relation from Zipf and Heaps laws.","key_machinery":"The argument is carried by two empirical laws of quantitative linguistics: Zipf's rank-frequency law $f = k/r^a$ with $a\\simeq 1$ over the kernel vocabulary, and Heaps' law $V = \\alpha L^\\beta$ for vocabulary growth. The asymptotic expansion of the harmonic sum gives $H \\sim \\frac{1}{2}\\ln V + \\ln \\ln V$, and substituting Heaps law yields both the entropy-vs-length formula and, through $TTR = V/L$, the final $H$–$TTR$ relation. The load-bearing numerical fact is that ranks in the second Zipf regime contribute roughly 20% of the entropy, so treating the whole distribution as a single $a=1$ power law is a good approximation.","core_discovery":"The central claim is that for $TTR \\ll 1$, the Shannon word entropy of a natural-language text follows $H \\sim \\frac{\\beta}{2(1-\\beta)} \\ln(1/TTR) + \\ln \\ln(1/TTR)$, where $\\beta$ is the Heaps exponent. This expression results from combining the large-vocabulary entropy of a Zipf distribution with exponent one, $H \\sim \\frac{1}{2}\\ln V + \\ln \\ln V$, with the Heaps-law observation that $TTR \\sim L^{\\beta-1}$. The authors verify the functional form on six gigaword corpora, showing that the same curve describes books, web media, and tweets in English, Spanish, and Turkish, with per-corpus fitting parameters absorbing the residual differences. They also show that the second, steeper regime of the Zipf distribution contributes only about 20% of the total entropy, which justifies the single-exponent derivation.","pith_inferences":["Editorial inference: The derivation uses only Zipf and Heaps laws, so the same $H$–$TTR$ relation should appear in other heavy-tailed systems with type counts and sublinear type growth, such as ecological abundance data or genomic k-mer counts; this is a testable cross-domain prediction the authors list only as future work.","Editorial inference: The fitted slopes are well below the asymptotic prefactor, implying that real corpora of $10^9$ tokens have not yet reached the clean limit; practical entropy estimates from TTR should therefore be calibrated per corpus rather than used with the theoretical prefactor.","Editorial inference: The relation can serve as a distributional diagnostic for text generators: if a language model's long outputs trace a different $H$–$TTR$ curve than human corpora, that signals a systematic over- or under-repetition that perplexity alone would miss.","Editorial inference: A natural next test is to vary tokenization (subword vs. word) — the relation should still hold with modified exponents, since subword units also obey Zipf-like and Heaps-like laws in many languages."],"forward_implications":["In the asymptotic regime $TTR \\ll 1$, the entropy and the type-token ratio carry the same information: the Heaps exponent and one of the two quantities determine the other.","Very large corpora can be cross-validated: estimates of $H$ from word-frequency tables should agree with TTR-based estimates through Eq. (15) once corpus-specific constants are calibrated.","Any corpus in any language that obeys Zipf and Heaps laws is predicted to fall on the same universal $H$–$TTR$ curve, with only the parameters $\\beta$ and the additive constant varying.","The over-all (length-dependent) definition of TTR is essential; a mean-segmental TTR would produce a positive instead of negative correlation with entropy, as the paper notes."],"supporting_citations":[{"why":"Zipf's law, the rank-frequency power law used to compute the entropy asymptotic.","marker":"[23]"},{"why":"Heaps (Herdan) law, which links vocabulary size to text length and yields the TTR scaling.","marker":"[24, 25]"},{"why":"Review of statistical laws of language that motivates the assumption that Zipf and Heaps hold broadly.","marker":"[26]"},{"why":"Standardized Project Gutenberg Corpus and its preprocessing, the source and cleaning of one of the six datasets.","marker":"[40]"},{"why":"Asymptotic entropy of the Zipf distribution, the source of $H \\sim \\frac{1}{2}\\ln V + \\ln \\ln V$.","marker":"[82-84]"},{"why":"NSB entropy estimator, the independent estimator used to confirm the plug-in entropy measurements.","marker":"[88]"},{"why":"Prior work showing a positive correlation between segmental TTR and entropy, clarifying why the over-all TTR convention is necessary.","marker":"[91]"},{"why":"Cross-linguistic word entropy results, providing context for the language-specific entropy differences.","marker":"[12]"}],"fun_headline_variants":["One law connects entropy and type-token ratio across six corpora","Zipf and Heaps laws unite entropy and type-token ratio","Gigaword test verifies entropy-type-token formula","Three languages, one curve: entropy vs type-token ratio"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the assumption that a single Zipf law with exponent exactly one captures enough of the word-frequency distribution that the higher-rank tail contributes only about a fifth of the entropy, and that this tail contribution stays roughly constant; if the tail dominates or varies with text length, the derived prefactors and the whole asymptotic relation change.","fun_headline_variants_meta":{"raw":{"variants":["One law connects entropy and type-token ratio across six corpora","Zipf and Heaps laws unite entropy and type-token ratio","Gigaword test verifies entropy-type-token formula","Three languages, one curve: entropy vs type-token ratio"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000874,"raw_usage":{"total_tokens":3753,"prompt_tokens":890,"completion_tokens":2863,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":506,"completion_tokens_details":{"reasoning_tokens":2790}},"tokens_in":506,"tokens_out":2863,"duration_ms":21218,"temperature":1.0,"reasoning_tokens":2790,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:50:00.133178+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute $H$ and $TTR$ on a corpus with a Zipf exponent in the kernel regime clearly different from one (for example a highly isolating language) and test whether the data obey $H = \\frac{\\beta}{2(1-\\beta)}\\ln(1/TTR) + \\ln\\ln(1/TTR)$ with no fitted multiplicative parameter; a systematic failure of this parameter-free prediction would falsify the universality of the relation.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Zipf's law, the rank-frequency power law used to compute the entropy asymptotic."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Review of statistical laws of language that motivates the assumption that Zipf and Heaps hold broadly."},{"cited_title":"Gerlach and F","cited_arxiv_id":null,"evidence_quote":"Standardized Project Gutenberg Corpus and its preprocessing, the source and cleaning of one of the six datasets."},{"cited_title":"Nemenman, F","cited_arxiv_id":null,"evidence_quote":"NSB entropy estimator, the independent estimator used to confirm the plug-in entropy measurements."},{"cited_title":"Bentz, T","cited_arxiv_id":null,"evidence_quote":"Prior work showing a positive correlation between segmental TTR and entropy, clarifying why the over-all TTR convention is necessary."}],"review_version":1}