REVIEW 3 major objections 3 minor 99 references
The paper claims that the neural scaling law of foundation models can be derived deductively from Zipf's law via Heaps' law and Hilberg's hypothesis, under explicit information-theoretic assumptions.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 16:25 UTC pith:IENNYOUG
load-bearing objection A rigorous lower-bound chain from Zipf to neural scaling, with the abstract overselling a one-sided result; the non-stationary step survives the stress-test. the 3 major comments →
From Zipf's Law to Neural Scaling through Heaps' Law and Hilberg's Hypothesis
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is Proposition 16: if a stochastic text satisfies the differential Hilberg law — conditional excess entropy per symbol bounded below by a power law in the horizon — and if a trained model's entropy is bounded by compute and parameter budgets, then the worst-case expected cross entropy exceeds the entropy rate by at least the explicit expression in Eq. (90). This yields exponent bounds γ_T ≤ 1−β and γ_N ≤ 1/β−1. The paper further derives the differential Hilberg law from a differential Heaps law for stationary Santa Fe processes, and derives that Heaps law from approximate Zipf distributions under mixing conditions. The conclusion is that the observed neural scaling law can
What carries the argument
The load-bearing object is the differential Hilberg law, Eq. (86): the conditional block entropy per symbol, sup_{k≥t} H(X_{k+1}^{k+s}|X_1^t)/s − h, is bounded below by C (t+s)^{β−1}. It is a strengthened, horizon-dependent version of the plain Hilberg law, and it is what makes the entropy-budget argument go through. Alongside it sit the entropy budgets (87)–(88), which model compute and parameter counts as caps on conditional and total Shannon entropy, and the Santa Fe process, a toy source in which each token is a pair (K_t, Z_{K_t}) of a Zipf-distributed index and a copied knowledge bit, which lets Heaps-type vocabulary growth be converted into Hilberg-type entropy growth.
Load-bearing premise
The load-bearing premise is the differential Hilberg law, Eq. (86) — a horizon-wise power-law lower bound on conditional excess entropy that strengthens the empirically studied plain Hilberg law and is assumed rather than derived from it; if real text satisfies only the plain law, the neural-scaling conclusion does not follow.
What would settle it
On a large corpus, compute sup_{k≥t} H(X_{k+1}^{k+s}|X_1^t)/s − h for a range of t and s; if this conditional excess entropy decays faster than C (t+s)^{β−1} for any β close to the compression-based estimate 0.8, or vanishes over long horizons, Proposition 16's premise fails and the chain from Hilberg to neural scaling is not applicable to that corpus.
If this is right
- If the differential Hilberg law holds with exponent β, the neural scaling exponents satisfy γ_T ≤ 1−β and γ_N ≤ 1/β−1, so token-scaling and parameter-scaling exponents are determined by the same language-level exponent.
- Under the empirical compression-based estimate β≈0.8, the bounds give γ_T ≤ 0.2 and γ_N ≤ 0.25, which are loose compared with the commonly cited values γ_T≈0.095 and γ_N≈0.076; the paper suggests internet-scale corpora may have a larger β.
- The derivation predicts underparameterization (γ_T < γ_N) as the optimal regime when parameters have bounded entropy, in contrast to the overparameterization reported in practice.
- The Santa Fe process shows a Zipf-distributed IID narration over random knowledge bits satisfies Heaps' law and Hilberg's law, so Hilberg's hypothesis does not require intuitively complex structure.
- Because the final implication (C) allows arbitrary non-stationary processes, the neural-scaling step is the most robust link in the chain once its differential premise is granted.
Where Pith is reading between the lines
- The paper leaves open whether real corpora satisfy the differential Hilberg law; measuring that conditional excess entropy directly on large datasets would test the crucial premise without relying on the plain-law proxy.
- If the entropy-budget identification of parameter and compute counts is replaced by resource-bounded Kolmogorov complexity, the exponent bounds might tighten and possibly explain why overparameterized models appear better.
- A testable extension: across languages or domains with different measured β, the framework predicts correspondingly different neural scaling exponents — a cross-linguistic prediction that current scaling-law studies do not yet examine.
- The two-regime Zipf distributions and log-log convex vocabulary growth documented in quantitative linguistics would break the simple single-β picture; the framework suggests they should show up as piecewise or varying scaling exponents in model loss.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper attempts a systematic deductive chain from Zipf's law to Heaps' law, from Heaps' law to Hilberg's hypothesis, and from Hilberg's hypothesis to the neural scaling law, with Santa Fe processes as a running example. The main mathematical contributions are: (i) Propositions 13–14, deriving differential Heaps laws from Zipf-type tail conditions; (ii) Proposition 15, transferring differential Heaps laws to differential Hilberg laws for Santa Fe processes; and (iii) Proposition 16, deriving a lower bound on excess cross-entropy from a differential Hilberg law and entropy-budget constraints on the model, leading to the exponent inequalities γ_T ≤ 1−β and γ_N ≤ 1/β−1. The paper is written as a formal proof-based consolidation rather than an empirical study, and it candidly lists open problems about tightness and the meaning of compute.
Significance. If the stated assumptions are granted, the proofs provide a useful formal baseline connecting four well-known empirical regularities. The paper is honest about the main limitations: the derived neural-scaling statement is one-sided, and the differential Hilberg law is stronger than the empirically studied plain Hilberg law. Strengths include the explicit assumption-by-assumption organization, the absence of curve-fitting in the derivation, the Santa Fe process as a concrete non-vacuous example, and the clear identification of what would be needed to close the gaps. The work is best viewed as a formal lower-bound theory for scaling exponents rather than a derivation of the equality-form neural scaling law. Its significance is therefore conditional, but it is a useful contribution to the theoretical literature on scaling laws.
major comments (3)
- [§3.3, Eq. (90); Abstract] Proposition 16 proves only a lower bound on the excess expected cross-entropy. It does not prove the equality-form neural scaling law in Eqs. (8)–(10), and the resulting inequalities γ_T ≤ 1−β and γ_N ≤ 1/β−1 are one-sided. The abstract and conclusion nevertheless state that 'the neural scaling law is a consequence' and that the constraints 'produce the neural scaling law.' The open problems in §4 concede that tightness is unresolved. The paper should be recast as deriving testable lower bounds and upper bounds on exponents under explicit assumptions, not as a derivation of the empirical scaling law itself.
- [§3.3, Eq. (86); §1, Eq. (15)] Assumption (86) is a differential Hilberg law, which is substantially stronger than the empirically studied Hilberg hypothesis (6): it requires a uniform bound on conditional block entropies for all future starting points, not just the unconditional block entropy H(X_1^t). It is not derived from (6), and Proposition 15 derives it only for stationary Santa Fe processes. Thus the chain advertised in the abstract—Hilberg's hypothesis ⇒ neural scaling—does not follow for the plain Hilberg law. The paper needs either a derivation of (86) from (6) for a relevant class of processes or an explicit statement that the neural-scaling result depends on an unverified strengthening.
- [§3.3, Eqs. (90)–(95)] The expression in (90) contains the term ((1−y)/(1+y))^{1−β} with y = (c t^{−β}/(1−β))^{1/2}. If y ≥ 1, the base 1−y is non-positive and the non-integer power is not real; moreover, the definition of s_max in (89) can become negative. The proposition as stated is therefore not a valid real inequality for all c, t, n. The statement should explicitly restrict to y < 1 (e.g., c < (1−β)t^β), and state that the bound is vacuous otherwise. This does not affect the asymptotic regime of fixed c and t→∞, but the theorem statement needs a domain condition.
minor comments (3)
- [§3.2, Eq. (82)] In the proof of Proposition 15(1), the step sup_{k≥t} H(K_{k+1}^{k+s}|K_1^t)/s ≥ h might appear to assume stationarity. It is actually a consequence of Proposition 6's inf-characterization; adding a pointer would remove ambiguity.
- [§3.3, Eq. (14)] Equation (14) is displayed before the variables f, y, and the entropy-budget assumptions are introduced. Consider moving the display after Proposition 16 or defining the terms inline.
- [General] There are occasional typos, e.g., 'equvalent' in the Conclusion. More importantly, the abstract and introduction should consistently say 'lower bound on excess cross-entropy' rather than 'neural scaling law' in places where only the lower-bound result is meant.
Circularity Check
No significant circularity: the derivation is a conditional theorem chain from explicit assumptions, with no fitted input renamed as a prediction.
full rationale
The paper's central claim (Prop. 16, Eq. 90) is a conditional lower bound: given the differential Hilberg law (86) on the data's conditional block entropy and the entropy budgets (87)-(88) for the model Q, the model's worst-case excess cross entropy is bounded below by the displayed expression. This is not equivalent to its input by construction: (86) concerns the data's conditional entropy, whereas (90) concerns a model's cross entropy, and the proof bridges the two using the source coding inequality (31) and information-theoretic triangle inequalities (101)-(102). No parameter is fitted to data and no empirical value is inserted; the exponents gamma_T and gamma_N are read off as bounds implied by the derived lower bound, not as fitted outputs. The Heaps-to-Hilberg step (Prop. 15) and the Zipf-to-Heaps steps (Props. 13-14) are likewise explicit implications from stated conditions on the narration's cardinality rate, stationarity, mixing, and marginal tails. The self-citations ([26,27,28,30]) are used for the Santa Fe process as an illustrative example and for prior formulations, but the load-bearing inequalities (e.g., Eq. 80) are derived in the present paper rather than imported. No uniqueness theorem is invoked to forbid alternatives, and no ansatz is smuggled in through a citation. The paper's own acknowledged limitations, such as the coarse information-theoretic modeling of compute and parameter counts (12)-(13), concern empirical adequacy rather than circularity.
Axiom & Free-Parameter Ledger
free parameters (2)
- β (Heaps/Hilberg exponent) =
not fitted in this paper; empirical estimate ≈0.8 from PPM corpora (Takahira et al.)
- Multiplicative constants C0–C9 =
unspecified
axioms (6)
- standard math Standard information-theoretic and analytic background: Shannon entropy inequalities, Fekete lemma, Hausdorff moment theorem, Gamma-function tail bounds.
- domain assumption Approximate Zipf law p_k ≍ k^{-1/β}, or tail conditions (66)/(70)/(17).
- domain assumption Conditional repeat-probability bound p_k(t)/p_k ≤ C3 (condition 18/71), i.e., sufficiently strong mixing or finite-state/IID.
- domain assumption Santa Fe decomposition X_t = (K_t, Z_{K_t}) with independent knowledge bits H(Z_k)∈[C7,C8], and stationarity of the narration for implication B.
- ad hoc to paper Differential Heaps law (16) and differential Hilberg law (15)/(86).
- ad hoc to paper Entropy resource constraints H(Q_tnc|X_1^t) ≤ C9 c and H(Q_tnc) ≤ C9 n.
read the original abstract
We inspect the deductive connection between the neural scaling law and Zipf's law -- two statements discussed in machine learning and quantitative linguistics. The neural scaling law describes how the cross entropy rate of a foundation model -- such as a large language model -- changes with respect to the amount of training tokens, parameters, and compute. By contrast, Zipf's law posits that the distribution of tokens exhibits a power law tail. Whereas similar claims have been made in more specific settings, we show that the neural scaling law is a consequence of Zipf's law under certain broad assumptions that we reveal systematically. The derivation steps are as follows: We derive Heaps' law on the vocabulary growth from Zipf's law, Hilberg's hypothesis on the entropy scaling from Heaps' law, and the neural scaling from Hilberg's hypothesis. We illustrate these inference steps by a toy example of the Santa Fe process that satisfies all four statistical laws.
Reference graph
Works this paper leans on
-
[1]
A. Achille and S. Soatto. AI agents as universal task solvers, 2025. https://arxiv.org/abs/2510.12066
arXiv 2025
-
[2]
N. I. Akhiezer.The Classical Moment Problem and Some Related Ques- tions in Analysis. Society for Industrial and Applied Mathematics, 2021
2021
-
[3]
D. J. Aldous. Exchangeability and related topics. In ´Ecole d’ ´Et´ e de Probabilit´ es de Saint-Flour XIII — 1983, volume 1117 ofLecture Notes in Mathematics, pages 1–198. Springer, 1985
1983
-
[4]
E. G. Altmann, J. B. Pierrehumbert, and A. E. Motter. Beyond word frequency: Bursts, lulls, and scaling in the temporal distributions of words.PLoS ONE, 4:e7678, 2009
2009
-
[5]
R. H. Baayen.Word frequency distributions. Kluwer Academic Publish- ers, 2001
2001
-
[6]
Y. Bahri, E. Dyer, J. Kaplan, J. Lee, and U. Sharma. Explaining neural scaling laws.http://arxiv.org/abs/2102.06701, 2021
Pith/arXiv arXiv 2021
-
[7]
P. L. Bartlett, P. M. Long, G. Lugosi, and A. Tsigler. Benign overfitting in linear regression.Proc. Nat. Acad. Sci., 117(48):30063–30070, 2020
2020
-
[8]
M. Belkin. Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation.https://arxiv.org/ abs/2105.14368, 2021
Pith/arXiv arXiv 2021
-
[9]
Belkin, D
M. Belkin, D. Hsu, S. Ma, and S. Mandal. Reconciling modern machine- 25 learning practice and the classical bias-variance trade-off.Proc. Nat. Acad. Sci., 116(32):15849–15854, 2019
2019
-
[10]
Bengio, Y
Y. Bengio, Y. LeCun, and G. E. Hinton. Deep learning for AI.Comm. ACM, 64(7):58–65, 2021
2021
-
[11]
S. N. Bernstein. Sur les fonctions absolument monotones.Acta Math., 52:1–66, 1928
1928
-
[12]
Bialek, I
W. Bialek, I. Nemenman, and N. Tishby. Complexity through nonex- tensivity.Physica A, 302:89–99, 2001
2001
-
[13]
R. C. Bradley. Basic properties of strong mixing conditions. a survey and some open questions.Probab. Surveys, 2:107–144, 2005
2005
-
[14]
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, G. K. Ariel Herbert-Voss, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, M. L. Eric Sigler, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei. La...
2020
-
[15]
V. Cabannes, E. Dohmatob, and A. Bietti. Scaling laws for associative memories, 2024.https://arxiv.org/abs/2310.02984
Pith/arXiv arXiv 2024
-
[16]
A. Chacoma and D. H. Zanette. Heaps’ law and Heaps functions in tagged texts: Evidences of their linguistic relevance, 2020.https: //arxiv.org/abs/2001.02178
Pith/arXiv arXiv 2020
-
[17]
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Brad- bury, J. Austin, M. Isard, G. Gur-Ari, P. Yin, T. Duke, A. Lev- skaya, S. Ghemawat, S. De...
Pith/arXiv arXiv 2022
-
[18]
J. G. Cleary and I. H. Witten. Data compression using adaptive coding and partial string matching.IEEE Trans. Comm., 32:396–402, 1984
1984
-
[19]
E. U. Condon. Statistics of vocabulary.Science, 67(1733):300–300, 1928
1928
-
[20]
T. M. Cover and J. A. Thomas.Elements of Information Theory, 2nd ed.Wiley & Sons, 2006
2006
-
[21]
J. P. Crutchfield and D. P. Feldman. Regularities unseen, randomness 26 observed: The entropy convergence hierarchy.Chaos, 15:25–54, 2003
2003
-
[22]
V. Davis. Types, tokens, and hapaxes: A new heap’s law.Glottotheory, 9(2):113–129, 2018
2018
-
[23]
Devlin, M
J. Devlin, M. Chang, K. Lee, and K. Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. In J. Burstein, C. Doran, and T. Solorio, editors,Proceedings of the 2019 Conference of the North American Chapter of the Association for Com- putational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN,...
2019
-
[24]
Dingemanse, D
M. Dingemanse, D. E. Blasi, G. Lupyan, M. H. Christiansen, and M. P. Arbitrariness, iconicity, and systematicity in language.Trends Cogn. Sci., 19(10):603–615, 2015
2015
-
[25]
Dębowski
Ł. Dębowski. On Hilberg’s law and its links with Guiraud’s law.J. Quantit. Linguist., 13:81–109, 2006
2006
-
[26]
Dębowski
Ł. Dębowski. A general definition of conditional information and its application to ergodic decomposition.Statist. Probab. Lett., 79:1260– 1268, 2009
2009
-
[27]
Dębowski
Ł. Dębowski. On the vocabulary of grammar-based codes and the logical consistency of texts.IEEE Trans. Inform. Theory, 57:4589–4599, 2011
2011
-
[28]
Dębowski.Information Theory Meets Power Laws: Stochastic Pro- cesses and Language Models
Ł. Dębowski.Information Theory Meets Power Laws: Stochastic Pro- cesses and Language Models. Wiley & Sons, 2021
2021
-
[29]
Dębowski
Ł. Dębowski. A refutation of finite-state language models through Zipf’s law for factual knowledge.Entropy, 23:1148, 2021
2021
-
[30]
Ł. Dębowski. A simplistic model of neural scaling laws: Multiperiodic Santa Fe processes.https://arxiv.org/abs/2302.09049v1, 2023
Pith/arXiv arXiv 2023
-
[31]
Dębowski
Ł. Dębowski. Corrections of Zipf’s and Heaps’ laws derived from hapax rate models.J. Quantit. Linguist., 32(2):128–165, 2025
2025
-
[32]
Ebeling and G
W. Ebeling and G. Nicolis. Word frequency and entropy of symbolic sequences: a dynamical perspective.Chaos Sol. Fract., 2:635–650, 1992
1992
-
[33]
J. B. Estoup.Gammes st´ enographiques. Paris: Institut Stenographique de France, 1916
1916
-
[34]
F. Fan. An asymptotic model for the English hapax/vocabulary ratio. Comput. Linguist., 36(4):631–637, 2010
2010
-
[35]
M. Fekete. ¨Uber die Verteilung der Wurzeln bei gewissen algebraischen Gleichungen mit ganzzahligen Koeffizienten.Math. Z., 17:228–249, 1923
1923
-
[36]
Ferrer-i-Cancho and R
R. Ferrer-i-Cancho and R. V. Sol´ e. Two regimes in the frequency of words and the origins of complex lexicons: Zipf’s law revisited.J. Quan- tit. Linguist., 8(3):165–173, 2001
2001
-
[37]
Font-Clos and A
F. Font-Clos and A. Corral. Log-log convexity of type-token growth in 27 Zipf’s systems.Phys. Rev. Lett., 114:238701, 2015
2015
-
[38]
R. Futrell and M. Hahn. Linguistic structure from a bottleneck on sequential information processing.Nature Hum. Behav., 2025.https: //doi.org/10.1038/s41562-025-02336-w
-
[39]
R. Futrell and K. Mahowald. How linguistics learned to stop worrying and love the language models.https://arxiv.org/abs/2501.17047, 2025
arXiv 2025
-
[40]
G´ acs and J
P. G´ acs and J. K¨ orner. Common information is far less than mutual information.Probl. Contr. Inform. Theory, 2:119–162, 1973
1973
-
[41]
Gerlach and E
M. Gerlach and E. G. Altmann. Stochastic model for the vocabulary growth in natural languages.Phys. Rev. X, 3:021006, 2013
2013
-
[42]
Guiraud.Les caract` eres statistiques du vocabulaire
P. Guiraud.Les caract` eres statistiques du vocabulaire. Paris: Presses Universitaires de France, 1954
1954
-
[43]
Harremo¨ es and F
P. Harremo¨ es and F. Topsøe. Zipf’s law, hyperbolic distributions and entropy loss.Electr. Notes Disc. Math., 21:315–318, 2005. General Theory of Information Transfer and Combinatorics
2005
-
[44]
Hausdorff
F. Hausdorff. Momentprobleme f¨ ur ein endliches Intervall.Math. Z., 16: 220–248, 1923
1923
-
[45]
H. S. Heaps.Information Retrieval—Computational and Theoretical Aspects. Academic Press, 1978
1978
-
[46]
T. Henighan, J. Kaplan, M. Katz, M. Chen, C. Hesse, J. Jackson, H. Jun, T. B. Brown, P. Dhariwal, S. Gray, et al. Scaling laws for autoregressive generative modeling.https://arxiv.org/abs/2010.14701, 2020
Pith/arXiv arXiv 2010
-
[47]
Herdan.Quantitative Linguistics
G. Herdan.Quantitative Linguistics. Butterworths, 1964
1964
-
[48]
D. Hernandez, J. Kaplan, T. Henighan, and S. McCandlish. Scaling laws for transfer.https://arxiv.org/abs/2102.01293, 2021
Pith/arXiv arXiv 2021
-
[49]
W. Hilberg. Der bekannte Grenzwert der redundanzfreien Information in Texten — eine Fehlinterpretation der Shannonschen Experimente? Frequenz, 44:243–248, 1990
1990
-
[50]
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre. Training compute-optimal large language models, 2022. https://arxiv.org/abs/2...
Pith/arXiv arXiv 2022
-
[51]
G. Hu. On the amount of information.Teor. Verojat. Primenen., 4: 439–447, 1962
1962
-
[52]
M. Hutter. Learning curve theory.https://arxiv.org/abs/2102.040 74, 2021
2021
-
[53]
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws 28 for neural language models.https://arxiv.org/abs/2001.08361, 2020
Pith/arXiv arXiv 2001
-
[54]
S. Karlin. Central limit theorems for certain infinite urn schemes.J. Math. Mech., 17(4):373–401, 1967
1967
-
[55]
Khmaladze
E. Khmaladze. The statistical analysis of large number of rare events. Technical Report MS-R8804. Centrum voor Wiskunde en Informatica, Amsterdam, 1988
1988
-
[56]
Kobayashi and K
T. Kobayashi and K. Tanaka-Ishii. Taylor’s law for human linguistic sequences. In I. Gurevych and Y. Miyao, editors,Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1138–1148, Melbourne, Australia, 2018. Association for Computational Linguistics
2018
-
[57]
Kuraszkiewicz and J
W. Kuraszkiewicz and J. Łukaszewicz. The number of different words as a function of text length.Pamiętnik Literacki, 42(1):168–182, 1951. In Polish
1951
-
[58]
U. Lai, G. S. Randhawa, and P. Sheridan. Heaps’ law in GPT-Neo large language model emulated corpora, 2023.https://arxiv.org/abs/23 11.06377
2023
-
[59]
L. A. Levin. Universal sequential search problems.Probl. Inform. Transm., 9(3):265–266, 1973
1973
-
[60]
L. A. Levin. Randomness conservation inequalities: Information and in- dependence in mathematical theories.Inform. Control, 61:15–37, 1984
1984
-
[61]
L´ evy.Processus stochastiques et mouvement brownien
P. L´ evy.Processus stochastiques et mouvement brownien. Paris: Gau- thier Villars, 1948
1948
-
[62]
H. Li, W. Zheng, Q. Wang, Z. Ding, H. Wang, Z. Wang, S. Xuyang, N. Ding, S. Zhou, X. Zhang, and D. Jiang. Farseer: A refined scaling law in large language models.https://arxiv.org/abs/2506.10972, 2025
Pith/arXiv arXiv 2025
-
[63]
W. Li. References on Zipf’s law.https://wli-zipf.upc.edu/, 2021
2021
-
[64]
Louart, Z
C. Louart, Z. Liao, and R. Couillet. A random matrix approach to neural networks.Ann. Appl. Probab., 28(2):1190–1248, 2018
2018
-
[65]
A. Maloney, D. A. Roberts, and J. Sully. A solvable model of neural scaling laws.https://arxiv.org/abs/2210.16859, 2022
Pith/arXiv arXiv 2022
-
[66]
Mandelbrot
B. Mandelbrot. Structure formelle des textes et communication.Word, 10:1–27, 1954
1954
-
[67]
Mehri and M
A. Mehri and M. Jamaati. Variation of Zipf’s exponent in one hundred live languages: A study of the Holy Bible translations.Phys. Lett. A, 381(31):2470–2477, 2017
2017
-
[68]
E. J. Michaud, Z. Liu, U. Girit, and M. Tegmark. The quantization model of neural scaling.https://arxiv.org/abs/2303.13506, 2023
Pith/arXiv arXiv 2023
-
[69]
Mikolov, I
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean. Dis- 29 tributed representations of words and phrases and their compositional- ity. In C. J. C. Burges, L. Bottou, Z. Ghahramani, and K. Q. Weinberger, editors,Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems 2013. Proceeding...
2013
-
[70]
Miliˇ cka
J. Miliˇ cka. Type-token & hapax-token relation: A combinatorial model. Glottotheory, 2(1):99–110, 2009
2009
-
[71]
Miliˇ cka
J. Miliˇ cka. Rank-frequency relation & type-token relation: Two sides of the same coin. In I. Obradović, E. Kelih, and R. K¨ ohler, editors,Meth- ods and Applications of Quantitative Linguistics—Selected papers of the 8th International Conference on Quantitative Linguistics (QUALICO), pages 163–171. Belgrade: Academic Mind, 2013
2013
-
[72]
G. A. Miller. Some effects of intermittent silence.Amer. J. Psych., 70: 311–314, 1957
1957
-
[73]
M. A. Montemurro and D. H. Zanette. New perspectives on Zipf’s law in linguistics: from single texts to large corpora.Glottometrics, 4:87–99, 2002
2002
-
[74]
O. Neumann and C. Gros. Alphazero neural scaling and zipf’s law: a tale of board games and power laws, 2025.https://arxiv.org/abs/ 2412.11979
arXiv 2025
-
[75]
Z. Pan, S. Wang, P. Liao, and J. Li. Understanding LLM behaviors via compression: Data generation, knowledge acquisition and scaling laws, 2025.https://arxiv.org/abs/2504.09597
arXiv 2025
-
[76]
T. Pearce and J. Song. Reconciling Kaplan and Chinchilla scaling laws, 2024.https://arxiv.org/abs/2406.12907
Pith/arXiv arXiv 2024
-
[77]
Petersen, J
A. Petersen, J. Tenenbaum, S. Havlin, H. E. Stanley, and M. Perc. Languages cool as they expand: Allometric scaling and the decreasing need for new words.Sci. Rep., 2:943, 2012
2012
-
[78]
Pitman and M
J. Pitman and M. Yor. The two-parameter Poisson–Dirichlet distribu- tion derived from a stable subordinator.Ann. Probab., 25(2):855–900, 1997
1997
-
[79]
Radford, J
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsupervised multitask learners.https://open ai.com/blog/better-language-models/, 2019
2019
-
[80]
D. A. Roberts, S. Yaida, and B. Hanin.The Principles of Deep Learning Theory. Cambridge University Press, 2022
2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.