REVIEW 4 major objections 5 minor 58 references
The distinct flavors of Zipf's law in the rank-size and in the size-distribution representations, and its maximum-likelihood fitting
T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Maximum-likelihood tests of Zipf's law should fit the size distribution, not the rank-size relation.
desk verdict A solid simulation study showing ML fitting of Zipf's law should target the size distribution rather than ranks, but the abstract overstates the scope by ignoring truncated power laws. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the pair of discrete representations of a Zipf system joined by the relation between them. A pure rank-size power law $n(r)=A r^{-\alpha}$ translates into a survivor function $S(n)\propto n^{-\beta}$ with $\beta=1/\alpha$, so the probability mass function $f(n)=S(n)-S(n+1)$ is only asymptotically a power law, with exponent $\gamma=1+\beta=1+1/\alpha$. The inferential machinery is maximum-likelihood estimation of the exponent from the tail above a lower cut-off $a$, paired with a Kolmogorov-Smirnov goodness-of-fit test whose p-values come from repeated simulation of the fitted model; the paper applies this to the rank variable and to the size variable for each simulated system. The mechanism that drives the paper's result is that rank, unlike size, is not a random variable defined before sampling: it is assigned a posteriori to the sorted sample, and the resulting distortion biases the ML exponent upward and makes the power-law fit fail the goodness-of-fit test.
What would settle it
Simulate a system from Zipf's law for types, then fit its rank-size relation with a power law that includes an upper truncation at a maximum rank $r_b$ and compare models with an information criterion such as BIC; if such a fit accepts the rank-size power law with p>0.20 and recovers the input $\alpha$ without bias, the claim that the rank-size representation is not adequate for fitting would be overturned.
Extended reading notes
Core claim
Using synthetic systems generated either from a power law in the hidden rank variable, $n(z)\propto z^{-\alpha}$ (Zipf's law for types), or from a discrete power law in size, $f(n)\propto n^{-\gamma}$ (Zipf's law for sizes), the paper applies maximum-likelihood estimation with a Kolmogorov-Smirnov goodness-of-fit test to both representations. For every lower cut-off considered, the rank-size relation rejects the power-law hypothesis, in both discrete and continuous versions of the estimator, even when the generating model was a true power law. The reason is that rank is assigned after sorting the sample, so the observed rank-size curve is distorted by sampling fluctuations and develops flat tails from tied smallest sizes, which no normalized non-truncated power law can reproduce. Fitting $f(n)$ instead accepts the power-law hypothesis for suitable lower cut-offs and yields exponents close to the simulated values: for example, $\hat\gamma\simeq 1.86$ when the simulated asymptotic value is $\gamma = 1.833$ in the types version, and $\hat\gamma\simeq 1.835$ in the sizes version. The paper's summary claim is that whatever version of Zipf's law holds, maximum-likelihood estimation has to be applied taking size as the random variable.
Load-bearing premise
The conclusion assumes that the only null model that matters is a non-truncated power law fitted by maximum likelihood and judged by a Kolmogorov-Smirnov test with a p-value threshold of 0.20; truncated power laws and alternative goodness-of-fit procedures are deliberately not considered.
Editorial extensions
If this is right
- Empirical Zipf exponents obtained by fitting rank-size plots with maximum likelihood—even for genuine power-law systems—are biased upward and fail the goodness-of-fit test, so published exponents from rank fits should not be read as reliable power-law estimates.
- The two versions of Zipf's law predict different distributions of tokens into types for discrete data, connected by $\gamma=1+1/\alpha$ only asymptotically; evidence for one version is not evidence for the other.
- For any maximum-likelihood test of Zipf's law in real systems, the size variable should be treated as the random variable and $f(n)$ fitted, choosing the lower cut-off $n_a$ by the same goodness-of-fit procedure.
- When data follow Zipf's law for types, fitting $f(n)$ gives a slight positive bias in $\hat\gamma$ (e.g. about 1.86 instead of 1.833 in the simulated example), so recovered exponents need to be translated through the asymptotic relation rather than compared directly with rank-based values.
- In systems where $f(n)$ is exactly a power law, discrete maximum-likelihood estimation is preferable to the continuous approximation because it accepts a much smaller lower cut-off, keeping more data and giving tighter uncertainties.
Reading between the lines
- Beyond the paper: if the conclusion holds, published Zipf exponents estimated from rank-frequency plots by maximum likelihood would need re-estimation from size distributions, and re-analysis of word-frequency, city-size, and income data would show how much prior conclusions change.
- Beyond the paper: the paper deliberately excludes truncated power laws; a natural extension is to allow an upper cut-off in rank fits and use model selection criteria to compare representations fairly, since the authors note that truncation can inflate p-values.
- Beyond the paper: the inequivalence is driven by discreteness, so continuous analogues of these quantities should not show the same split; a testable prediction is that discretizing a continuous power-law variable and then fitting the rank-size curve will reproduce the rejection, whereas fitting the original continuous variable will not.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that for systems obeying Zipf's law, the rank-size representation n(r) and the size-distribution representation f(n) are not equivalent under discreteness, and that maximum-likelihood (ML) estimation should always be applied to the size distribution rather than to ranks. The authors generate synthetic systems with a pure power law in either the rank variable or the size variable, then apply the Deluca-Corral ML procedure with a Kolmogorov-Smirnov goodness-of-fit test to both representations, in both discrete and continuous forms. Their central empirical finding is that the non-truncated power-law null for ranks is rejected in all simulated cases, whereas the same null applied to sizes is not rejected and recovers the generating exponents, with some bias. The paper concludes with a normative recommendation that ML fitting of Zipf's law must take size as the random variable, regardless of which version of Zipf's law holds.
Significance. If the central claim is correct, it would have a practical impact on how empirical Zipf's law is tested: researchers would be directed to fit f(n) rather than n(r), and rank-based ML exponent estimates would be considered unreliable. The paper's strengths are its transparent simulation design, full algorithmic detail in the appendix, explicit distinction between two non-equivalent definitions of Zipf's law, and candid acknowledgement of limitations (e.g., truncated power laws are not considered). The main weakness is that the unqualified conclusion exceeds the evidence, which covers only non-truncated power-law fits with a specific goodness-of-fit procedure, and the comparison between representations is not matched in sample size. These gaps undermine the universality of the abstract and Discussion statements, though the core negative result for non-truncated rank fits appears robust.
major comments (4)
- [Discussion, Table IV, and Abstract] The central claim 'no matter which random variable is power-law distributed, the rank-size representation is not adequate for fitting' and the normative statement in the Discussion that 'the application of maximum likelihood estimation has to be done taking the size as the random variable' are broader than the evidence presented. The simulations only treat non-truncated power laws for the rank variable; truncated power laws are explicitly excluded ('We have not considered truncated power laws in this article'), and the Discussion reports only preliminary, non-systematic evidence that truncated rank-size fits yield inflated p-values. If a truncated power-law rank fit with an estimated upper cutoff were to recover the true α or γ in the same synthetic systems, the conclusion would require qualification. Please restrict the abstract and Table IV to non-truncated power laws or provide a systematic analysis of truncated rank-size fits.
- [Simulations of Zipf's systems; Tables II, III, and V] The comparison between the two representations is not matched in the number of data points. Table V defines the number of data for the rank-size representation as L (tokens) and for the size distribution as V (types). In Table II, Ltot = 10^6 while Vtot is about 1.3 × 10^5, so the rank-based KS test uses roughly 7.5 times more observations than the f(n)-based test. Because the power of the KS test increases with sample size, the universal rejection of rank fits could be an artifact of the larger effective sample rather than of the representation. The authors should address this by matching sample sizes (for example, fitting ranks on a random subsample of tokens) or by demonstrating that the rejection persists when the number of observations is equal.
- [Fitting and testing discrete power laws (Appendix); Simulations section] The goodness-of-fit threshold and Monte Carlo simulation size are arbitrary and are reported inconsistently. The main text states that the KS statistic is calculated from 100 Monte-Carlo simulations ('p-value greater than 0.2 ... from 100 Monte-Carlo simulations'), while the Appendix derives the standard deviation using '1000 simulations' and defines the selection rule as a* = min{a such that p > 0.20}. The p > 0.20 threshold is not standard and no sensitivity analysis is provided. This matters because the chosen cutoff na determines the estimated exponent and the reported bias. Please justify the threshold or show the robustness of the conclusions to it, and correct the simulation-count inconsistency.
- [Section 3.2, Fig. 1 and surrounding text] The conceptual argument that 'the rank is not a proper random variable' is central, but the demonstration only shows that rankings obtained by ordering a finite sample deviate from the population power law. The KS test applied to ranks uses a null model that treats ranks as i.i.d. draws from a power law, which is not the actual generative process for ranks. The paper should state this misspecification explicitly and clarify that the conclusion is limited to ML+KS fitting with non-truncated power laws, not to all possible rank-based estimation procedures.
minor comments (5)
- [Introduction, first paragraph] The citation to family names is empty: 'or their family names [], etc.'; a reference is missing.
- [Fig. 2] The y-axis label 'Simulated value' is unclear; it should be 'ML estimated exponent α̂' or similar.
- [Abstract] The phrase 'may be with some bias' is vague; please specify the direction and magnitude of the bias observed in Tables II and III.
- [Introduction, last paragraph] The claim that the article 'can be considered a complement or an alternative to Ref. [33]' is not elaborated; a sentence or two explaining the relationship would help the reader.
- [Appendix, Simulation procedure] The statement 'we were not aware of them at the time of writing and running our code' is a personal note rather than a technical justification; a neutral remark or omission would be more appropriate.
Circularity Check
No significant circularity: the simulations are self-contained; the rank-size conclusion is a model-scope limitation, not a reduction of output to input.
full rationale
The paper's central claim is that maximum-likelihood estimation should be applied to the size variable rather than to ranks. The supporting work is self-contained: synthetic systems are generated from either the type model, Eq. (4), or the size model, Eq. (8), and the ML plus Kolmogorov-Smirnov fitting procedure is fully described in the Appendix. No step in the derivation takes the conclusion as an input or renames a fitted parameter as a prediction; the reported recovery of simulated exponents is a standard simulation-validation check. The rank-size rejection arises from fitting a non-truncated power law to finite, rank-ordered data, and the paper explicitly limits this finding: 'We have not considered truncated power laws in this article' (Discussion). It also notes that a goodness-of-fit test comparing empirical rank-size data to simulated rank-size data is not contemplated in the standard algorithms and is left for future research. These are scope limitations and overstatements in the abstract, not circular reductions. The self-citations (e.g., Refs. [10] and [44]) supply the fitting algorithm, but because the algorithm is re-derived in the Appendix and applied symmetrically to both representations, the central claim does not reduce to those citations. Score 2 reflects the presence of methodological self-citations that are not load-bearing, not an actual circular derivation.
Assumptions & free parameters
free parameters (2)
- p-value threshold for acceptance =
0.20
- Monte Carlo sample size for GoF =
100 in main text, 1000 in appendix
assumptions (3)
- domain assumption The two generative models (hidden-rank power law, Eq. (14), and size-distribution power law, Eq. (8)) are adequate representatives of real Zipf systems.
- domain assumption Maximum likelihood estimation plus KS GoF with the Deluca-Corral procedure is a valid criterion for adequacy of a power-law fit.
- standard math The Euler-Maclaurin approximation to the Hurwitz zeta function and the rejection sampling algorithm are numerically accurate.
Cite this review
Pith. "Pith review of The distinct flavors of Zipf's law in the rank-size and in the size-distribution representations, and its maximum-likelihood fitting." pith.science (2026). https://pith.science/paper/3TVEWJSJ
@misc{pith2026190801398,
author = {Pith},
title = {Pith review of: The distinct flavors of Zipf's law in the rank-size and in the size-distribution representations, and its maximum-likelihood fitting},
year = {2026},
howpublished = {\url{https://pith.science/paper/3TVEWJSJ}},
note = {Machine review of arXiv:1908.01398}
}
read the original abstract
In the last years, researchers have realized the difficulties of fitting power-law distributions properly. These difficulties are higher in Zipf's systems, due to the discreteness of the variables and to the existence of two representations for these systems, i.e., two versions about which is the random variable to fit. The discreteness implies that a power law in one of the representations is not a power law in the other, and vice versa. We generate synthetic power laws in both representations and apply a state-of-the-art fitting method (based on maximum-likelihood plus a goodness-of-fit test) for each of the two random variables. It is important to stress that the method does not fit the whole distribution, but the tail, understood as the part of a distribution above a cut-off that separates non-power-law behavior from power-law behavior. We find that, no matter which random variable is power-law distributed, the rank-size representation is not adequate for fitting, whereas the representation in terms of the distribution of sizes leads to the recovery of the simulated exponents, may be with some bias.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
P. Bak. How Nature Works: The Science of Self-Organized Criticality . Copernicus, New York, 1996
work page 1996
- [2]
-
[3]
M. Mitzenmacher. A brief history of generative models for power law and lognormal distribu- tions. Internet Math., 1 (2):226–251, 2004
work page 2004
-
[4]
M. E. J. Newman. Power laws, Pareto distributions and Zipf’s law. Cont. Phys., 46:323 –351, 2005
work page 2005
-
[5]
M.V. Simkin and V.P. Roychowdhury. Re-inventing Willis. Physics Reports, 502(1):1 – 35, 2011
work page 2011
-
[6]
H. Bauke. Parameter estimation for power-law distributions by maximum likelihood methods. Eur. Phys. J. B , 58:167–173, 2007
work page 2007
-
[7]
E. P. White, B. J. Enquist, and J. L. Green. On estimating the exponent of power-law frequency distributions. Ecol., 89:905–912, 2008
work page 2008
-
[8]
A. Clauset, C. R. Shalizi, and M. E. J. Newman. Power-law distributions in empirical data. SIAM Rev., 51:661–703, 2009
work page 2009
Show all 58 references
-
[9]
Corral, F
A. Corral, F. Font, and J. Camacho. Non-characteristic half-lives in radioactive decay. Phys. Rev. E, 83:066103, 2011
2011
-
[10]
Deluca and A
A. Deluca and A. Corral. Fitting and goodness-of-fit test of non-truncated and truncated power-law distributions. Acta Geophys., 61:1351–1394, 2013
2013
-
[11]
Corral and A
A. Corral and A. Gonz´ alez. Power-law distributions in geoscience revisited. Earth Space Sci., submitted:submitted, 2018
2018
-
[12]
Scale-free Networks Well Done
Ivan Voitalov, Pim van der Hoorn, Remco van der Hofstad, and Dmitri Krioukov. Scale-free Networks Well Done. arXiv, 1811.02071, 2018
2018 arXiv
-
[13]
G. K. Zipf. Human behaviour and the principle of least effort . Addison-Wesley, Cambridge (MA), USA, 1949
1949
-
[14]
L. A. Adamic and B. A. Huberman. Zipf’s law and the Internet. Glottom., 3:143–150, 2002. 28
2002
-
[15]
Furusawa and K
C. Furusawa and K. Kaneko. Zipf’s law in gene expression. Phys. Rev. Lett., 90:088102, 2003
2003
-
[16]
R. L. Axtell. Zipf distribution of U.S. firm sizes. Science, 293:1818–1820, 2001
2001
-
[17]
Serr` a, A
J. Serr` a, A. Corral, M. Bogu˜ n´ a, M. Haro, and J. Ll. Arcos. Measuring the evolution of contemporary western popular music. Sci. Rep., 2:521, 2012
2012
-
[18]
H. Baayen. Word Frequency Distributions. Kluwer, Dordrecht, 2001
2001
-
[19]
M. Baroni. Distributions in text. In A. L¨ udeling and M. Kyt¨ o, editors, Corpus linguistics: An international handbook, Volume 2 , pages 803–821. Mouton de Gruyter, Berlin, 2009
2009
-
[20]
S. T. Piantadosi. Zipf’s law in natural language: a critical review and future directions. Psychon. Bull. Rev. , 21:1112–1130, 2014
2014
-
[21]
Malevergne, V
Y. Malevergne, V. Pisarenko, and D. Sornette. Testing the Pareto against the lognormal distributions with the uniformly most powerful unbiased test applied to the distribution of cities. Phys. Rev. E , 83:036111, 2011
2011
-
[22]
L¨ u, Z.-K
L. L¨ u, Z.-K. Zhang, and T. Zhou. Zipf’s law leads to Heaps’ law: Analyzing their relation in finite-size systems. PLoS ONE, 5(12):e14139, 12 2010
2010
-
[23]
D. Zanette. Statistical patterns in written human language . 2012
2012
-
[24]
Mandelbrot
B. Mandelbrot. On the theory of word frequencies and on related Markovian models of discourse. In R. Jakobson, editor, Structure of Language and its Mathematical Aspects , pages 190–219. American Mathematical Society, Providence, RI, 1961
1961
-
[25]
Y. Pawitan. In All Likelihood: Statistical Modelling and Inference Using Likelihood . Oxford UP, Oxford, 2001
2001
-
[26]
Gerlach and E
M. Gerlach and E. G. Altmann. Stochastic model for the vocabulary growth in natural languages. Phys. Rev. X , 3:021006, 2013
2013
-
[27]
Baixeries, B
J. Baixeries, B. Elvev˚ ag, and R. Ferrer-i-Cancho. The evolution of the exponent of Zipf’s law in language ontogeny. PLoS ONE, 8(3):e53227, 2013
2013
-
[28]
J. Tuldava. The frequency spectrum of text and vocabulary. J. Quantitative Linguistics , 3(1):38–50, 1996
1996
-
[29]
V. K. Balasubrahmanyan and S. Naranan. Quantitative linguistics and complex system stud- ies. J. Quantitative Linguistics , 3(3):177–228, 1996
1996
-
[30]
Ferrer i Cancho and R
R. Ferrer i Cancho and R. V. Sol´ e. Two regimes in the frequency of words and the origin of complex lexicons: Zipf’s law revisited. J. Quant. Linguist. , 8(3):165–173, 2001
2001
-
[31]
A. M. Petersen, J. N. Tenenbaum, S. Havlin, H. E. Stanley, and M. Perc. Languages cool 29 as they expand: Allometric scaling and the decreasing need for new words. Sci. Rep., 2:943, 2012
2012
-
[32]
Ferrer-i-Cancho and R
R. Ferrer-i-Cancho and R. Gavald` a. The frequency spectrum of finite samples from the in- termittent silence process. Journal of the American Association for Information Science and Technology, 60(4):837–843, 2009
2009
-
[33]
E. G. Altmann and M. Gerlach. Statistical laws in linguistics. In M. D. Esposti, E. G. Altmann, and F. Pachet, editors, Creativity and Universality in Language. Lecture Notes in Morphogenesis. Springer, 2016
2016
-
[34]
Font-Clos, G
F. Font-Clos, G. Boleda, and A. Corral. A scaling law beyond Zipf’s law and its relation with Heaps’ law. New J. Phys. , 15:093033, 2013
2013
-
[35]
H. S. Heaps. Information retrieval: computational and theoretical aspects . Academic Press, 1978
1978
-
[36]
Abramowitz and I
M. Abramowitz and I. A. Stegun, editors. Handbook of Mathematical Functions. Dover, New York, 1965
1965
-
[37]
N. L. Johnson, A. W. Kemp, and S. Kotz. Univariate Discrete Distributions . Wiley- Interscience, New Jersey, 3rd edition, 2005
2005
-
[38]
Peters, A
O. Peters, A. Deluca, A. Corral, J. D. Neelin, and C. E. Holloway. Universality of rain event size distributions. J. Stat. Mech. , P11030, 2010
2010
-
[39]
M. L. Goldstein, S. A. Morris, and G. G. Yen. Problems with fitting to the power-law distribution. Eur. Phys. J. B , 41:255–258, 2004
2004
-
[40]
A keystone mutualism drives pattern in a power function
S. Pueyo and R. Jovani. Comment on “A keystone mutualism drives pattern in a power function”. Science, 313:1739c–1740c, 2006
2006
-
[41]
Hergarten
S. Hergarten. Self-Organized Criticality in Earth Systems . Springer, Berlin, 2002
2002
-
[42]
S. M. Burroughs and S. F. Tebbens. Power-law scaling and probabilistic forecasting of tsunami runup heights. Pure Appl. Geophys., 162:331–342, 2005
2005
-
[43]
Casella and R
G. Casella and R. L. Berger. Statistical Inference. Duxbury, Pacific Grove CA, 2nd edition, 2002
2002
-
[44]
Corral, A
A. Corral, A. Deluca, and R. Ferrer-i-Cancho. A practical recipe to fit discrete power-law distributions. ArXiv, 1209:1270, 2012
2012
-
[45]
L. Devroye. Non-Uniform Random Variate Generation . Springer-Verlag, New York, 1986
1986
-
[46]
Christensen and N
K. Christensen and N. R. Moloney. Complexity and Criticality . Imperial College Press, 30 London, 2005
2005
-
[47]
Deluca and A
A. Deluca and A. Corral. Scale invariant events and dry spells for medium-resolution local rain data. Nonlinear Proc. Geophys., 21:555–567, 2014
2014
-
[48]
A. N. Kolmogorov. Foundations of the Theory of Probability . Chelsea Publising Company, New York, 2nd edition, 1956
1956
-
[49]
Bouchaud and A
J.-P. Bouchaud and A. Georges. Anomalous diffusion in disordered media: statistical mecha- nisms, models and physical applications. Phys. Rep., 195:127–293, 1990
1990
-
[50]
A. Corral. Scaling in the timing of extreme events. Chaos. Solit. Fract., 74:99–112, 2015
2015
-
[51]
Corral and F
A. Corral and F. Font-Clos. Dependence of exponents on text length versus finite-size scaling for word-frequency distributions. Phys. Rev. E , 96:022318, 2017
2017
-
[52]
K. P. Burnham and D. R. Anderson. Model selection and multimodel inference. A practical information-theoretic approach. Springer, New York, 2nd edition, 2002
2002
-
[53]
Ferrer i Cancho
R. Ferrer i Cancho. The variation of Zipf’s law in human language. Eur. Phys. J. B , 44:249– 257, 2005
2005
-
[54]
W. H. Press, S. A. Teukolsky, W. T. Vetterling, and B. P. Flannery. Numerical Recipes in FORTRAN. Cambridge University Press, Cambridge, 2nd edition, 1992
1992
-
[55]
L. Vepstas. An efficient algorithm for accelerating the convergence of oscillatory series, useful for computing the polylogarithm and Hurwitz zeta functions. ArXiv, page math/0702243, 2007
2007 arXiv
-
[56]
H¨ ormann and G
W. H¨ ormann and G. Derflinger. Rejection-inversion to generate variates from monotone discrete distributions. ACM Trans. Model. Comput. Simul
-
[57]
R. D. Malmgren, D. B. Stouffer, A. E. Motter, and L. A. N. Amaral. A Poissonian explanation for heavy tails in e-mail communication. Proc. Natl. Acad. Sci. USA , 105:18153–18158, 2008. 31 Ltot = 10 3 Ltot = 10 4 Ltot = 10 5 Ltot = 10 6 Simulated value Lower cut-off of the fit ra ...
2008
-
[58]
4: ML exponent from the distribution-of-sizes representation, ˆ γ, as a function of the cut- off in size, na, for a synthetic system fulfilling Zipf’s law for types, Eq
4 FIG. 4: ML exponent from the distribution-of-sizes representation, ˆ γ, as a function of the cut- off in size, na, for a synthetic system fulfilling Zipf’s law for types, Eq. (14), using the (exact) discrete fitting procedure. The original value of α is equal to 1.2 (represente...
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.