{"id":"c56283e0-aae5-450b-88cf-2f823aec8842","arxiv_id":"1908.01398","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Maximum-likelihood fitting of Zipf's law should use the distribution of sizes, not the rank-size relation, because only the former recovers simulated exponents.","lead":"This paper tests two ways of checking Zipf's law: fitting the classic rank-size plot or fitting the distribution of sizes. It finds that maximum-likelihood fitting of the size distribution recovers the true exponent in simulations, while rank-based fitting fails even when the data were generated from a rank-size power law.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Rank-size inadequacy is demonstrated only for non-truncated power laws; truncated power laws, a standard alternative for rank data, are explicitly excluded, so the broad claim needs qualification.","rationale":"The reader's weakest_assumption identifies the same scope limitation: the paper establishes the inadequacy of rank-size fitting only for non-truncated power laws with a specific ML-plus-KS procedure, while truncated power laws are common for rank-size data and are explicitly excluded. This is the most load-bearing soft spot because the strongest claim is phrased universally ('Whatever version of Zipf's law might hold in real systems...'), and the abstract repeats it without the qualification that appears only in the Discussion. The paper's analytical derivations in Sec. 2 showing the non-equivalence of the two representations are solid, and the simulation evidence for f(n) fitting is convincing for the tested model class. The concern is not internal inconsistency but scope of inference: the empirical conclusion may not generalize to truncated power-law fits, and the manuscript itself flags this. A systematic truncated-power-law test on the same simulations would settle whether the broad recommendation survives. The minor 100-versus-1000 Monte Carlo inconsistency and the choice of p>0.20 threshold are secondary to this scope issue. Since the reader's CONDITIONAL verdict already reflects this qualification, no verdict change is needed.","tokens_in":20312,"tokens_out":5322,"duration_ms":61187,"concrete_test":"Repeat the simulation protocol of Sec. 3.2 and Table II (e.g., α=1.2, Ltot=10^6, 20 replicates) and additionally fit each empirical rank-size relation with a truncated discrete power law, n(r)=A/r^α for 1≤r≤rb, estimating both α and rb by maximum likelihood. Assess goodness of fit with a Kolmogorov-Smirnov test whose null distribution is generated by simulating data from the fitted truncated law (not from the non-truncated law), and report the distribution of α̂ and the acceptance rate at p>0.20 or via BIC. If the median α̂ is within about 10% of the true α and the acceptance rate is near the nominal level, the claim that rank-size fitting is inadequate would need to be restricted to non-truncated power laws. Repeat the same check for systems simulated from the size distribution, Eq. (8), to test the 'no matter which random variable' part of the claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central normative claim, stated in the Discussion as \"the application of maximum likelihood estimation has to be done taking the size as the random variable,\" is broader than the evidence presented. The simulations show that a non-truncated power-law null fitted to ranks via ML plus a Kolmogorov-Smirnov test with p>0.20 is rejected, while fitting f(n) recovers the simulated exponents. However, the rank-size fits use only the non-truncated model of Eq. (4); truncated power laws, which are a standard model for bounded rank-size data, are explicitly excluded: \"We have not considered truncated power laws in this article.\" The abstract and the strongest claim are unqualified, while Table IV's \"PL rejected\" row for the rank representation refers only to this single model class. The Discussion even reports preliminary evidence that truncated rank-size fits produce p-values biased toward high values, but that is not a systematic test. If a truncated power-law ML fit with an estimated upper cut-off rb recovers the true α (or γ via Eq. 6) in the same synthetic systems, then the conclusion that rank-size fitting is inadequate would need at least the qualification the Discussion already reserves, and the abstract would overstate the scope.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that for systems obeying Zipf's law, the rank-size representation n(r) and the size-distribution representation f(n) are not equivalent under discreteness, and that maximum-likelihood (ML) estimation should always be applied to the size distribution rather than to ranks. The authors generate synthetic systems with a pure power law in either the rank variable or the size variable, then apply the Deluca-Corral ML procedure with a Kolmogorov-Smirnov goodness-of-fit test to both representations, in both discrete and continuous forms. Their central empirical finding is that the non-truncated power-law null for ranks is rejected in all simulated cases, whereas the same null applied to sizes is not rejected and recovers the generating exponents, with some bias. The paper concludes with a normative recommendation that ML fitting of Zipf's law must take size as the random variable, regardless of which version of Zipf's law holds.","tokens_in":20581,"tokens_out":5266,"duration_ms":57849,"significance":"If the central claim is correct, it would have a practical impact on how empirical Zipf's law is tested: researchers would be directed to fit f(n) rather than n(r), and rank-based ML exponent estimates would be considered unreliable. The paper's strengths are its transparent simulation design, full algorithmic detail in the appendix, explicit distinction between two non-equivalent definitions of Zipf's law, and candid acknowledgement of limitations (e.g., truncated power laws are not considered). The main weakness is that the unqualified conclusion exceeds the evidence, which covers only non-truncated power-law fits with a specific goodness-of-fit procedure, and the comparison between representations is not matched in sample size. These gaps undermine the universality of the abstract and Discussion statements, though the core negative result for non-truncated rank fits appears robust.","major_comments":[{"comment":"The central claim 'no matter which random variable is power-law distributed, the rank-size representation is not adequate for fitting' and the normative statement in the Discussion that 'the application of maximum likelihood estimation has to be done taking the size as the random variable' are broader than the evidence presented. The simulations only treat non-truncated power laws for the rank variable; truncated power laws are explicitly excluded ('We have not considered truncated power laws in this article'), and the Discussion reports only preliminary, non-systematic evidence that truncated rank-size fits yield inflated p-values. If a truncated power-law rank fit with an estimated upper cutoff were to recover the true α or γ in the same synthetic systems, the conclusion would require qualification. Please restrict the abstract and Table IV to non-truncated power laws or provide a systematic analysis of truncated rank-size fits.","section":"Discussion, Table IV, and Abstract"},{"comment":"The comparison between the two representations is not matched in the number of data points. Table V defines the number of data for the rank-size representation as L (tokens) and for the size distribution as V (types). In Table II, Ltot = 10^6 while Vtot is about 1.3 × 10^5, so the rank-based KS test uses roughly 7.5 times more observations than the f(n)-based test. Because the power of the KS test increases with sample size, the universal rejection of rank fits could be an artifact of the larger effective sample rather than of the representation. The authors should address this by matching sample sizes (for example, fitting ranks on a random subsample of tokens) or by demonstrating that the rejection persists when the number of observations is equal.","section":"Simulations of Zipf's systems; Tables II, III, and V"},{"comment":"The goodness-of-fit threshold and Monte Carlo simulation size are arbitrary and are reported inconsistently. The main text states that the KS statistic is calculated from 100 Monte-Carlo simulations ('p-value greater than 0.2 ... from 100 Monte-Carlo simulations'), while the Appendix derives the standard deviation using '1000 simulations' and defines the selection rule as a* = min{a such that p > 0.20}. The p > 0.20 threshold is not standard and no sensitivity analysis is provided. This matters because the chosen cutoff na determines the estimated exponent and the reported bias. Please justify the threshold or show the robustness of the conclusions to it, and correct the simulation-count inconsistency.","section":"Fitting and testing discrete power laws (Appendix); Simulations section"},{"comment":"The conceptual argument that 'the rank is not a proper random variable' is central, but the demonstration only shows that rankings obtained by ordering a finite sample deviate from the population power law. The KS test applied to ranks uses a null model that treats ranks as i.i.d. draws from a power law, which is not the actual generative process for ranks. The paper should state this misspecification explicitly and clarify that the conclusion is limited to ML+KS fitting with non-truncated power laws, not to all possible rank-based estimation procedures.","section":"Section 3.2, Fig. 1 and surrounding text"}],"minor_comments":[{"comment":"The citation to family names is empty: 'or their family names [], etc.'; a reference is missing.","section":"Introduction, first paragraph"},{"comment":"The y-axis label 'Simulated value' is unclear; it should be 'ML estimated exponent α̂' or similar.","section":"Fig. 2"},{"comment":"The phrase 'may be with some bias' is vague; please specify the direction and magnitude of the bias observed in Tables II and III.","section":"Abstract"},{"comment":"The claim that the article 'can be considered a complement or an alternative to Ref. [33]' is not elaborated; a sentence or two explaining the relationship would help the reader.","section":"Introduction, last paragraph"},{"comment":"The statement 'we were not aware of them at the time of writing and running our code' is a personal note rather than a technical justification; a neutral remark or omission would be more appropriate.","section":"Appendix, Simulation procedure"}],"recommendation":"major_revision","confidential_remarks":"The paper relies on the authors' own fitting method (Ref. [10]) and several self-citations; this is not disqualifying given the self-contained simulations. The main editorial concern is the gap between the unqualified abstract/Discussion claims and the explicit exclusion of truncated power laws. A systematic truncation analysis or a clearly scoped claim would strengthen the paper considerably."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know: the paper makes a real practical case that when you fit Zipf's law by ML, you should fit the size distribution f(n), not the rank-size relation n(r); and that case is solid for non-truncated power laws, which is a narrower scope than the abstract claims. The simulations are transparent: they generate synthetic systems from both flavors, fit both representations with discrete and continuous ML, and consistently show rank fits rejected while f(n) fits recover the input exponents. The appendix gives full algorithmic detail, so the work is reproducible.\n\nCredit where due: the systematic comparison across both representations and both simulation directions is new. Mandelbrot, Baayen, and Newman noted the non-equivalence of the two Zipf forms, but nobody had run the ML fitting comparison at this scale. The four-case summary in Table IV is clean and the message is practically useful.\n\nNow the soft spots, in proportion. The strongest normative sentence in the Discussion—“Whatever version of Zipf’s law might hold in real systems... the application of maximum likelihood estimation has to be done taking the size as the random variable”—is too strong as written. The evidence covers non-truncated power laws only, and truncated power laws are the standard workaround for rank data with a finite maximum rank. The paper explicitly says “We have not considered truncated power laws in this article,” but the abstract states the conclusion without that qualification. The stress-test note is fair: rank-size rejection is demonstrated for one model class, so the universal claim needs qualification. A secondary, minor issue is the goodness-of-fit procedure: the p>0.20 acceptance threshold is arbitrary, and the text is inconsistent about whether the KS distribution is built from 100 or 1000 Monte Carlo simulations. These are minor because the main pattern—rejections for ranks, acceptances for sizes—holds across tables and parameter values. The self-citation of the Deluca-Corral procedure is not a flaw: the method is described in the appendix and the simulations are self-contained.\n\nWho gets value: anyone doing power-law or Zipf fitting in quantitative linguistics, complex systems, or econophysics. I would cite it with the scope caveat. The paper deserves serious peer review; revision should soften the abstract and add at least a discussion—ideally a simulation—of truncated rank-size fits.","headline":"A solid simulation study showing ML fitting of Zipf's law should target the size distribution rather than ranks, but the abstract overstates the scope by ignoring truncated power laws.","tokens_in":21076,"tokens_out":1851,"would_cite":true,"duration_ms":19691,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Maximum-likelihood tests of Zipf's law should fit the size distribution, not the rank-size relation.","keywords":["Zipf's law","rank-size relation","size distribution","maximum likelihood estimation","power-law tails","discrete power laws","goodness-of-fit test","statistical inference"],"falsifier":"Simulate a system from Zipf's law for types, then fit its rank-size relation with a power law that includes an upper truncation at a maximum rank $r_b$ and compare models with an information criterion such as BIC; if such a fit accepts the rank-size power law with p>0.20 and recovers the input $\\alpha$ without bias, the claim that the rank-size representation is not adequate for fitting would be overturned.","tokens_in":20148,"feed_emoji":"📊","tokens_out":9950,"duration_ms":92528,"temperature":0.7,"pith_summary":"Zipf's law is usually written in two ways: as a power-law decay of the rank-size curve, $n(r)\\propto r^{-\\alpha}$, and as a power-law decay of the size distribution, $f(n)\\propto n^{-\\gamma}$. For discrete systems these two forms are not equivalent: a pure power law in one representation is only an asymptotic power law in the other. The paper simulates systems of both kinds and applies maximum-likelihood estimation with a goodness-of-fit test to each representation. In every simulated case the rank-size fit rejects the power-law hypothesis even when the data were generated from a true power law, while the size-distribution fit accepts it and recovers the input exponent up to a small bias. The paper's conclusion is that maximum-likelihood fitting of Zipf's law must be performed on the distribution of sizes, taking size rather than rank as the random variable.","feed_headline":"Fit Zipf's law to sizes, not ranks","feed_subtitle":"Even genuine rank-size power laws fail ML fits, while the size distribution recovers the true exponent.","key_machinery":"The central object is the pair of discrete representations of a Zipf system joined by the relation between them. A pure rank-size power law $n(r)=A r^{-\\alpha}$ translates into a survivor function $S(n)\\propto n^{-\\beta}$ with $\\beta=1/\\alpha$, so the probability mass function $f(n)=S(n)-S(n+1)$ is only asymptotically a power law, with exponent $\\gamma=1+\\beta=1+1/\\alpha$. The inferential machinery is maximum-likelihood estimation of the exponent from the tail above a lower cut-off $a$, paired with a Kolmogorov-Smirnov goodness-of-fit test whose p-values come from repeated simulation of the fitted model; the paper applies this to the rank variable and to the size variable for each simulated system. The mechanism that drives the paper's result is that rank, unlike size, is not a random variable defined before sampling: it is assigned a posteriori to the sorted sample, and the resulting distortion biases the ML exponent upward and makes the power-law fit fail the goodness-of-fit test.","core_discovery":"Using synthetic systems generated either from a power law in the hidden rank variable, $n(z)\\propto z^{-\\alpha}$ (Zipf's law for types), or from a discrete power law in size, $f(n)\\propto n^{-\\gamma}$ (Zipf's law for sizes), the paper applies maximum-likelihood estimation with a Kolmogorov-Smirnov goodness-of-fit test to both representations. For every lower cut-off considered, the rank-size relation rejects the power-law hypothesis, in both discrete and continuous versions of the estimator, even when the generating model was a true power law. The reason is that rank is assigned after sorting the sample, so the observed rank-size curve is distorted by sampling fluctuations and develops flat tails from tied smallest sizes, which no normalized non-truncated power law can reproduce. Fitting $f(n)$ instead accepts the power-law hypothesis for suitable lower cut-offs and yields exponents close to the simulated values: for example, $\\hat\\gamma\\simeq 1.86$ when the simulated asymptotic value is $\\gamma = 1.833$ in the types version, and $\\hat\\gamma\\simeq 1.835$ in the sizes version. The paper's summary claim is that whatever version of Zipf's law holds, maximum-likelihood estimation has to be applied taking size as the random variable.","pith_inferences":["Beyond the paper: if the conclusion holds, published Zipf exponents estimated from rank-frequency plots by maximum likelihood would need re-estimation from size distributions, and re-analysis of word-frequency, city-size, and income data would show how much prior conclusions change.","Beyond the paper: the paper deliberately excludes truncated power laws; a natural extension is to allow an upper cut-off in rank fits and use model selection criteria to compare representations fairly, since the authors note that truncation can inflate p-values.","Beyond the paper: the inequivalence is driven by discreteness, so continuous analogues of these quantities should not show the same split; a testable prediction is that discretizing a continuous power-law variable and then fitting the rank-size curve will reproduce the rejection, whereas fitting the original continuous variable will not."],"forward_implications":["Empirical Zipf exponents obtained by fitting rank-size plots with maximum likelihood—even for genuine power-law systems—are biased upward and fail the goodness-of-fit test, so published exponents from rank fits should not be read as reliable power-law estimates.","The two versions of Zipf's law predict different distributions of tokens into types for discrete data, connected by $\\gamma=1+1/\\alpha$ only asymptotically; evidence for one version is not evidence for the other.","For any maximum-likelihood test of Zipf's law in real systems, the size variable should be treated as the random variable and $f(n)$ fitted, choosing the lower cut-off $n_a$ by the same goodness-of-fit procedure.","When data follow Zipf's law for types, fitting $f(n)$ gives a slight positive bias in $\\hat\\gamma$ (e.g. about 1.86 instead of 1.833 in the simulated example), so recovered exponents need to be translated through the asymptotic relation rather than compared directly with rank-based values.","In systems where $f(n)$ is exactly a power law, discrete maximum-likelihood estimation is preferable to the continuous approximation because it accepts a much smaller lower cut-off, keeping more data and giving tighter uncertainties."],"supporting_citations":[{"why":"provides the standard maximum-likelihood fitting procedure for non-truncated power-law tails that the paper's procedure adapts and tests.","marker":"[8]"},{"why":"supplies the maximum-likelihood plus Kolmogorov-Smirnov goodness-of-fit method applied to both the rank-size and size-distribution representations.","marker":"[10]"},{"why":"derives the size distribution that results when tokens are sampled from a power-law rank-generating model, grounding the asymptotic exponent relation between the two Zipf forms.","marker":"[24]"},{"why":"gives the discrete power-law fitting approach that the appendix generalizes to tail selection and goodness-of-fit testing.","marker":"[44]"},{"why":"provides the rejection algorithm used to simulate discrete power-law distributed hidden ranks and sizes.","marker":"[45]"}],"fun_headline_variants":["Fit Zipf's law to sizes, not ranks — MLE study","Rank-size fitting fails even for true power laws","Zipf's law: rank fits mislead, size fits work","Maximum likelihood says: sizes beat ranks for Zipf"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion assumes that the only null model that matters is a non-truncated power law fitted by maximum likelihood and judged by a Kolmogorov-Smirnov test with a p-value threshold of 0.20; truncated power laws and alternative goodness-of-fit procedures are deliberately not considered.","fun_headline_variants_meta":{"raw":{"variants":["Fit Zipf's law to sizes, not ranks — MLE study","Rank-size fitting fails even for true power laws","Zipf's law: rank fits mislead, size fits work","Maximum likelihood says: sizes beat ranks for Zipf"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000636,"raw_usage":{"total_tokens":2964,"prompt_tokens":1008,"completion_tokens":1956,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":624,"completion_tokens_details":{"reasoning_tokens":1887}},"tokens_in":624,"tokens_out":1956,"duration_ms":14480,"temperature":1.0,"reasoning_tokens":1887,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:14:03.746856+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a system from Zipf's law for types, then fit its rank-size relation with a power law that includes an upper truncation at a maximum rank $r_b$ and compare models with an information criterion such as BIC; if such a fit accepts the rank-size power law with p>0.20 and recovers the input $\\alpha$ without bias, the claim that the rank-size representation is not adequate for fitting would be overturned.","supporting_citations":[{"cited_title":"Clauset, C","cited_arxiv_id":null,"evidence_quote":"provides the standard maximum-likelihood fitting procedure for non-truncated power-law tails that the paper's procedure adapts and tests."},{"cited_title":"Deluca and A","cited_arxiv_id":null,"evidence_quote":"supplies the maximum-likelihood plus Kolmogorov-Smirnov goodness-of-fit method applied to both the rank-size and size-distribution representations."},{"cited_title":"Mandelbrot","cited_arxiv_id":null,"evidence_quote":"derives the size distribution that results when tokens are sampled from a power-law rank-generating model, grounding the asymptotic exponent relation between the two Zipf forms."},{"cited_title":"Corral, A","cited_arxiv_id":null,"evidence_quote":"gives the discrete power-law fitting approach that the appendix generalizes to tail selection and goodness-of-fit testing."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the rejection algorithm used to simulate discrete power-law distributed hidden ranks and sizes."}],"review_version":1}