REVIEW 3 major objections 4 minor 26 references
Transforming the Erd\H{o}s-Kac theorem
T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper transforms the Erdős–Kac theorem to build interval estimates for the prime-divisor count, with a square-root variance stabilizer, a three-quarters width optimizer, and a trained score interval for small integers.
desk verdict The delta-method/Box-Cox part of this paper is correct and genuinely useful, but the trained-interval reliability claims sit on a calibration that the paper never explains. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by the delta method fitted into the probabilistic framework of additive arithmetic functions. Starting from the central limit theorem for additive functions, the paper introduces the remainder function $r(x) = (g(x)-g(1))/(x-1) - g'(1)$, which is continuous at $1$, and uses Slutsky's theorem under the discrete uniform measure to prove that $g(f_n/A_n)$ is asymptotically normal whenever the additive function $f_n$ is (Theorem 2.2). Specializing $f_n=\omega$, $A_n=\ell_2 + O(1)$, $B_n^2=\ell_2 + O(1)$ gives Theorem 1.1. The Box-Cox transformation $y_\lambda(x) = (x^\lambda-1)/\lambda$ (or $\log x$ at $\lambda=0$) turns this into an explicit $\lambda$-family of normal approximations; the asymptotic expansion of the interval width, whose first correction term contains the factor $(\lambda-1)(2\lambda-1)$, identifies $\lambda=3/4$ as the width-minimizer. For small integers, the paper replaces ordinary coverage probability with a fuzzy coverage probability that partially credits the fractional upper and lower endpoints of an interval, and it estimates local adjustment functions $f_{\mu,\lambda}$ and $f_{\sigma,\lambda}$ by fitting power functions of linear functions of $\ell_2$ to smoothed $\omega$. The score interval estimate comes from solving $(\omega-\ell_2)/\sqrt{\omega} = \pm z$ for $\omega$, in the manner of Wilson's interval for a binomial proportion.
What would settle it
Take the actual values of $\omega(m)$ for $m$ near $10^8$ and $10^{12}$, compute the smoothed mean and standard deviation of $\omega^\lambda$ with the same window and powers as Section 7, and compare them with the fitted functions in Table 5; a deviation of more than a few percent would show that the power-of-a-linear-function model does not extrapolate, and the trained interval estimates' claimed out-of-sample reliability would fail. An even more direct check is to compute the exact fuzzy coverage probabilities of the trained score and Poisson intervals at $m=10^{13}$ and $m=10^{14}$ and see whether they remain inside Bradley's liberal band around the nominal $0.6319$.
Extended reading notes
Core claim
The central discovery is Theorem 1.1: if $g$ is differentiable at $1$ and $g'(1)\neq 0$, then under the discrete uniform measure on $\{1,\dots,n\}$, the transformed ratio $(g(\omega/\ell_2)-g(1))/(g'(1)/\sqrt{\ell_2})$ converges in distribution to the standard normal $\Phi$. With $g$ equal to the Box-Cox power $y_\lambda$, this becomes $(\omega^\lambda - \ell_2^\lambda)/(\lambda \ell_2^{\lambda-1/2}) \Rightarrow \Phi$, and the same argument gives versions with local adjustment functions $f_\mu$ and $f_\sigma$ in place of $\ell_2$ and $\sqrt{\ell_2}$. The two distinguished powers are $\lambda=1/2$, where the denominator no longer depends on $m$ so the variance is stabilized, and $\lambda=3/4$, which minimizes the asymptotic width of the two-sided interval because the first nonconstant term in the width's expansion is proportional to $(\lambda-1)(2\lambda-1)$. Solving the quadratic obtained by standardizing with $\omega$ in the denominator gives a score interval estimate in the spirit of Wilson, and numerical work with fuzzy coverage probabilities shows that the score interval, and the Poisson interval based on Landau's formula, keep their coverage closest to nominal for $m$ between $10^5$ and $10^{14}$ once their means and standard deviations are estimated from $\omega$ values on $[10^4,10^6]$. The same transformation machinery is applied to the Erdős–Pomerance theorem, where it yields variance stabilization at $\lambda=1/4$ while the optimal width remains at $\lambda=3/4$.
Load-bearing premise
The load-bearing premise outside the asymptotic theory is the Section 7 working assumption that the mean and standard deviation of the transformed prime-divisor count can be represented as fixed powers of linear functions of $\log\log m$, with coefficients fitted on $m\in[10^4,10^6]$ and trusted to extrapolate to $m=10^{14}$ and beyond; if that empirical model drifts, the claimed reliability of the trained score and Poisson intervals for out-of-sample $m$ does not follow.
Editorial extensions
If this is right
- The square-root transformation yields a simple rarity rule: integers with $|\sqrt{\omega(m)} - \sqrt{\log\log m}| > 1.5$ make up roughly 1 in 400 of all integers, and >2.0 roughly 1 in 16,000, giving an easily remembered scale for interpreting $\omega$.
- The three-quarters power gives the asymptotically narrowest two-sided $100(1-\alpha)\%$ interval for $\omega$ near large $m$: $[\ell_2(1 - 3z/(4\sqrt{\ell_2}))^{4/3}, \ell_2(1 + 3z/(4\sqrt{\ell_2}))^{4/3}]$, a direct improvement over the usual $\ell_2 \pm z\sqrt{\ell_2}$ interval.
- The score interval estimate of $\omega$, obtained by solving the quadratic $(\omega-\ell_2)^2/\omega = z^2$, is the most reliable of the normal-based intervals after training, with fuzzy coverage probabilities close to nominal for $m \in [10^5, 10^{14}]$.
- The Poisson interval estimate, centered at $\ell_2(m)+1$ with width calibrated to the nominal level, is relatively reliable even without training, and it tends to be narrower than the transformed Erdős–Kac intervals (e.g., at $m\approx 10^{70}$ it gives $[4.52, 7.65]$ versus Billingsley's $[3.05, 7.11]$).
- The transformation results transfer to the Erdős–Pomerance theorem: $\omega(\varphi(m))$ is normal after a Box-Cox-type transform with variance stabilization at $\lambda=1/4$ and asymptotically optimal interval width at $\lambda=3/4$.
Reading between the lines
- Because variance stabilization at $\lambda=1/2$ makes deviations comparable across different $m$, one could define a universal rarity score $z = (\sqrt{\omega} - \sqrt{\ell_2})/0.5$ and rank integers of vastly different sizes on one scale; the paper does not propose such a score but its own Table 1 is the seed of it.
- The training procedure is only validated on $m$ up to $10^{14}$; a natural robustness test is to retrain on shifted windows (e.g., $10^6$–$10^8$) and see whether the fitted power-law forms in Table 5 drift, which would indicate that the choice of training range, rather than the asymptotic theory, drives the reported coverage.
- The relative success of the Poisson interval estimate supports a shifted-Poisson model as a better finite-sample description of $\omega$ than the normal law; the paper's discussion of Landau's formula and the shifted-Poisson mass function marks the Poisson approximation as a natural target for a rate-of-convergence theorem.
- The fuzzy-coverage device is transferable: any interval for a lattice-valued statistic whose endpoints fall between integers suffers the same jump problem, so the same fractional-credit definition could be applied to binomial, Poisson, or hypergeometric intervals, not just to $\omega$.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper applies the delta method to Billingsley's probabilistic version of the Erdős–Kac theorem and obtains a central limit theorem for g(ω/ℓ2) when g'(1)≠0. Choosing g as a Box–Cox power gives a transformed Erdős–Kac theorem; the paper identifies λ=1/2 as the variance-stabilizing transformation and λ=3/4 as the power that asymptotically minimizes the width of the two-sided interval estimate for ω. It then introduces local adjustment functions to refine the intervals for finite m, develops fuzzy coverage probabilities to handle the discreteness of ω, and compares five interval estimates (λ=1/2, 3/4, 1, Poisson, and score) for m∈[10^5,10^14]. After fitting mean and standard deviation functions on m∈[10^4,10^6], it claims that the trained score and Poisson intervals are reliable. The paper also states analogous transformed Erdős–Pomerance theorems.
Significance. If the main claims hold, the theoretical part is a clean and useful contribution: it gives a principled family of Erdős–Kac-based intervals, and the variance-stabilization and width-optimality results are natural and well-motivated. The asymptotic expansion behind Theorem 4.1 is correct, and the fuzzy-coverage framework is an appropriate response to the discreteness of ω. The claimed numerical advantage of the trained score and Poisson intervals, however, is not supported by the calibration evidence as presented: the fitted standard deviations in Table 5 are markedly smaller than the Erdős–Kac variance even inside the training range, so the reliability conclusion rests on unverified numerics. The paper would also be strengthened by providing the omitted proofs and by making the numerical evaluation reproducible.
major comments (3)
- [Section 7, Table 5, Figure 8] The fitted scale parameters in Table 5 appear inconsistent with the Erdős–Kac variance on the training range. At m=10^5 (ℓ2≈2.44), the λ=1 value gives f̂σ,1≈0.0499+0.3677·2.44≈0.95, whereas the Erdős–Kac variance of ω at that scale is ℓ2≈2.44, so the empirical standard deviation should be about 1.56, not 0.95. For λ=1/2, the fitted value is about 0.60 instead of the variance-stabilized value 1. This is not a small finite-sample correction. Since the trained interval widths in Figure 8 are built from these f̂σ,λ values, the reported reliability of the trained Box-Cox, score, and Poisson intervals is not established. Please reconcile Table 5 with a direct computation of the residual standard deviations on the training range, or supply the code/data so that the reader can verify the calibration.
- [Theorems 3.1, 8.1–8.3] Theorem 3.1 is the central result used for all subsequent interval constructions, and Theorems 8.1–8.3 are stated in Section 8, but their proofs are omitted 'for brevity.' At least for Theorem 3.1, a short derivation from Theorem 1.1 should be included; for Section 8, either provide the analogous delta-method derivations or state explicitly that they follow from (25) by the same argument. As written, the reader cannot verify the λ=0 cases or the rates without reconstructing the calculations.
- [Theorem 8.1 and Theorem 8.2] The displayed denominator in Theorem 8.1, 2g'(1)/(√3 ℓ2), is not consistent with Theorem 8.2 or with (25). Substituting g(x)=yλ(x) into that denominator and into the numerator gives a ratio with a denominator containing ℓ2^{2λ−1}, not ℓ2^{2λ−1/2} as in Theorem 8.2; the λ=1 case would not reduce to (25). The correct denominator should be 2g'(1)/(√3√ℓ2). The same notational correction is needed in the λ=0 statement of Theorem 8.2 and in the following sentence, so that the variance-stabilizing power λ=1/4 is derived from the correct rate.
minor comments (4)
- [Section 1] There are several typographical slips, including 'the the' in the introduction and 'the the vicinity' in Section 4; a careful proofreading pass is needed.
- [Section 6] The sentence beginning 'when ⌈Lλ,α(m)⌉=⌊Uλ,α(m)⌋+1, we may further assume that at least one of ... is included' is not reflected in the fuzzy coverage formula that follows; the intended convention should be stated precisely or removed.
- [Section 7] The numerical results in Tables 4–5 and Figures 4–8 are not reproducible from the text alone because no code or detailed data-processing pipeline is given; providing the code or a detailed pseudocode would substantially help the reader check the calibration issue raised above.
- [Section 9] The claim that 'all the theoretical results follow even if we replace ω with Ω' is stated without comment; since Ω is also an additive function satisfying the relevant Billingsley conditions, a one-sentence justification would make the remark self-contained.
Circularity Check
No significant circularity: the central CLT and optimization results derive from external theorems (Billingsley, Mertens, Hardy-Ramanujan, Landau, Erdős-Pomerance); the Section 7 training is empirical and transparently evaluated.
full rationale
The derivation chain starts from Billingsley's Erdős-Kac theorem and Mertens' estimates, then applies standard delta-method/Slutsky arguments (Theorems 2.2, 1.1, 3.1). The variance-stabilizing choice λ=1/2 and the width-optimal λ=3/4 are obtained by algebraically minimizing the asymptotic interval width derived from Theorem 3.1, not by fitting to data. The score interval is obtained by inverting the asymptotic normal pivot in Corollary 3.2.2, following Politis (2024), an external source. The Poisson interval is calibrated to the Landau asymptotic formula (16), and its performance is then measured by fuzzy coverage probabilities computed from data; this measurement is not forced by the calibration. The Section 7 training fits f̂_{μ,λ} and f̂_{σ,λ} to m∈[10^4,10^6] and then evaluates fuzzy coverage for m∈[10^5,10^14]; although some evaluation points lie inside the training range, the paper explicitly distinguishes in-sample from out-of-sample performance, and the fitted parameters do not directly set the empirical coverage to the nominal level. The only self-citation (Noguchi and Ward 2024) is a supporting reference for the standard square-root variance-stabilizing property and plays no load-bearing role. Concerns that the fitted standard deviations are miscalibrated relative to the Erdős-Kac variance are empirical correctness risks, not circularity.
Assumptions & free parameters
free parameters (6)
- beta_{0,lambda}, beta_{1,lambda} (mean fit) =
lambda=1/2: (0.4300, 0.9152); lambda=3/4: (0.4284, 0.9343); lambda=1: (0.4270, 0.9527)
- gamma_{0,lambda}, gamma_{1,lambda} (standard deviation fit) =
lambda=1/2: (-0.0109, 0.1498); lambda=3/4: (0.0196, 0.2695); lambda=1: (0.0499, 0.3677)
- eta_0, eta_1 (trained score interval variance fit) =
(-0.1136, 0.3855)
- Window j and grid increment =
j=2000, increment=1000, training range [10^4,10^6]
- Power selection q_mu, q_sigma =
q_mu=1, q_sigma=1/lambda
- kappa_alpha(m) (Poisson interval half-width) =
varies with m
assumptions (6)
- domain assumption Billingsley's Theorem 2.1 and Theorem 3.1 on additive functions
- standard math Mertens' second theorem
- domain assumption Hardy-Ramanujan theorem on the normal order of omega
- domain assumption Landau's asymptotic formula for rho_d(m)
- domain assumption Erdős-Pomerance theorem
- standard math Slutsky's theorem and the delta method under P_n
invented entities (1)
-
Local adjustment functions f_mu,lambda and f_sigma,lambda
independent evidence
Cite this review
Pith. "Pith review of Transforming the Erd\H{o}s-Kac theorem." pith.science (2026). https://pith.science/paper/U2GZED6T
@misc{pith2026250608503,
author = {Pith},
title = {Pith review of: Transforming the Erd\Hos-Kac theorem},
year = {2026},
howpublished = {\url{https://pith.science/paper/U2GZED6T}},
note = {Machine review of arXiv:2506.08503}
}
read the original abstract
Transforming the Erd\H{o}s-Kac theorem provides more flexibility in how the theorem can be utilized as an interval estimate for the prime omega function, which counts the number of distinct prime divisors. Here, we consider a direct transformation by the delta method. Then, we demonstrate that the square-root and three-quarters power asymptotically achieve variance stabilization and an optimal width, respectively. Furthermore, by adjusting the denominator of the theorem, we derive the score interval estimate. To make these interval estimates reliable for small positive integers, we examine performances of various interval estimates for the prime omega function using fuzzy coverage probabilities. The results indicate that the score interval estimate performs well even for small positive integers after training the mean and standard deviation using the prime omega function. Moreover, the Poisson interval estimate is relatively reliable with or without training. Additional theoretical results on the Erd\H{o}s-Pomerance theorem are also provided.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in ":" * " " * FUNCTION f...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in ":" * " " * FUNCTION f...
-
[3]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in ":" * " " * FUNCTION f...
-
[4]
Bera, A. K. and Koley, M. (2023). A history of the delta method and some new results. Sankhya B , 85(2):272--306
work page 2023
-
[5]
Billingsley, P. (1969). On the central limit theorem for the prime divisor function. The American Mathematical Monthly , 76(2):132--139
work page 1969
-
[6]
Billingsley, P. (1974). The probability theory of additive arithmetic functions. The Annals of Probability , 2(5):749--791
work page 1974
-
[7]
Box, G. E. P. and Cox, D. R. (1964). An analysis of transformations. Journal of the Royal Statistical Society Series B: Statistical Methodology , 26(2):211--243
work page 1964
-
[8]
Bradley, J. V. (1978). Robustness? British Journal of Mathematical and Statistical Psychology , 31(2):144--152
work page 1978
Show all 26 references
-
[9]
Elliott, P. D. T. A. (1980). Probabilistic Number Theory II: Central Limit Theorems . Springer-Verlag, New York, NY
1980
-
[10]
and Kac, M
Erd o s, P. and Kac, M. (1939). On the G aussian law of errors in the theory of additive functions. Proceedings of the National Academy of Sciences , 25(4):206--207
1939
-
[11]
and Kac, M
Erd o s, P. and Kac, M. (1940). The G aussian law of errors in the theory of additive number theoretic functions. American Journal of Mathematics , 62(1):738--742
1940
-
[12]
and Pomerance, C
Erd o s, P. and Pomerance, C. (1985). On the normal number of prime factors of (n). The Rocky Mountain Journal of Mathematics , 15(2):343--352
1985
-
[13]
Geyer, C. J. and Meeden, G. D. (2005). Fuzzy and randomized confidence intervals and p-values. Statistical Science , 20(4):358--366
2005
-
[14]
Hardy, G. H. and Ramanujan, S. (1917). On the normal number of prime factors of a number n. Quarterly Journal of Mathematics , 48:76--92
1917
-
[15]
Harper, A. J. (2009). Two new proofs of the E rd \"o s-- K ac T heorem, with bound on the rate of convergence, by S tein's method for distributional approximations. Mathematical Proceedings of the Cambridge Philosophical Society , 147(1):95--114
2009
-
[16]
Kowalski, E. (2021). An Introduction to Probabilistic Number Theory , volume 192. Cambridge University Press, Cambridge, United Kingdom
2021
-
[17]
Landau, E. (1900). Sur quelques probl \`e mes relatifs \`a la distribution des nombres premiers. Bulletin de la Soci \'e t \'e Math \'e matique de France , 28:25--38
1900
-
[18]
Loyd, K. (2023). A dynamical approach to the asymptotic behavior of the sequence. Ergodic Theory And Dynamical Systems , 43(11):3685--3706
2023
-
[19]
and Ward, M
Noguchi, K. and Ward, M. C. (2024). Asymptotic optimality of the square-root transformation on the gamma distribution using the Kullback--Leibler information number criterion. Statistics & Probability Letters , 210:110118
2024
-
[20]
Politis, D. N. (2024). Studentization versus variance stabilization: A simple way out of an old dilemma. Statistical Science , 39(3):409--427
2024
-
[21]
Ramsey, P. H. (1980). Exact type 1 error rates for robustness of S tudent's t test with unequal variances. Journal of Educational Statistics , 5(4):337--349
1980
-
[22]
Rudin, W. (1976). Principles of Mathematical Analysis . McGraw-Hill, New York, NY, third edition
1976
-
[23]
Serfling, R. J. (2009). Approximation Theorems of Mathematical Statistics . John Wiley & Sons, New York, NY
2009
-
[24]
Tenenbaum, G. (2015). Introduction to Analytic and Probabilistic Number Theory , volume 163. American Mathematical Society, Providence, RI, third edition
2015
-
[25]
Wilson, E. B. (1927). Probable inference, the law of succession, and statistical inference. Journal of the American Statistical Association , 22(158):209--212
1927
-
[26]
Yu, G. (2009). Variance stabilizing transformations of P oisson, binomial and negative binomial distributions. Statistics & Probability Letters , 79(14):1621--1629
2009
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.