Pith. sign in

REVIEW 4 major objections 6 minor 11 references

Arithmetic Bias in the Distribution of Mersenne Prime Exponents and the Divisor Structure of p-1

T0 review · 4 major / 6 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read Known Mersenne prime exponents carry a secondary arithmetic signal: divisor-rich p-1 appears more often than exponent size alone predicts.

desk verdict The S(p) enrichment is plausibly real but confounded by the known p≡1 mod 4 bias; the final law is a curve fit, not a prediction—worth refereeing, but only as a conditional accept. read the letter →

arxiv 2603.08994 v5 pith:AVTQ263Y submitted 2026-03-09 math.NT

classification math.NT MSC 11A4111N0511Y1162G10
keywords Mersenneprimesdivisorfunctioncyclotomicpolynomialsarithmeticbiasprimedistributionprobabilisticnumbertheoryexperimental
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that the distribution of known Mersenne prime exponents is not fully explained by exponent size. It introduces S(p) = log τ(p-1) / log log p and shows that known exponents, excluding the smallest cases, sit systematically higher in S(p) than nearby prime controls of comparable size. To explain this, the paper builds a heuristic model in which divisors of p-1 create cyclotomic layers that filter candidate factorizations of 2^p − 1, leading to a refined probability law P(M_p prime | S) ≈ C (log p)^{S(p)} / p. If correct, the classical size-only heuristic remains the asymptotic law, but a finite-scale arithmetic bias tied to the divisor structure of p-1 is real and measurable.

What carries the argument

The normalized divisor parameter S(p) = log τ(p-1) / log log p; the cyclotomic identity 2^{p-1} − 1 = ∏_{d|p-1} Φ_d(2); the two-factor diagonal residue classes that concentrate all prime factors of M_p into one residue class modulo a; and the empirical scaling ω_{<p}(A_p) ≈ k τ(p-1) with k ≈ 0.684, which converts divisor count into effective filter count.

What would settle it

Record the S(p) values of the next Mersenne prime exponents discovered beyond 136,279,841; if they are not predominantly above the local median of S among nearby primes, the bias claim fails. Alternatively, partially factor A_p = (2^{p-1} − 1)/p for p in 10^6–10^7 and check whether ω_{<p}(A_p)/τ(p-1) stays near 0.684; a sharp drop would invalidate the layer-counting step.

Watch

Extended reading notes

Core claim

On the paper's own terms: Mersenne prime exponents have elevated divisor complexity of p-1 relative to local prime controls, and this elevation survives matching by exponent size. The mean percentile rank of S(p) in a 5000-prime window is 0.648 versus a null expectation of 0.5; sign, Wilcoxon, and permutation tests reject the null with p-values around 0.002. The paper interprets this as evidence for a structural refinement of the standard probability estimate for M_p being prime, preserving the log p / p scale after averaging over S.

Load-bearing premise

The model stands on the premise that the number of cyclotomic layers that actually constrain factorizations is proportional to τ(p-1) (L ≈ k τ(p-1)), a proportionality measured only below p ≈ 10^6 and assumed to continue at larger scales.

Editorial extensions

If this is right

  • If the bias is real, size-only estimates under-predict the probability for exponents whose p-1 has many divisors and over-predict for sparse ones, without changing the average behavior.
  • Search strategies for new Mersenne primes could rank candidate exponents by S(p) within a given size range.
  • The model yields testable predictions: the split of new exponents between S(p) > 1 and S(p) < 1 should remain skewed rather than approach 50/50.
  • The empirical k ≈ 0.684 relation gives a direct way to estimate the number of effective modular filters from τ(p-1) alone.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The finite sample of 48 known exponents makes the statistical signal fragile; the strongest confirmation would come from the next discovered exponent, not from reanalysis of the same data.
  • If the bias persists, S(p) becomes a cheap prefilter for computational searches; if it does not, the effect is likely a selection artifact of how the known set was assembled.
  • The k ≈ log 2 saturation argument suggests the layer-density coefficient may be bounded by a capacity effect, which would make the bias weaker or saturating at very large p.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This paper reports an empirical bias in the divisor structure of p−1 for known Mersenne prime exponents. It defines S(p)=log τ(p−1)/log log p and compares S(p) for the 48 known exponents p≥13 against nearby prime controls. The authors report moderate enrichment (mean percentile ~0.65, Cohen's d≈0.56) with p-values ≈0.001–0.002 across sign, Wilcoxon, KS, conditional logistic, and permutation tests. To interpret this, they develop a heuristic model based on cyclotomic divisors of 2^{p−1}−1, aggregating congruence constraints to propose P(M_p prime|S)≈C (log p)^{S(p)}/p, which they claim reproduces the observed S>1 versus S<1 imbalance while preserving the Wagstaff scale. The paper is explicitly framed as heuristic and finite-scale.

Significance. Strengths: the statistical analyses are carefully described and reproducible, using distribution-free procedures and stratified permutation tests; the paper is honest about limitations. If, after proper conditioning, the enrichment survives, it would be an interesting secondary arithmetic signal in the distribution of Mersenne exponents. Weaknesses: the central empirical claim is currently confounded by the known p≡1 mod 4 bias, which mechanically inflates S(p); the model is fitted on the same data and extrapolated beyond its measured range. The paper's contribution therefore hinges on whether the enrichment persists within residue classes.

major comments (4)
  1. [Appendix A.1–A.4 and §3] The control sets are not stratified by p mod 4. For p≡3 mod 4, p−1=2·odd, so τ(p−1)=2τ(odd); for p≡1 mod 4, p−1 has v2≥2, so τ(p−1)≥3τ(odd′). Hence S(p)=logτ/loglogp is systematically larger for p≡1 mod 4. Table 1 shows a 31/17 bias of the known Mersenne exponents toward p≡1 mod 4 (the Wagstaff residue effect recalled in §4.1). Since the controls in A.1 include both residues, the reported sign test (p=0.0021), Wilcoxon (p=0.0014), KS (p=0.00046), and permutation (p=0.0012) may reflect this known residue effect rather than a new divisor-structure signal. The analyses must be rerun with controls restricted to the same residue class (and/or with p mod 4 as a covariate in the conditional logistic model).
  2. [Appendix A.6–A.7] The key scaling L(p)≈kτ(p−1) is measured only for p<10^6. The entries for p∼10^7 and 10^8 (k≈0.684, θ≈0.82) are explicitly described as 'finite-scale projections' rather than direct factor-enumeration results. Nevertheless, Table A.7 uses these projected values to generate predicted counts for p<10^7 and p<10^8, including the claimed agreement with the observed 28/6 and 37/10 splits. Because this extrapolation is load-bearing for the largest-range claims, the authors should either restrict the model to the directly computed range or supply independent evidence for the proportionality beyond p=10^6.
  3. [§5.4 and A.7] The proposed law P∝(logp)^S/p is introduced after the empirical enrichment is observed, and its parameters (k, w_2, a_trunc, ε) are calibrated on the same Mersenne-prime data. The matched 'prediction' in A.7 is therefore an in-sample consistency check, not a falsifiable out-of-sample prediction. To claim predictive power, all parameters should be fixed using exponents below a threshold (or a random subset) and then evaluated on held-out exponents. As it stands, the agreement between the model and the observed S>1/S<1 split is expected by construction.
  4. [§4.2–§5.5] The refined law P(M_p prime|S)≈C(logp)^S/p omits the mod-4 factor that Wagstaff's refinement (§4.1) requires. Since S(p) is mechanically higher for p≡1 mod 4, the model may simply be re-encoding the established residue bias as a 'structural' effect. The paper should show that the S-effect is not a proxy for mod 4, e.g., by demonstrating the enrichment within each residue class, or by including both S(p) and p mod 4 in the conditional model and reporting the partial effect of S.
minor comments (6)
  1. [General] No equation numbers are used, which makes precise reference cumbersome; adding numbers would improve readability.
  2. [§2.4] The statement that log τ(p−1) 'fluctuates on a scale comparable to log log p' is informal; a reference to known distribution results for τ(n) would help.
  3. [§5.3] The upper limit t in the sum ∑_{n=2}^t is never defined; it should be specified as the maximum factorization length.
  4. [Table 1] The caption calls π_{1000} and π_{5000} 'percentile ranks' but the values are proportions in [0,1]; clarify the intended scale.
  5. [A.1] The sentence 'Exponents below p=13 are excluded' could explicitly state that the four excluded Mersenne exponents are 2, 3, 5, and 7.
  6. [A.6] The regression table presents 'Pearson r, Spearman ρ, k 95%CI' in a single column; separate columns would be clearer.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the structural model's parameters are estimated from auxiliary arithmetic data, and the A.7 prediction is a genuine out-of-sample consistency check.

full rationale

The paper's derivation chain is: (1) empirically demonstrate elevated S(p) among Mersenne exponents; (2) propose a cyclotomic-layer mechanism; (3) calibrate the layer-count relation ω<p(Ap)≈kτ(p−1) on all primes, not on the Mersenne response; (4) derive P∝(logp)^S/p; (5) test the derived model against the observed Mersenne split. The fitted coefficient k is estimated from a dataset that does not include the Mersenne-prime indicator, so the model's prediction of the G+/G− imbalance is not forced by the response variable. Although the final law is algebraically equivalent to P∝τ(p−1)/p (since (logp)^S=τ(p−1)), this is a definitional identity, not a circular input. The model's parameters are not fitted to the Mersenne data; the observed exponents are used only in the final comparison stage (Appendix A.7 explicitly states this). The paper also disclaims first-principles derivation and acknowledges the heuristic, finite-scale status of the framework. The main unresolved concern is the potential confounding of S(p) with the known p≡1 mod 4 bias, but that is a statistical validity issue rather than a circularity of derivation.

Assumptions & free parameters 4 free parameters · 9 assumptions · 1 invented entities

The paper's central empirical finding is independent of these inputs, but its explanatory model leans on several fitted or ad hoc ingredients, most importantly k_eff and the unproved filtering mechanism.

free parameters (4)
  • k_eff(p) (layer-density proportionality) = 0.73–0.94 for a_trunc=10^6; 0.684 for a_trunc=p; projected to 10^8
    Fitted by regression ω<atrunc(Ap)≈k τ(p−1) in A.6; the final model's S-dependence enters through this relation.
  • w_2(p) (two-factor pattern weight) = unmodeled, assumed ≪1
    Introduced in §5.3 to keep the exponential structural aggregate subprincipal; its value is not derived or fitted but affects normalization.
  • a_trunc (modulus truncation) = 10^6 or p
    Chosen visibility cutoff in A.6; changes k and the set of active moduli.
  • ε (visibility threshold) = ~10^-6
    Heuristic threshold tying a_trunc to activation probability in A.6.
assumptions (9)
  • standard math Zsigmondy's theorem: for d>1, 2^d−1 has a primitive prime divisor except d=6
    Used in §2.3 to argue each cyclotomic layer contributes at least one modulus.
  • standard math Erdős–Kac theorem on normal order of ω(n)
    Used in §2.4 to justify normalizing log τ(p−1) by log log p.
  • standard math Dirichlet's theorem on primes in arithmetic progressions
    Used in §2.5 to assert equidistribution of factors across residue classes.
  • standard math Mertens' theorem
    Used in §4.1 to derive the logarithmic factor in the Wagstaff sieve.
  • domain assumption Wagstaff/Bateman–Horn heuristic as the true baseline
    The refinement is anchored to the unproved baseline P∼C log p/p.
  • ad hoc to paper Cyclotomic layers act as independent multiplicative filters on composite factorizations
    §4.2–5.4: the aggregation model assumes layer contributions combine multiplicatively and that diagonal congruence classes reduce composite density; no proof or quantitative derivation is given.
  • ad hoc to paper Effective modulus count L(p)≈k τ(p−1) holds beyond p<10^6
    A.6 fits k up to 10^6 and projects k≈0.684 for p up to 10^8; table marks these with an asterisk.
  • domain assumption Known 52 Mersenne exponents are a representative sample
    All inference uses the currently known finite list; selection effects from GIMPS search history are not modeled.
  • ad hoc to paper Typical regime S(p)≈1
    Assumed in §4.2 and §6.1; by the cited Erdős–Kac normal order, log τ(n)∼log2 log log n, so asymptotically S→log2≈0.693, not 1.
invented entities (1)
  • Effective cyclotomic modulus (set A_p^eff)
    purpose: Quantifies how many cyclotomic layers contribute observable congruence constraints below a truncation scale
    Defined via arbitrary truncation and the fitted relation ω<atrunc≈kτ; no independent handle beyond the calibration data.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Arithmetic Bias in the Distribution of Mersenne Prime Exponents and the Divisor Structure of p-1." pith.science (2026). https://pith.science/paper/AVTQ263Y

@misc{pith2026260308994,
  author       = {Pith},
  title        = {Pith review of: Arithmetic Bias in the Distribution of Mersenne Prime Exponents and the Divisor Structure of p-1},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AVTQ263Y}},
  note         = {Machine review of arXiv:2603.08994}
}
abstract

According to the classical Wagstaff heuristic, the probability that a Mersenne number $M_p=2^p-1$ is prime depends primarily on the size of the exponent $p$. We investigate whether the divisor structure of $p-1$ produces detectable secondary variation within this aggregate probability scale. We introduce the normalized divisor parameter $S(p)=\frac{\log\tau(p-1)}{\log\log p}$, which provides a scale-adjusted measure of the divisor complexity of $p-1$. Using the currently known Mersenne prime exponents, excluding the smallest cases, we compare $S(p)$ against nearby prime controls of comparable size. Across several complementary statistical analyses, Mersenne prime exponents exhibit elevated values of $S(p)$. To interpret this empirical bias, we develop a reduced coarse-grained structural model based on the cyclotomic decomposition $2^{p-1}-1=\prod_{d\mid(p-1)}\Phi_d(2)$. The divisor structure of \(p-1\) generates cyclotomic layers associated with modular constraints on candidate factorizations. Their cumulative filtering effect motivates a structural refinement of the classical Wagstaff heuristic of the form $\mathbb P(M_p\ {\rm prime}\mid S)\approx C(p,S)\frac{\log p}{p}$, where $C(p,S)$ denotes the finite-scale structural factor. The resulting model predicts a redistribution toward higher values of $S(p)$, consistent with the observed imbalance across the explored exponent ranges, while preserving the aggregate Wagstaff probability scale after marginalization over $S$. The proposed framework is heuristic and finite-scale, and is intended as a possible structural interpretation of the observed arithmetic bias rather than as a derivation from first principles or a modification of the classical Wagstaff asymptotic scale.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

11 extracted references

  1. [1]

    P. T. Bateman and R. A. Horn, A heuristic asymptotic formula concerning the distribution of prime numbers, Mathematics of Computation 16 (1962), 363–367

  2. [2]

    S. S. Wagstaff Jr., The distribution of Mersenne primes, Mathematics of Computation 39 (1982), 397–403

  3. [3]

    Zsigmondy, Zur Theorie der Potenzreste, Monatshefte für Mathematik und Physik 3 (1892), 265–284

    K. Zsigmondy, Zur Theorie der Potenzreste, Monatshefte für Mathematik und Physik 3 (1892), 265–284

  4. [4]

    Ireland and M

    K. Ireland and M. Rosen, A Classical Introduction to Modern Number Theory, Graduate Texts in Mathematics, Vol. 84, Springer, 1990

  5. [5]

    G. H. Hardy and E. M. Wright, An Introduction to the Theory of Numbers, 6th ed., Oxford University Press, 2008

  6. [6]

    Tenenbaum, Introduction to Analytic and Probabilistic Number Theory, 3rd ed., Graduate Studies in Mathematics, Vol

    G. Tenenbaum, Introduction to Analytic and Probabilistic Number Theory, 3rd ed., Graduate Studies in Mathematics, Vol. 163, American Mathematical Society, 2015

  7. [7]

    Crandall and C

    R. Crandall and C. Pomerance, Prime Numbers: A Computational Perspective, 2nd ed., Springer, 2005

  8. [8]

    Erd o s and M

    P. Erd o s and M. Kac, The Gaussian law of errors in the theory of additive number theoretic functions, American Journal of Mathematics 62 (1940), no. 1, 738--742

Show all 11 references
  1. [9]

    D. S. Dummit and R. M. Foote, Abstract Algebra, 3rd ed., Wiley, 2004

  2. [10]

    McFadden, Conditional logit analysis of qualitative choice behavior, in Frontiers in Econometrics, Academic Press, 1974, pp

    D. McFadden, Conditional logit analysis of qualitative choice behavior, in Frontiers in Econometrics, Academic Press, 1974, pp. 105–142

  3. [11]

    The Great Internet Mersenne Prime Search (GIMPS), Mersenne Prime Database, https://www.mersenne.org, accessed February 2026

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.