{"id":"091776f7-6089-4cc3-92e8-55368d9e8820","arxiv_id":"2608.00615","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Statistical analysis of 250,000 Boyd Mahler-measure constants reveals p-adic valuation laws (P(v_p = r) ≈ p^{-r} for p ≥ 5) and a lower-bound conjecture at p = 2, with a conditional proof on a subfamily.","lead":"Boyd conjectured that Mahler measures of a family of two-variable polynomials are rational multiples of L-function derivatives. This paper computes 250,000 examples, finds empirical laws for the 2-adic, 3-adic, and p≥5 valuations of those rationals, and tests machine-learning models on them.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"p=5 already rejects the headline p≥5 law: Table 3.1 gives χ²/ν=12.8, so the strongest empirical claim is false as stated; a local-conditioned refinement is needed.","rationale":"I read the paper as an honest, data-rich heuristic study: the authors repeatedly label their assertions as empirical, report near-zero exact ML prediction, and give a rigorous Lemma 3.1 plus a conditional Proposition 3.4. Those parts deserve credit. However, the single most load-bearing assertion in the abstract and in the reader's strongest_claim is the p≥5 valuation law. The paper's own Table 3.1 rejects it at p=5: χ²=76.84 with 6 degrees of freedom is not a small deviation, and Table 3.1b shows the excess is spread across several positive valuations. The paper's wording 'strong deviation' is accurate, but the abstract still says the law holds 'for p≥5' without qualification. This is an internal-data inconsistency, not a disagreement with external consensus, so it should be treated as a correctness risk. The reader's chosen weakest_assumption, the 19-digit LLL rational recognition, is a legitimate but secondary risk: for the observed range n_k is likely no larger than about 10^15–10^16, so 19 digits should generally be sufficient, and a sample recomputation would likely confirm this. The p=5 failure, by contrast, is already visible in the published tables. I would not reject the paper: the 2-adic and 3-adic structures, the conditional Bloch–Kato result, and the ML experiments are separate contributions that survive even if the p≥5 law needs a local-conditioned refinement. But the central empirical claim should be revised before acceptance, either by conditioning on bad-reduction classes or by explicitly excluding p=5 from the law. This keeps the CONDITIONAL verdict; hence UNCHANGED relative to the reader's assessment.","tokens_in":30830,"tokens_out":15966,"duration_ms":223224,"concrete_test":"For all k≤250,000, partition the v_5(n_k) data by the three local residue classes: (a) 5 | k, (b) k ≡ ±2 mod 5 (i.e. 5 | k²−16), and (c) 5 ∤ k(k²−16). Compute the observed counts for v_5=r in each class and compare with M·5^{-r} (or with the appropriate conditional model) using a χ² test. If the good-reduction class (c) fits 5^{-r} while (a)+(b) do not, the unconditional p≥5 law must be replaced by a local-conditioned version; if no class fits, the p≥5 law as stated is empirically false. As a secondary check, recompute a random sample of the 250,000 rational recognitions at 40+ digit precision to confirm the valuations themselves are not an artifact of the 19-digit LLL step.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central statistical claim (abstract, §3.2, and the reader's strongest_claim) is that for all p≥5, P(v_p(n_k)=r)=p^{-r} for r≥1. The paper's own Table 3.1 disproves this at p=5: χ²=76.84 with ν=6, i.e. χ²/ν=12.81, an overwhelming rejection with N=249,999. Table 3.1b shows the deviations are systematic rather than a tail artifact: v_5=2,3,4,6 are all 6–19% above the p^{-r} prediction. Since p=5 is the prime with by far the most data, the unqualified 'p≥5' law is not merely unproved; it is contradicted by the paper's main dataset. The paper honestly flags the deviation in the text, but the abstract and the stated central claim retain the unqualified law. Moreover, the paper's own Bloch–Kato discussion (Remark 3.5) predicts v_p(n_k) should depend on whether p divides the conductor, so an unconditional p^{-r} law is surprising and needs to be checked against local conditions such as p | k or p | k²−16. If the p=5 deviation concentrates in bad-reduction classes, the law as formulated is incomplete; if it persists in good-reduction classes, the law is simply false at p=5. Either way, the strongest empirical claim requires revision or a precise exclusion of p=5.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper computes Boyd Mahler measure quotients r_k (equivalently n_k = 1/r_k) for the first 250,000 integer k, combining PARI/GP numerical computation with LLL rational recognition, statistical analysis, and transformer experiments on the Axolver/Int2Int platform. The main statistical claims are: the sign of r_k is governed by the root number; |n_k| has order N_k/log(k) via Lemma 3.1; for p≥5 the positive p-adic valuations of n_k follow P(v_p(n_k)=r)=p^{-r}; the 3-adic valuations obey congruence and parity laws on large subfamilies; and the 2-adic valuations are modeled by a weighted count of odd primes dividing k and k^2-16, plus corrections c(k), s(k), and a negative-binomial residual. A conditional proof of the 2-adic lower-bound conjecture (Conjecture 3.2) is given for a special family under Bloch-Kato-type assumptions from [DGdJK26]. The machine-learning experiments show that the model can learn the size of n_k from the conductor, partially learn v_3 and v_2, but collapses to the trivial prediction for p=5,7.","tokens_in":31152,"tokens_out":5943,"duration_ms":75620,"significance":"If the empirical regularities survive scrutiny, this is a substantial contribution to the arithmetic of Boyd's Mahler measure conjectures: it proposes concrete, locally structured laws for the rational factors r_k and connects them to conductors, bad reduction, and Bloch-Kato Selmer groups. The paper's strengths are its scale, the explicit chi-square diagnostics, the public data/code, and the honest reporting of the neural-network limitations (e.g., §5.2). Lemma 3.1 is clean and the conditional Proposition 3.4 is a serious attempt to connect the data to the Bloch-Kato framework. However, the headline p≥5 law is already contradicted by the paper's own Table 3.1 at p=5, and the rational-recognition step at only 19 decimal digits is not validated for the largest conductors. Both points are load-bearing for the empirical claims and require correction before the paper's central assertions can be accepted.","major_comments":[{"comment":"The unqualified law P(v_p(n_k)=r)=p^{-r} for all p≥5 is contradicted by the paper's own data at p=5. Table 3.1a reports χ²=76.84 with ν=6, i.e. χ²/ν=12.81, and Table 3.1b shows systematic excesses at r=2,3,4,6 (ratios 1.065, 1.071, 1.178, 1.188). Since p=5 has by far the largest sample in the table, this is not a tail artifact. The abstract, §3.2, and §6 all retain the unqualified p≥5 statement. The paper must either state the law only for p≥7, or formulate and test a p=5 refinement, e.g. conditioning on p|k, p|(k²−16), or the reduction type of E_k. As written, the headline empirical claim is false at p=5.","section":"§3.2, Table 3.1; Abstract"},{"comment":"All valuation statistics in §3 inherit the correctness of the LLL rational recognition of r_k from 19 decimal digits. No higher-precision check is reported. For large k, n_k is of size ≍ N_k/log(k), and the corresponding rational r_k can have large numerator and denominator; 19 digits may not be sufficient to determine such a rational uniquely. A misidentified r_k would silently bias every empirical law in Section 3. Please recompute a random sample of values (especially large k and large conductors) with, say, 50–100 digits and report the fraction of r_k that change, or provide a denominator bound justifying 19 digits.","section":"§2, numerical recognition"},{"comment":"The statistical models in §3.3–3.4 are fitted and evaluated on the same 250,000 values: the weights W_q(k), corrections c(k) and s(k), the negative-binomial parameters, the mode shift in (3.12), and the exceptional sets in Table 3.5 are all estimated from the full dataset. No out-of-sample split or holdout validation is reported. Moreover, Figure 3.4 shows χ²/ν=7.9 for the final 2-adic model, which for N≈250,000 is a poor absolute fit despite the similar visual shape. Please add a validation protocol (e.g., fit on one half, test on the other), report out-of-sample χ² for the valuation distributions, and quantify the NB goodness-of-fit rather than relying on visual agreement.","section":"§3.3–3.4, Table 3.5, Figure 3.4"}],"minor_comments":[{"comment":"The caption reads 'χ 2/ν=7.9' with broken typography; define χ² and ν explicitly at that point, since the main definition appears only in §3.2.","section":"Figure 3.4 caption"},{"comment":"Law (3.6) is stated to hold on A 'with the exception of k=1', but k=1 belongs to A; please clarify whether the exception is the single k=1 or a different value, and make the statement in (3.6) consistent with the table.","section":"§3.3.1, Table 3.3"},{"comment":"The text says the model learned fastest with N_k, 'closely followed by factor(Δk) and closely followed by factor(Nk)'; the figure legend lists the factored conductor peak at 99.5% and factored discriminant at 99.7%. Please align the text with the plotted peaks.","section":"§5.1, Figure 5.2"},{"comment":"The validation set is described as not used for hyperparameter tuning, but model selection across epochs is performed on validation accuracy. Please note this explicitly, as it is a form of model selection.","section":"§4"},{"comment":"Several displayed equations have broken math spacing (e.g., 'N_k/4π^2 L(E_k,2)' in (3.3), and 'NB(3,3/5)' vs 'NB(r=3, p=3/5)'). A careful proofread of the LaTeX rendering would improve clarity.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is a serious large-scale empirical study with a valuable data release, and the conditional Proposition 3.4 is a useful bridge to the Bloch-Kato literature. The main obstacle is not the ML part but the validity of the statistical headline: the authors' own p=5 chi-square test rejects the central p≥5 law, and the 19-digit rational recognition is not checked for large conductors. I would be willing to reconsider after a revision that (i) re-formulates or refines the p≥5 law, (ii) validates the rational recognition at higher precision on a sample, and (iii) adds out-of-sample validation for the fitted models. I do not see a reason to reject outright, since these are fixable within the scope of the manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a data-rich empirical study with one honest theorem and a lot of plausible numerology. The 250,000 values of r_k is a real step beyond Boyd's 40, and Lemma 3.1 is a clean rigorous bound that gives the order-of-magnitude estimate |n_k| ≍ N_k/log k. The 2-adic model with weighted bad-prime counts, the s(k) correction near ±8, and the negative-binomial residual is a well-specified empirical description, and Conjecture 3.2 is concrete enough to be tested further. The 3-adic parity laws on the sets A, B, C are genuinely new and show structure that comes from local reduction data. Proposition 3.4 is a real conditional proof, modest but solid: it derives Conjecture 3.2 on a restricted family from the work of Dummigan–Golyshev–de Jeu–Kerr.\n\nThe soft spots are significant. Most importantly, the abstract and conclusion state an unqualified 'p ≥ 5' law P(v_p(n_k)=r) = p^{-r}, but the paper's own Table 3.1a gives χ²/ν = 12.81 for p=5 with N ≈ 250,000. That is an overwhelming rejection, and Table 3.1b shows systematic excess at v_5 = 2,3,4,6. The paper texts flags this, but the central claim should be revised: either exclude p=5 or refine by local conditions, e.g. whether p divides k or k²−16. The Bloch–Kato formula in Remark 3.5 suggests such local dependence is expected, so an unconditional law is surprising on its face.\n\nSecond, the numerical rational recognition: both µ_k and L'(E_k,0) are computed to 19 decimal digits and then r_k recovered via LLL. As k and the conductor grow, the rational r_k can have large numerator/denominator, and 19 digits may not be enough. The authors report no higher-precision cross-check. If even a small fraction of r_k are misidentified, all valuation statistics inherit that noise. This needs to be addressed directly.\n\nThird, the statistical models are fit and evaluated on the same dataset. The √m nonsquare check and the larger-k spot checks give some independent evidence, but the fitted weights and distributions are not validated out of sample. The ML experiments are honestly reported—near-zero exact prediction, trivial collapse at p=5,7—but the claim that the p=2 network learns 'beyond' the explicit predictor is not established; the correlation drop and accuracy gain could be overfitting.\n\nWho is this for? Computational number theorists and anyone working on Mahler measures, Boyd's conjectures, or empirical arithmetic statistics. It deserves peer review, not desk rejection, but with major revision: revise the p≥5 law, verify precision, temper the abstract. I'd cite it for the dataset and the 2-adic structure, cautiously.","headline":"Valuable dataset and some real conjectures, but the headline p≥5 law is contradicted by the paper's own p=5 chi-square, and the 19-digit rational recognition deserves a hard look before the statistics are trusted.","tokens_in":31799,"tokens_out":1962,"would_cite":true,"duration_ms":28864,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["11G05","11F67","11G40"],"pacs":[],"model":"deepseek-v4-flash","headline":"The rational factor in Boyd's Mahler-measure identities is arithmetically structured: for p≥5 its p-adic valuation obeys a geometric law, and at 2 and 3 it is tied to the elliptic curve's bad primes.","keywords":["Mahler measure","Boyd conjecture","elliptic curves","L-function special values","p-adic valuations","Bloch–Kato conjecture","arithmetic statistics","machine learning"],"falsifier":"Recompute $\\mu_k$ and $L'(E_k,0)$ for 500 randomly chosen $k$ near 250,000 with 100-digit precision and rerun rational reconstruction; if any recovered $r_k$ differs from the 19-digit value, or if the reconstruction failure rate grows with the conductor, the empirical p-adic laws collapse. A complementary check: tally $v_5(n_k)$ on a fresh 250,000-value window, since the reported $\\chi^2/\\nu\\approx 12.8$ for $p=5$ already flags this prime as the outlier.","tokens_in":30634,"feed_emoji":"🧮","tokens_out":9459,"duration_ms":100894,"temperature":0.7,"texified_at":"2026-08-05T21:58:53.456633+00:00","pith_summary":"This paper tries to show that the rational factor $r_k$ in Boyd's conjectural formula—Mahler measure of a one-parameter family equals $r_k$ times the $L'$-value of an associated elliptic curve—is not an arbitrary rational number but carries systematic arithmetic structure. Working with the first 250,000 integers $k$, the authors find that the reciprocal $n_k = 1/r_k$ is almost always an integer, that its size is governed by the conductor, and that its p-adic valuations obey distinct laws for $p\\ge 5$, $p=3$, and $p=2$. For $p\\ge 5$ the data indicate $P(v_p(n_k)=r)\\approx p^{-r}$; for $p=3$ the valuation is tied to congruence classes of $k$ and the parity of a count of bad primes; for $p=2$ the valuation is approximately a weighted count of odd prime divisors of $k$ and $k^2-16$ plus a small negative-binomial residual. The paper states a precise lower-bound conjecture $\\hat{v}(k)\\le v_2(n_k)$ and proves it conditionally, under Bloch–Kato hypotheses, on an infinite subfamily.","texify_model":"deepseek-v4-flash","texify_usage":{"total_tokens":11512,"prompt_tokens":967,"completion_tokens":10545,"prompt_tokens_details":{"cached_tokens":0},"prompt_cache_hit_tokens":0,"prompt_cache_miss_tokens":967,"completion_tokens_details":{"reasoning_tokens":9545}},"feed_headline":"Mahler-measure ratios follow p-adic laws across 250,000 cases","feed_subtitle":"The rational factor linking Mahler measures to elliptic-curve L-values hides structured valuations at 2, 3, and beyond.","key_machinery":"The central object is the reciprocal $n_k = 1/r_k$ of Boyd's rational factor. The main working identity is the empirical decomposition at the prime 2: $v_2(n_k) = \\omega_{odd}(k) + 2\\omega_{odd}(k^2-16) - c(k) + s(k) + X_k$, where $\\omega_{odd}$ counts distinct odd prime divisors, $c(k)$ is 4 or 2 depending on whether $E_k$ is semistable or additive at 2, $s(k)$ is a deterministic correction that spikes when $k$ is 2-adically close to $\\pm 8$, and $X_k$ follows $NB(3,3/5)$. At the prime 3 the analogous mechanism is the parity law $v_3(n_k)=0 \\Rightarrow \\alpha(k)$ odd and $v_3(n_k)=1 \\Rightarrow \\alpha(k)$ even, with $\\alpha(k) = \\#S_k + v_3(k) + 1_{8|k}$. At primes $p\\ge 5$ the mechanism is the geometric law $P(v_p(n_k)=r)=p^{-r}$. The load-bearing bridge to arithmetic is the Bloch–Kato","core_discovery":"Boyd's conjecture asserts $\\mu_k = r_k L'(E_k,0)$ for an elliptic curve $E_k$. The paper's central discovery: for all but seven exceptional $k$ up to 250,000, $r_k$ is the reciprocal of an integer $n_k$ whose p-adic valuations are locally determined by $E_k$. For $p\\ge 5$, $P(v_p(n_k)=r)=p^{-r}$. For $p=3$, valuations obey congruence and parity laws tied to bad primes $\\equiv 1 \\pmod{3}$. For $p=2$, $v_2(n_k) \\approx \\omega_{odd}(k)+2\\omega_{odd}(k^2-16)-c(k)+s(k)+X_k$ with $X_k\\approx NB(3,3/5)$. The paper conjectures $\\hat{v}(k)\\le v_2(n_k)$ for $k\\ne 4,8$ and proves it conditionally, under Bloch–Kato hypotheses, on an infinite subfamily $k=4u$.","pith_inferences":["The p=5 χ² deviation, while small, is a signature that the p≥5 law may fail at small primes; a natural test is to recompute the p=5 tallies on a disjoint 250,000-value window or under the nonsquare k=√m regime.","If the residual X_k is genuinely NB(3,3/5) and independent of k, then the sharpness cases v_2(n_k)=\\hat v(k) should have density dictated by that distribution; the network's improved detection of those cases suggests a deterministic hidden feature—possibly a Selmer 2-rank or regulator index equal to 1—that a dedicated classifier could try to isolate.","The same valuation statistics could be tested on other one-parameter Mahler-measure families to see whether the 2-adic weighted count and negative-binomial residual are universal or special to this family."],"forward_implications":["If the p≥5 law holds, an integer k has n_k divisible by p with probability 1/(p−1), not 1/p, and higher powers of p then follow the classical geometric distribution.","If Conjecture 3.2 is true, v_2(n_k) is always at least the weighted count of odd bad primes minus a bounded 2-adic correction; in particular its average order is at least 5 log log k.","Under Bloch–Kato, valuations of n_k carry information about Tate-twisted Selmer groups and regulator indices; conditional verification on the k=4u family shows Mahler-measure data can certify lower bounds on Selmer 2-ranks.","The near-integrality statement (n_k integer except k∈{3,4,5,8,12,16,32}) is recovered and sharpened, and the size law |n_k| ≍ N_k/log k gives a first-order prediction of the magnitude of n_k."],"supporting_citations":[{"why":"States the conjectural identity tested here and the near-integrality observation that the paper refines.","marker":"[Boy98]"},{"why":"Supplies the Beilinson-style prediction linking Mahler measures to elliptic L-values, the conceptual motivation for the identity.","marker":"[Den97]"},{"why":"Gives the modular-Mahler-measure framework and connects the Weierstrass model to k^2, used for the root-number sign and the nonsquare regime.","marker":"[R V99]"},{"why":"Provides the explicit hypergeometric formula for μ_k and the proof for k=1; this formula is the basis of the numerical dataset.","marker":"[RZ14]"},{"why":"Supplies the L'(E_k,0) computation and the rational recognition routine that produces every r_k in the dataset.","marker":"[PAR25]"},{"why":"Provides the Bloch–Kato formula for the 2-part of v_2(n_k), enabling the conditional proof of Conjecture 3.2 on the k=4u subfamily.","marker":"[DGdJK26]"},{"why":"Gives the functional equation used to define μ_k in the nonsquare k=√m regime and the proof from k=2.","marker":"[LR07]"}],"fun_headline_variants":["Boyd's Mahler ratios p-adic laws in 250k elliptic curves","Except seven, r_k is integer reciprocal: p-adic laws emerge","p-adic odds for Mahler ratios: 1/p^m for primes ≥5","Neural nets and stats decode Mahler measure p-adic structure","Boyd's r_k is 1/n: valuations at 2 and 3 show extra structure"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The whole dataset depends on recognizing $r_k$ as a rational number from 19 decimal digits of $\\mu_k$ and $L'(E_k,0)$; if that precision is insufficient for large $k$, some of the 250,000 quotients are misidentified, and every valuation law inherits the error.","fun_headline_variants_meta":{"raw":{"variants":["Boyd's Mahler ratios p-adic laws in 250k elliptic curves","Except seven, r_k is integer reciprocal: p-adic laws emerge","p-adic odds for Mahler ratios: 1/p^m for primes ≥5","Neural nets and stats decode Mahler measure p-adic structure","Boyd's r_k is 1/n: valuations at 2 and 3 show extra structure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001938,"raw_usage":{"total_tokens":7475,"prompt_tokens":858,"completion_tokens":6617,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":602,"completion_tokens_details":{"reasoning_tokens":6510}},"tokens_in":602,"tokens_out":6617,"duration_ms":47861,"temperature":1.0,"reasoning_tokens":6510,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T00:31:55.254509+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute $\\mu_k$ and $L'(E_k,0)$ for 500 randomly chosen $k$ near 250,000 with 100-digit precision and rerun rational reconstruction; if any recovered $r_k$ differs from the 19-digit value, or if the reconstruction failure rate grows with the conductor, the empirical p-adic laws collapse. A complementary check: tally $v_5(n_k)$ on a fresh 250,000-value window, since the reported $\\chi^2/\\nu\\approx 12.8$ for $p=5$ already flags this prime as the outlier.","supporting_citations":[],"review_version":1}