{"id":"86d2b9ad-13ff-47f1-a40c-350809b766f8","arxiv_id":"1908.08925","paper_version":4,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Digit block frequencies in 600 billion digits of Catalan's constant and the lemniscate arc length match the expected random distribution within about two standard deviations.","lead":"This paper counts how often each one-, two-, and three-digit block appears in the first 600 billion digits of Catalan's constant and the arc length of a lemniscate, in decimal and hexadecimal. The digit blocks appear with the frequencies expected for a random-looking number, adding evidence to the conjecture that these constants are normal.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The conclusion depends on the authenticity of the archived 600-billion-digit prefixes; the paper provides no independent checksum, third-party verification for Catalan, or full recomputation check.","rationale":"The paper's statistical machinery is internally consistent: the predicted variances use binomial dispersion for the mean of the variance estimator, and the small inflation of the error bars due to overlapping blocks does not materially change the reported deviations. The interpretation of a maximum 1.8-sigma deviation as consistency rather than proof is appropriate, so I do not see a statistical flaw that would invalidate the central claim. The one place where the argument can silently fail is if the input data are not the claimed constants. World-record computations are complex, and the paper does not provide an independent verification trail: no hashes, no comparison against a second implementation for a large segment, and the acknowledgment section does not state that the Catalan computation was externally verified. Because all conclusions are conditional on the input, this is the load-bearing concern. The suggested recomputation test would settle it. The reader's weakest assumption identifies the same issue, so I agree with the reader's assessment. The verdict remains conditional on making verification artifacts available.","tokens_in":5523,"tokens_out":21744,"duration_ms":219575,"concrete_test":"Download a randomly selected 1-billion-digit block from each of the four archived strings (Catalan decimal, Catalan hex, lemniscate decimal, lemniscate hex) and recompute that block with an independent implementation, for example y-cruncher using a different algorithm or an independent library for a smaller prefix. Compare the recomputed digits byte-for-byte with the archived files, and also compare the SHA-256 of the full archived files against hashes published at the time of computation (Kim 2019a,b). If the recomputed block matches, the data-authenticity risk is substantially reduced.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section 4 is that the frequency variances for the first 600 billion digits are persistent with normality of the two constants. This holds only if the archived decimal and hexadecimal strings are actually the prefixes of Catalan's constant and the lemniscate arc length. Section 2 states that single-digit counts were cross-checked with y-cruncher's Digit Viewer, but this is a spot check, not a validation of the full files. The Data Availability Statement gives Internet Archive links but no checksums or a generation script. The acknowledgments mention Dr. Cutress verifying the lemniscate world record, but not the Catalan computation. If an archived file is truncated, corrupted, or mislabeled, the variance values would describe a different string, and the normality conclusion would not apply to the constants. This is the weakest load-bearing step because every downstream number inherits the data's correctness.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes digit frequency distributions in the first 600 billion decimal and hexadecimal digits of Catalan's constant and the arc length of a lemniscate with a=1, using data from the author's own y-cruncher world-record computations. For each of the two constants, each base (10 and 16), and each block length 1, 2, and 3, it computes the sample variance of the frequencies of all b^k digit combinations and compares this with the variance expected under a binomial model with success probability b^{-k}. The measured variances are reported to lie within 1.8 standard deviations of the predicted values, and the paper concludes that the data are 'persistent to' (i.e., consistent with) the conjecture that these constants are normal in bases 10 and 16. The paper includes histograms of the frequency distributions and tables of predicted and actual variances.","tokens_in":5663,"tokens_out":7883,"duration_ms":74757,"significance":"If the archived digit strings are indeed the correct prefixes, the paper provides a straightforward, parameter-free statistical check on the normality conjecture for two important constants at an unprecedentedly large scale. The core comparison is simple and transparent, and the conclusion is appropriately cautious, using 'persistent' rather than claiming a proof. The data are made publicly available on the Internet Archive, which supports reproducibility. The main limitations are that only block lengths 1 to 3 are tested, so the evidence for normality is very weak, and that the integrity of the data is not fully established; both issues are addressable in revision. The paper is a useful data-driven sanity check for world-record constant computations.","major_comments":[{"comment":"The entire analysis depends on the correctness of the 600-billion-digit files archived at the Internet Archive, but the only verification described is a cross-check of single-digit counts using y-cruncher's Digit Viewer, which is a spot check rather than a validation of the full files. No checksums, independent recomputation results, or repeatable generation scripts are provided, and the acknowledgment mentions verification of the lemniscate record but not of the Catalan record. Please either (a) provide a full verification protocol, such as checksums for both files, an independent recomputation of a substantial portion, or a script that regenerates the data, or (b) explicitly restrict the conclusion to the archived strings so that the reader can judge whether the evidence bears on the constants themselves rather than on possibly corrupted or mislabeled files.","section":"Section 2 and Data Availability Statement"},{"comment":"The paper reports a 'predicted variance and error of frequencies' for each case, for example (1.500±0.707)×10^-13 for decimal length 1, and uses the quoted error to compute the 'Deviation [σ]' column, but it gives no formula for the error on the predicted variance. The quoted values indicate that the error was computed as the standard error of a sample variance from independent normal observations, whereas the b^k block counts are jointly multinomial and are not independent. Please state the exact estimator used for the sample variance, its expectation under the null, and the formula for its standard error (or cite the precise equation in Trueb 2016), and justify the normal approximation for small numbers of cells such as 10. Without this information, the reported 1–2σ thresholds cannot be reproduced or interpreted reliably.","section":"Section 2, Tables 1 and 2"}],"minor_comments":[{"comment":"The caption for Figure 10 says 'digits 0–9' for a hexadecimal length-1 histogram, but it should say 'digits 0–F'.","section":"Figure 10 caption"},{"comment":"The caption for Figure 11 says '00–99' for a hexadecimal length-2 histogram, but it should say '00–FF'.","section":"Figure 11 caption"},{"comment":"The caption for Figure 12 says '000–999' for a hexadecimal length-3 histogram, but it should say '000–FFF'.","section":"Figure 12 caption"},{"comment":"The phrase 'persistent to the conjecture' is nonstandard; 'consistent with the conjecture' would be clearer. The Discussion should also explicitly state that the analysis covers only block lengths 1 through 3 and that this provides only weak evidence for normality, since a number can pass these tests while failing for longer block lengths.","section":"Abstract and Discussion"},{"comment":"The expected variance of the frequencies is stated to follow from the binomial model, but the paper does not explain how the 'error of the predicted variance' is derived; please add a sentence or a reference to a specific equation in Trueb (2016).","section":"Section 2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a modest but potentially useful statistical report on two world-record constant computations. The central statistical comparison is straightforward and the measured deviations are small, but the two issues identified in the major comments—data integrity and missing derivation of the error on the predicted variance—need to be resolved before the manuscript can be accepted. The limited block-length coverage is a weakness that should be acknowledged more prominently, but it does not by itself invalidate the paper's cautious conclusion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid data-check paper, not a theorem paper. Kim computed 600 billion digits of Catalan's constant and the lemniscate arc length, then ran Peter Trueb's binomial-variance test on digit blocks of length 1 to 3 in base 10 and 16. All observed variances sit within about 1.8 sigma of the theoretical variance, so the data are consistent with normality in those bases. The statistical result is real and the conclusion is appropriately cautious — \"persistent to\" rather than \"proof of.\"\n\nWhat's new is the data: this is the first time these two constants have been pushed to 6e11 digits and subjected to this test. The method is standard, but for a constants-and-normality question that's the normal state of play: you can't prove normality with finite data, so the best you can do is run the same weak test on bigger prefixes. The paper credits Trueb and links the archived data on the Internet Archive, and the computation itself was done with y-cruncher, which is a known quantity.\n\nThe soft spots are mostly about provenance and power. The archived files have no checksums, and the only cross-check is a spot check of single-digit counts using y-cruncher's Digit Viewer. If a file is truncated or mislabeled, every variance in Tables 1 and 2 describes a different string. That is the load-bearing assumption, and the paper doesn't fully pin it down. It would be nice to have a generation script or a reproducible recomputation of a prefix. Second, the test only covers block lengths 1–3. That is a weak test: it would miss, for example, a number that is normal in short blocks but has long-range correlations. The paper doesn't oversell it, but readers should know the discriminating power is limited. Also, the claim in Section 2 that \"integrity of arbitrary length digit combinations are ensured if single digit counts are correct\" is a bit of a non-sequitur unless the counting script treats overlapping windows in a particular way; it doesn't matter much for the result, but it's worth asking.\n\nWho should read it: people working on normality of specific constants, and anyone who wants a clean example of how to run Trueb's test on new data. It's not a paper that changes the mathematical picture, but it is a legitimate contribution to the accumulating empirical evidence. A serious editor should send it to peer review — the data is important, the method is standard, and the flaws are fixable with better documentation rather than being fatal.","headline":"A modest but honest data-check paper: new 600-billion-digit prefixes for Catalan's constant and the lemniscate arc length, analyzed with Trueb's variance test, all consistent with normality.","tokens_in":6157,"tokens_out":2391,"would_cite":false,"duration_ms":23432,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["11-XX","62-XX","68-XX"],"pacs":[],"model":"deepseek-v4-flash","headline":"The first 600 billion digits of Catalan's constant and the lemniscate arc length show digit frequencies consistent with normality in bases 10 and 16.","keywords":["Catalan's constant","lemniscate constant","normal number","digit analysis","digit frequency","binomial variance","base 10","base 16"],"falsifier":"Recompute a contiguous prefix of each constant with an independent algorithm and implementation—for example, binary-splitting the series $\\sum_{n=0}^{\\infty}(-1)^n/(2n+1)^2$ for Catalan's constant—and compare digit-by-digit with the archived strings; any mismatch would show the analyzed data are not the true prefix. A second check: run the same variance computation on block lengths 4 and 5 of the archived strings; deviations of several standard deviations there would contradict the paper's claim that the streams behave like a normal number.","tokens_in":5319,"feed_emoji":"🎲","tokens_out":11736,"duration_ms":109072,"temperature":0.7,"pith_summary":"Normal numbers—real numbers in which every finite block of digits appears with its equal limiting frequency—are conjectured for most familiar constants but proved for almost none. This paper tests that conjecture on the current world-record digit computations: $600{,}000{,}000{,}100$ decimal digits of Catalan's constant, $600{,}000{,}000{,}000$ decimal digits of the arc length of a lemniscate with $a=1$, and the matching hexadecimal strings. Counting every digit block of length 1, 2, and 3 in both bases, the author finds that the actual variance of the block frequencies stays within about $1.8$ standard deviations of the variance a binomial model predicts for a uniformly random digit stream. The conclusion is that the data are consistent with, though not proof of, the conjecture that both constants are normal numbers in bases 10 and 16.","feed_headline":"600B digits of two constants pass normality checks","feed_subtitle":"In bases 10 and 16, every one- to three-digit block matches random expectation within 1.8 sigma.","key_machinery":"Central to the argument is the variance of the frequency counts of all $b^k$ digit strings of a fixed length $k$ in base $b$. Under the null hypothesis that every string is equally likely, the count of any particular string among $N$ digit positions follows a binomial distribution with probability $b^{-k}$, giving a predicted variance of roughly $N b^{-k}(1-b^{-k})$. The paper compares the actual variance of the $b^k$ observed frequencies with that predicted variance (including its standard error), and displays the results as histograms with one- and two-standard-deviation bands. This variance comparison, rather than a direct check of each individual frequency, is the appropriate test because individual frequencies fluctuate as $N^{-1/2}$ even for a truly normal number. The approach is adopted from a previous digit-statistics analysis of the first $22.4$ trillion decimal digits of $\\pi$.","core_discovery":"The paper's central claim is that, at the scale of roughly 600 billion digits, neither constant shows a statistically meaningful departure from the digit-frequency behavior expected of a normal number. For each of the twelve cases—two constants, two bases (10 and 16), and block lengths 1–3—the sample variance of the observed frequencies lies within about $1.8$ standard deviations of the predicted variance. The largest absolute deviation is $1.798\\,\\sigma$, for two-digit decimal blocks in the lemniscate arc-length string; the hexadecimal results for the same constant deviate by at most $1.080\\,\\sigma$, and Catalan's constant by at most $1.349\\,\\sigma$ in either base. The paper states these results as 'persistent to' the conjecture that both constants are normal in bases 10 and 16, explicitly presenting the analysis as statistical evidence rather than a proof.","pith_inferences":["Extending the same variance test to four- and five-digit blocks would be a sharper check than anything in the paper, because the binomial predictions remain stable while the number of categories grows to $10^4$ or $10^5$ in base 10.","Hexadecimal normality is the more significant agreement for binary structure: a hexadecimal digit is exactly four bits, so uniform hex digits imply uniform bit strings, whereas uniform decimal digits give only indirect information about bits.","If the pattern persists, the digit streams of these constants could be used as deterministic pseudo-random sources; normality alone, however, would not certify them for cryptographic use without further randomness tests."],"forward_implications":["On the paper's evidence, the normality conjecture for Catalan's constant and the lemniscate arc length survives the largest digit sample yet examined, with no block of length 1–3 showing anomalous frequency variance.","Because the method is data-driven and uses archive-accessible digit strings, it can be applied immediately to the world-record prefixes of other constants to screen for early signs of non-normality.","At the 600-billion-digit scale, both constants' digit streams match a uniform random model closely enough that block-level statistical tests on these constants should expect near-uniform output.","If the data are genuine, these variance numbers furnish reference points for how a normal candidate behaves over 600 billion digits, useful for calibrating future tests."],"supporting_citations":[{"why":"Supplies the variance-comparison method, standard-error formulas, and the one/two-sigma plotting conventions adapted here from a pi digit analysis.","marker":"[Trueb, 2016]"},{"why":"Provides the world-record decimal and hexadecimal prefixes of Catalan's constant that are the subject of the analysis.","marker":"[Kim, 2019a]"},{"why":"Provides the world-record decimal and hexadecimal prefixes of the arc length of a lemniscate with $a=1$ that are analyzed.","marker":"[Kim, 2019b]"},{"why":"Documents the computation program and the Digit Viewer tool used to produce and cross-check the digit records.","marker":"[Yee, 2019]"}],"fun_headline_variants":["600B digits of Catalan and lemniscate constants look normal","Record digit counts: both constants pass normality checks in bases 10 and 16","1.8 sigma max deviation in 600B digits supports normality","World-record digits of two constants align with normality","Normality holds for Catalan and lemniscate constants in bases 10 and 16"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes that the archived 600-billion-digit strings are authentic, correctly computed prefixes of the two constants; if the computation or the archived copies contain any error, the frequency statistics describe a different number and the normality evidence is void.","fun_headline_variants_meta":{"raw":{"variants":["600B digits of Catalan and lemniscate constants look normal","Record digit counts: both constants pass normality checks in bases 10 and 16","1.8 sigma max deviation in 600B digits supports normality","World-record digits of two constants align with normality","Normality holds for Catalan and lemniscate constants in bases 10 and 16"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000835,"raw_usage":{"total_tokens":3794,"prompt_tokens":874,"completion_tokens":2920,"prompt_tokens_details":{"cached_tokens":768},"prompt_cache_hit_tokens":768,"prompt_cache_miss_tokens":106,"completion_tokens_details":{"reasoning_tokens":2825}},"tokens_in":106,"tokens_out":2920,"duration_ms":282683,"temperature":1.0,"reasoning_tokens":2825,"cache_read_input_tokens":768,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:45:37.444195+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute a contiguous prefix of each constant with an independent algorithm and implementation—for example, binary-splitting the series $\\sum_{n=0}^{\\infty}(-1)^n/(2n+1)^2$ for Catalan's constant—and compare digit-by-digit with the archived strings; any mismatch would show the analyzed data are not the true prefix. A second check: run the same variance computation on block lengths 4 and 5 of the archived strings; deviations of several standard deviations there would contradict the paper's claim that the streams behave like a normal number.","supporting_citations":[],"review_version":1}