{"id":"0cf3f6be-8f22-4d00-a26b-7c0ae9054685","arxiv_id":"2501.02105","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Machine learning predicts the Fricke sign of Maass forms from their first 1,000 Fourier coefficients with about 95% accuracy, and the predictions largely agree with heuristic Hejhal computations.","lead":"This paper uses machine learning on the first 1,000 Fourier coefficients of Maass forms to predict their Fricke sign, an arithmetic invariant, with about 95% accuracy, and applies the model to thousands of forms whose signs were unknown. The practical payoff is a fast heuristic for filling in missing data in the LMFDB and new evidence that statistical patterns in coefficients carry arithmetic information.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"High accuracy is established only on rigorously labeled forms; transfer to the 15,423 unknown-sign forms rests on agreement with a selective Hejhal subset whose representativeness is not shown.","rationale":"The paper has a genuine empirical finding: LDA on 1000 coefficients achieves 96.1% validation accuracy on labeled forms, and the zeroing experiment (Table 2.2, a'_n) shows most of the signal survives removal of the direct sign-encoding coefficients, so the result is not a trivial readout. The 95% agreement with the Hejhal heuristic on 4,595 unknown-sign forms is real independent support, and the neural network results are comparable. However, the paper's advertised use case is predicting signs for all 15,423 unknown forms, and that step is only heuristically validated. The unknown-sign set is missing signs precisely because numerical precision was insufficient; Figure 2.9 shows the unknown fraction grows with level, so the labeled set is biased toward easier, lower-level forms. The Hejhal-confirmed subset is a subset where a heuristic converged, and the paper does not compare its feature/level/resonance distribution to the rest of L0. If the confirmed subset is concentrated at low levels, the 95.45% agreement does not establish transfer to higher-level unknowns. The murmur-based validation (Figures 2.4, 2.5) is partially circular because the same coefficient averages inform both the classifier and the validation. The proposed test directly settles representativeness by stratifying the Hejhal-confirmed subset; this is a targeted, low-cost analysis the authors can run with their existing data. For these reasons, the reader's CONDITIONAL verdict is appropriate; no change in verdict is needed, but the paper should add this stratification.","tokens_in":10623,"tokens_out":12385,"duration_ms":127265,"concrete_test":"Stratify the 4,595 Hejhal-confirmed unknown-sign forms by level (e.g., quartiles) and by spectral parameter R; compute LDA's agreement with Hejhal within each stratum, and compare the level/R distribution of the confirmed subset with that of the 10,828 unconfirmed unknown forms. If agreement stays approximately 95% in high-level and low-R-spacing strata and the confirmed subset spans the full level range, transfer is supported. If agreement drops sharply at high levels or the confirmed subset under-represents them, the transfer assumption fails and the prediction accuracy on most unknown forms is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the classifier's high accuracy on the labeled subset (w=±1) carries over to the 15,423 forms with unknown w. This transfer is load-bearing because applying the model to L0 is a major output, and direct validation there is limited to 4,595 forms where a heuristic Hejhal algorithm converges. The paper notes (Figure 2.9) that the unknown fraction grows with level, so L1⊔L−1 is biased toward low levels. It does not show that the 4,595 Hejhal-confirmed forms are representative of the full L0; convergence of the heuristic is likely easier for lower levels and larger eigenvalue gaps, exactly the regimes where the labeled set is concentrated. Thus the 95.45% agreement with Hejhal (Table 2.4) may not reflect accuracy on higher-level unknowns. The paper's additional checks (murmuration plots of predicted signs) are partially self-fulfilling, since the classifier is trained on the same averaged coefficient patterns. If the Hejhal-confirmed subset skews low-level, the central claim that predictions on the full unknown set are 'reasonable' is unproven.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies supervised machine learning, specifically Linear Discriminant Analysis (LDA) and feed-forward neural networks, to the first 1000 Fourier coefficients of Maass newforms from the LMFDB in order to predict the Fricke sign. On the 19,993 forms with rigorously known sign, LDA is reported to reach about 96% validation accuracy. The trained model is then applied to the 15,423 forms with unknown Fricke sign, and the resulting predictions are checked in two ways: by comparing averaged coefficient patterns ('murmurations') for predicted signs with those for known signs, and by comparing with heuristic Hejhal guesses on a subset of 4,595 forms, where agreement is about 95%. The paper also addresses the concern that the Fricke sign is directly encoded in coefficients at primes dividing the level by repeating experiments with a modified feature in which such coefficients are zeroed.","tokens_in":10821,"tokens_out":7402,"duration_ms":78536,"significance":"If the transfer claim holds, the paper provides a useful empirical data-scientific tool for guessing missing Fricke signs in the LMFDB and evidence that coefficient vectors contain recoverable information beyond the direct local-factor readout. The paper is commendably explicit about the direct-encoding confound and tests a zeroed feature variant, and it compares its predictions with an independent heuristic algorithm. The main weaknesses are the lack of uncertainty quantification and, more importantly, the absence of evidence that the 4,595-form Hejhal-confirmed subset is representative of the full unknown-sign set, which is load-bearing for the claim that predictions on all unknown-sign forms are reasonable.","major_comments":[{"comment":"The abstract states 96% (resp. 94%) accuracy for even (resp. odd) parity, but Section 2.3 and Table 2.2 report 94.9% for even forms and 96.3% for odd forms. Since these are the headline accuracy figures, the abstract should be corrected to match the table.","section":"Abstract and Section 2.3/Table 2.2"},{"comment":"The application to the 15,423 unknown-sign forms is a central output of the paper, but the only direct validation is on the 4,595 forms for which the heuristic Hejhal algorithm converged. Figure 2.9 shows that the unknown fraction increases with level, and the convergence of a heuristic root-finding algorithm is likely easier at lower levels and for larger eigenvalue gaps. The paper does not report the level or spectral-parameter distribution of the 4,595 Hejhal-confirmed forms, nor accuracy stratified by level. Without such evidence, the claim that predictions on the full unknown-sign set are 'reasonable' is not established.","section":"Sections 2.3, 2.7, Table 2.4, Figure 2.9"},{"comment":"The sentence 'Similarly, when gcd(n,N)>1, we have a_n = 0' is inconsistent with Eq. (2.4), which gives the nonzero value a_p = -w_p/sqrt(p) for each prime p dividing N. The intended statement is presumably that entries not rigorously computed are set to zero in the database. This matters because the a'_n experiment is the main evidence that the classifier learns beyond the direct encoding, so the exact zeroing rule (which indices are zeroed, and whether coefficients with mixed prime factors are handled correctly) must be stated precisely.","section":"Section 2.4, Eq. (2.4)"},{"comment":"All accuracy figures are single-split point estimates with no confidence intervals, standard errors, or repeated-seed variation. The reported differences between feature sets, such as 0.9612 for a_n versus 0.9456 for a'_n, could be within sampling noise. Please provide bootstrap intervals or results over several random splits, especially for the comparison that supports the 'learning something more' claim.","section":"Tables 2.2-2.6 and Section 2.3"}],"minor_comments":[{"comment":"The text says these figures provide 'clear evidence' of separation, but no quantitative measure is given; consider adding a simple statistic such as the L2 distance between the averaged coefficient sequences or the area between the curves.","section":"Figures 2.1 and 2.2"},{"comment":"The training sizes 7772 and 5023 for even and odd forms are not tied to the described 80-20 splits; please clarify how these counts are obtained from Table 2.1.","section":"Section 2.3"},{"comment":"The use of Box's M test to 'satisfy' equal covariance is not rigorous: rejecting equality for only 33 of 1000 features is not the same as establishing equality, and the test is sensitive to sample size. A direct comparison of covariance matrices would be more appropriate.","section":"Section 2.3"},{"comment":"The phrase 'without any hyperparameter tuning' appears in Section 2.3, but the neural network experiments in Section 2.6 use Adam with a learning rate of 1e-3 and 4e4 iterations; the claim should be restricted to the LDA experiments.","section":"Sections 2.3 and 2.6"},{"comment":"There is a typo in the caption of Table 2.5: 'Maaass forms' should be 'Maass forms', and the space in 'F ricke sign' in the Section 2.3 heading should be removed.","section":"Section 2.7 and Table 2.5"},{"comment":"No code, data, or reproducibility statement is included. Since the experiments are computational, please provide a link to a repository with the exact data-processing steps, random seeds, and model configurations, or at least specify the seed and split procedure.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The labeled-data accuracy result is plausible and the a'_n experiment addresses the most obvious tautology, so I would not reject the paper outright. However, the transfer to unknown-sign forms is the part that would make the paper valuable, and it currently rests on an unexamined representative assumption. If the authors can provide stratified validation on the Hejhal subset and temper the claims about L0 accordingly, the paper would be acceptable for an experimental-math venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The main empirical finding is plausible and likely correct: LDA and a small neural net can predict the Fricke sign of a Maass form from its first 1000 coefficients at around 95% accuracy, even after zeroing the coefficients that directly encode the sign. The parity-corrected murmuration plots are a genuine improvement over [BLLD+24], and the zeroed-feature experiment in Section 2.4 is the right check. I think the paper is worth engaging with, but it needs cleanup.\n\nWhat is actually new: applying standard classifiers to Fricke signs, and the observation that including composite-index coefficients helps beyond primes alone. The latter is curious and not fully explained, but the experiments in Section 2.4 and Figure 2.7 are honest attempts to probe it.\n\nThe soft spots, in order of severity. First, the abstract swaps the even/odd accuracy figures: it says 96% even, 94% odd, while Table 2.2 gives 0.9488 even and 0.9633 odd. That is a concrete error that will embarrass the authors. Second, no code is shipped, and there are no confidence intervals or repeated-seed variance for the neural net, so the reader cannot tell how stable the 95% figures are. Third, and most substantive, the transfer from labeled to unknown-sign forms is not established. The unknown fraction grows with level (Figure 2.9), and the 4,595 Hejhal-confirmed forms are a selective subset; the paper does not show their level distribution or that Hejhal convergence is independent of level. Without that, the 95% agreement with Hejhal is suggestive but not proof that predictions on the full unknown set are reliable.\n\nThe direct-encoding issue is real but partially handled. For squarefree N up to 105, a_p for p|N gives w_p exactly, so the unzeroed a_n features contain a tautological readout. The a'_n experiment removes that and still gets 94-96%, which is the actual result. The paper says this clearly enough, but the abstract's emphasis on the unzeroed accuracy overstates the novelty.\n\nThe math looks solid, definitions are standard, and the citation pattern is fine. This is a computational paper, not a theorem paper, and the claims are framed as empirical. Who is it for? People working on murmurations, machine learning for L-functions, and the LMFDB Maass form database. It deserves a serious referee; I would send it to review after asking for the abstract fix, error bars, and a level-distribution check for the Hejhal subset.","headline":"A useful empirical paper whose core claim is probably right, but whose headline numbers are sloppy and whose transfer to unknown-sign forms is only heuristically supported.","tokens_in":11436,"tokens_out":2611,"would_cite":true,"duration_ms":26819,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["11F66","11Y35","68T05"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that supervised machine learning on the first 1,000 Fourier coefficients predicts the Fricke sign of a Maass form with 94–96% accuracy, and that predictions for forms with unknown signs agree with a heuristic algorithm…","keywords":["Maass forms","Fricke sign","murmurations","linear discriminant analysis","Fourier coefficients","machine learning","L-functions","neural networks"],"falsifier":"Take a random sample of the 15,423 forms with unknown Fricke sign, compute their signs rigorously at higher precision, and compare with the LDA predictions; if the error rate on that sample is near 50% rather than near 5%, the transfer claim is false.","tokens_in":10363,"feed_emoji":"📈","tokens_out":6566,"duration_ms":62058,"temperature":0.7,"pith_summary":"The paper tries to establish that the Fricke sign of a Maass newform—the eigenvalue under the Fricke involution that fixes the sign of the functional equation—can be recovered statistically from finitely many Fourier coefficients, without computing the form to the precision normally required. Averaging coefficients by sign reveals murmuration-like oscillations, and a simple linear classifier trained on these coefficients predicts the sign with 94–96% accuracy on forms where the sign is rigorously known. Applied to the 15,423 forms whose signs are unknown, the model's predictions reproduce the same averaging patterns, and on the 4,595 forms where a heuristic algorithm is confident, the predictions match roughly 95% of the time. The reason this matters is that Fricke signs are often unavailable precisely because coefficients cannot be computed accurately enough; a collective statistical readout would supply sign information cheaply across the whole database.","feed_headline":"94–96% accuracy: one classifier predicts Maass-form signs","feed_subtitle":"A simple linear model trained on 1,000 Fourier coefficients also fills in the 43% of signs the database leaves unknown.","key_machinery":"The central object is the normalized feature vector $$D = \\{((-1)^{\\$\\sigma$(f)} a_n)_{n=1}^{1000} : f \\in \\mathcal{L}\\},$$ in which each coefficient is multiplied by the parity sign so that even and odd forms are aligned by root number, together with the factorization of the Fricke sign into local factors $w_N = \\prod_{p\\mid N} w_p$. The machinery that carries the argument is Linear Discriminant Analysis (LDA), which fits a linear decision boundary under the assumption that the two sign classes share a covariance structure; the paper checks this assumption with a standard equal-covariance test. A second piece of machinery is the averaging operation that produces murmuration plots, which both motivates LDA and validates predictions by comparing average coefficients of predicted-sign forms with known-sign forms. The paper also uses a neural network with the spectral parameter $R$ appended, and a heuristic algorithm that guesses Fricke signs by solving approximate overdetermined linear systems; agreement with that heuristic provides the external check on the unknown-sign predictions.","core_discovery":"On the paper's own terms, the discovery is that the Fricke sign of a Maass form is learnable from the first 1,000 Fourier coefficients, with accuracy far above the 86% obtainable from prime-indexed coefficients alone, and robust to masking the coefficients whose indices share a factor with the level—the obvious place where the sign is encoded. The authors argue the classifier is not merely reading off $a_p$ for primes dividing the level, because setting those coefficients to zero leaves accuracy nearly unchanged for the best feature set and because level-1 forms, where no coefficient directly encodes the sign, are classified successfully. Instead, the full coefficient vector carries extra predictive signal: indices with one or two prime factors give 95.3% accuracy, so the multiplicative structure itself appears informative. This connects the predictive task to murmurations: the same sign-conditioned averages that oscillate in the plots are what the linear classifier exploits.","pith_inferences":["Because the root number is the product of parity and Fricke sign, a classifier with this accuracy also yields a root-number estimator; this could be used to prioritize candidates for rigorous certification.","The finding that full coefficient vectors beat prime-indexed ones suggests a testable hypothesis: the signal lives in the multiplicative semigroup structure of indices, so features derived from divisor counts should retain most of the accuracy.","Since unknown signs become more frequent at higher level (as the paper's own level-by-level plot shows), a level-stratified retraining experiment would be the cleanest way to test whether the transfer assumption holds.","The same averaging-plus-classifier pipeline could be tried on other expensive invariants of automorphic forms, such as symmetry type or eigenvalue location, wherever a labeled subset exists."],"forward_implications":["If the central claim holds, Fricke signs for the 15,423 unknown forms can be assigned probable values using only stored coefficients, matching independent heuristics on the subset where the heuristic is confident.","The accuracy advantage of full coefficient vectors over prime-indexed vectors (96.1% versus 86.2%) implies the predictive information is not contained only in $a_p$ for primes dividing the level, nor in primes generally; composite-index coefficients add real signal.","Because LDA works without hyperparameter tuning, the sign is nearly linearly separable in coefficient space after parity normalization, suggesting a simple statistical description rather than a deep structural one.","Because the model transfers across levels without being given the level, the learned boundary is a function of coefficient patterns rather than of the level or analytic conductor alone.","The connection to murmurations suggests sign-conditioned coefficient averages are a stable phenomenon, not an artifact of a particular classifier."],"supporting_citations":[{"why":"Supplies the 35,416 rigorously computed Maass forms, including the 15,423 with unknown Fricke signs.","marker":"[LMF24]"},{"why":"Establishes the murmuration phenomenon for Maass forms that motivates the averaging analysis.","marker":"[BLLD+24]"},{"why":"Is the reference for the Linear Discriminant Analysis classifier used throughout.","marker":"[HTF01]"},{"why":"Provides the original heuristic algorithm later adapted for the comparison baseline.","marker":"[Hej99]"},{"why":"Is the implementation of the heuristic algorithm that produced the 4,595-form comparison set.","marker":"[LD24]"},{"why":"Gives the explicit trace-formula method behind the rigorously computed coefficients and signs.","marker":"[SH22]"},{"why":"Ties murmurations of elliptic curves to machine-learned parity, the template for this study.","marker":"[HLOP24]"}],"fun_headline_variants":["94–96% accuracy: learning Maass form signs from coefficients","Fricke signs predicted with 96% accuracy from Fourier data","Murmuration-like patterns reveal Maass form signs","Data-driven discovery: Maass form signs are learnable","Classifier decodes Fricke signs beyond known database"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The classifier's transfer to the 15,423 unknown-sign forms assumes that forms whose signs are unknown because of computational difficulty are statistically similar, on the coefficient features used, to forms whose signs are known.","fun_headline_variants_meta":{"raw":{"variants":["94–96% accuracy: learning Maass form signs from coefficients","Fricke signs predicted with 96% accuracy from Fourier data","Murmuration-like patterns reveal Maass form signs","Data-driven discovery: Maass form signs are learnable","Classifier decodes Fricke signs beyond known database"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000351,"raw_usage":{"total_tokens":1906,"prompt_tokens":932,"completion_tokens":974,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":891}},"tokens_in":548,"tokens_out":974,"duration_ms":9718,"temperature":1.0,"reasoning_tokens":891,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:15:33.282827+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of the 15,423 forms with unknown Fricke sign, compute their signs rigorously at higher precision, and compare with the LDA predictions; if the error rate on that sample is near 50% rather than near 5%, the transfer claim is false.","supporting_citations":[],"review_version":1}