{"id":"2479076e-e152-4a45-9fde-29668f918080","arxiv_id":"2505.05549","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Feed-forward neural networks trained on Fourier coefficients can predict modular weights for negative-weight powers of eta and E2, and for simple Jacobi theta products, within the training range.","lead":"The authors trained simple neural networks to guess the \"weight\" of a modular form, a number that encodes its symmetry behavior, from the first few coefficients of its power series. The networks work well for the negative-weight forms that appear in black hole microstate counting, less well for positive-weight and more complex forms.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 2's E2 experiment appears to train on the eta exponent w rather than the true quasi-modular weight k = w + 2, so the reported negative-weight accuracy is an off-by-two artifact.","rationale":"The reader's verdict is CONDITIONAL, and I agree that the paper's evaluation is not yet sufficient to support the central claim as stated. However, the reader's weakest_assumption, the representativeness of the training distribution and the n_c filtering, is a valid external-generalization concern, while the more load-bearing issue I find is internal: for the E2/quasi-modular experiments, the target variable appears to be the eta exponent w rather than the total quasi-modular weight k = w + 2. Section 2 explicitly constructs the forms as 2 E2(q) Δ(q)^{w/12} and Appendix A writes E2 η^{2w} with weight parameter w, while the product's transformation under (1.7) has weight w + 2 because E2 carries weight 2. Table 2's range (-200, -1/2) is then labeled 'negative weights', but the right endpoint corresponds to total weight +3/2. This is not a matter of disagreement with a convention; it contradicts the paper's own definition and makes the reported 0.04% accuracy dependent on a shifted regression target. A simple recomputation with shifted labels would expose the issue. I therefore do not think the central claim is established for the quasi-modular black-hole counting functions, although the eta-function and Jacobi theta results could still be valid, so the verdict remains CONDITIONAL rather than REJECT. No ad hominem is intended; this is a concrete, testable labeling check. If the GitHub code reveals that the authors did use total weight k despite the text, the concern would be withdrawn. As it stands, the E2 pillar of the proof of concept needs re-derivation before the abstract's claim can be accepted.","tokens_in":13528,"tokens_out":16098,"duration_ms":179717,"concrete_test":"Inspect the data-generation code at the GitHub repository [27] and check whether the target label for Table 2 is w or w+2. Concretely, regenerate the 'Negative 30 coefficients' dataset, replace every target y by k = y + 2, and recompute the mean relative test error. If the error jumps from 0.04% to order 100% (or the network now predicts w+2 while the true labels are w+2), the reported accuracy is an artifact of the shifted target. As a secondary check, evaluate the trained network on 2 E2 / η^24 and verify whether it outputs -10, the true quasi-modular weight, rather than -12.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's Table 2 is the main evidence that neural networks predict weights of quasi-modular (E2-based) forms. In Section 2 and Appendix A, the data are generated from 2 E2(q) Δ(q)^{w/12} = 2 E2(q) η(τ)^{2w}. Since E2 is a quasi-modular form of weight 2 and η^{2w} is modular of weight w, the product is a quasi-modular form of total weight k = w + 2 under the paper's own definition (1.7). Yet Table 2's column header uses (w_min, w_max) = (-200, -1/2) for the 'Negative 30 coefficients' row and calls these negative weights. For w = -1/2, the actual quasi-modular weight is k = +3/2, so the supposedly negative-weight dataset contains positive-weight examples. The networks are therefore trained and evaluated on the label w, not on the physical modular weight k. This explains the very low reported errors: the target is shifted by 2 relative to the true weight. The claim in Section 3 that neural networks 'learn to predict the weights of ... negative powers of E2' is thus not supported until the labels are corrected. The same off-by-two issue propagates into the Kloosterman-sum experiments in Table 2, since the Rademacher formulas in (A.12)-(A.16) are written in terms of the same w parameter. This is a concrete internal inconsistency, not merely a generalization concern: a user applying the trained network to 2 E2/η^24, whose modular weight is -10, would receive a prediction near -12. The central claim about quasi-modular black-hole counting functions therefore rests on an artifact.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper trains feed-forward neural networks to predict the weight (or exponent) of modular, quasi-modular, and Jacobi forms from a finite number of Fourier coefficients. The datasets are generated from powers of the Dedekind eta function, products E2(τ)η^{2w}, and products of Jacobi theta functions. The authors report mean relative test errors below 1% for negative-weight eta and E2 families and for several Jacobi families, with larger errors for positive weights and more complicated products. They position this as a proof of concept for identifying modular symmetries in black hole counting functions.","tokens_in":13914,"tokens_out":13440,"duration_ms":127522,"significance":"If the reported results were valid, the paper would provide a useful proof of concept: a finite Fourier expansion suffices to identify the modular weight of a generating function, which could assist in pinning down unknown counting functions in string theory. The manuscript is honest about the failure on positive weights and outside the training range, and it provides code and data on GitHub, which is a strength. However, as detailed below, the labeling of the training targets is off for the E2 and Jacobi experiments, so the central claim is not currently supported.","major_comments":[{"comment":"The E2 experiments train on the exponent w in E2(τ)η^{2w} rather than the actual quasi-modular weight k = w+2. Since E2 has weight 2 and η^{2w} has weight w, the product is a quasi-modular form of weight w+2 under the paper's own definition (1.7). Table 2 labels the range (−200, −1/2) as 'negative', but for w = −1/2 the true weight is k = +3/2, so the 'negative-weight' dataset contains positive-weight examples. The networks are therefore learning to predict the exponent w, not the modular weight. This is not a harmless shift: the claim in Section 3 that the networks learn weights of 'negative powers of E2' (which are actually E2 times negative powers of η) is unsupported, and a user applying the trained model to E2/η^24, whose true weight is −10, would obtain a prediction near −12. The same issue propagates into the Kloosterman-sum experiments because (A.12)–(A.16) are written in terms of w. The tables and text must be revised with the correct target k = w+2, or the claims must be rescaled accordingly.","section":"Section 2, Table 2, Appendix A (Eqs. (A.12)–(A.16))"},{"comment":"The Jacobi experiments use the power k of θ_a as the prediction target, but the modular weight of θ^k_a is k/2, since θ_a is defined as a Jacobi form of weight 1/2 in Appendix A. Moreover, the statement in Section 2 that 'the coefficient of u^l in θ^k_a is a modular form of weight k + l/2' is inconsistent with this: for l=0 it would give weight k, not k/2. The simple u-expansion coefficient of a Jacobi form is not generally a modular form of SL(2,Z) (e.g., the u^1 coefficient of θ_3 is q^{1/2}, which is not a modular form). The tables report errors on predicting k, so the abstract's claim of predicting 'modular weights' from Jacobi data is not demonstrated. The authors need to either (a) define the target as the true modular weight (k/2 adjusted for the u^l coefficient and any η factor) and retrain, or (b) explicitly restrict the claim to predicting the exponent k and justify why that is the relevant physical quantity.","section":"Section 2, Tables 3–6, Appendix A"},{"comment":"The evaluation provides no baselines and no error bars. All reported errors are single mean relative errors without standard deviation across random seeds or comparisons to simple alternatives (e.g., linear regression on the same Fourier coefficients, or a constant predictor equal to the mean training weight). Because the labels are known analytically, a low error could simply reflect an easy regression problem rather than the network 'learning modular structure.' For a proof-of-concept claim, the authors should include at least one baseline and repeat the experiments a few times to report the spread.","section":"Section 2, Tables 1–6, Appendix B"}],"minor_comments":[{"comment":"'Negative powers of E2' is a misnomer; the experiments use E2 multiplied by negative powers of η, not negative powers of E2.","section":"Section 3"},{"comment":"The statement that imaginary-part experiments are not reported should be quantified (e.g., give the actual poor error) so readers can judge the omission.","section":"Section 2"},{"comment":"The notation '(−200, 30)' in the third column is ambiguous; it should clarify whether the first entry is the minimum weight or the negative of the maximum, and specify the step size in the half-integer and random sampling.","section":"Table 1"},{"comment":"The encoder Log[N[Abs[·]]] is undefined for zero Fourier coefficients; the paper should explain how zeros are handled in the data preprocessing.","section":"Equation (B.2)"},{"comment":"The paper uses 'mock-modular' and 'quasi-modular' interchangeably for E2; this is imprecise and should be corrected to 'quasi-modular' throughout.","section":"Throughout (abstract, Section 2, Appendix A)"},{"comment":"The GitHub repository is a strength, but the paper should list the random seed(s), software versions, and any preprocessing steps that affect the exact numerical results.","section":"Appendix B"},{"comment":"The n_c cut selects only forms with at least n_c non-zero coefficients; a sensitivity analysis (e.g., varying n_c) would clarify how the reported errors depend on this filtering.","section":"Section 2, Tables 3–6"}],"recommendation":"major_revision","confidential_remarks":"The off-by-two labeling error in the E2 experiments is easy to miss because the paper's notation is self-consistent within the w-parameter. The same issue affects the Jacobi experiments, where the target k is the exponent rather than the modular weight. The code and data are publicly available, so a correction and retraining are feasible. I recommend asking for a major revision rather than rejecting, because the underlying idea is sound and the errors are fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, what's worth knowing: this is an honest proof of concept. The idea—train a feed-forward net on truncated Fourier coefficients to predict weights of modular forms relevant to black hole counting—is sensible, and the specific experiments on powers of η, E2, and Jacobi theta products are new relative to the ML-in-number-theory papers they cite. The authors are also transparent about the failure modes: positive weights and complex theta products give large errors, and they flag that the network does not extrapolate beyond the training weight range.\n\nThe main problem is in Table 2. For the E2 experiments, the data are generated from 2 E2(q) Δ(q)^{w/12} = 2 E2(q) η(τ)^{2w}. Since E2 is a quasi-modular form of weight 2, the total quasi-modular weight of this product is w+2, not w. The column labels (w_min, w_max) = (-200, -1/2) are therefore not the actual weights: the dataset includes forms with positive weight (e.g., w=-1/2 gives k=+3/2). The networks are trained and evaluated on the shifted label w, which explains the tiny test errors. A user applying the trained model to 2 E2/η^24 (physical weight -10) would get a prediction near -12. This is not a cosmetic issue; the paper's summary explicitly says the networks learn weights of negative powers of E2, which is only true for the η-exponent, not the modular weight. The Rademacher/Kloosterman experiments in the same table inherit the same shift.\n\nOther soft spots are milder. There are no error bars or baselines; the n_c threshold in the theta tables filters out forms with sparse coefficients, which selects for easy cases; and the imaginary-part experiments are omitted with only a one-line note. The GitHub link has no commit hash, so reproducibility is not fully pinned down. These are normal for a proof of concept, but they raise the bar for trusting the protocol on genuinely unknown forms.\n\nWho should read it: people working on black-hole microstate counting who want a quick heuristic for guessing weights of candidate generating functions. Not yet a reliable tool for symmetry detection. The paper deserves a serious referee—the protocol is worthwhile and the reported negative results are useful—but the E2 labeling error needs a major revision.","headline":"The eta and Jacobi experiments give a credible proof of concept, but the E2 weight labels are off by two, so the quasi-modular headline claim does not hold as stated.","tokens_in":14440,"tokens_out":4988,"would_cite":false,"duration_ms":49715,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["11F11","11F50","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"Neural networks recover modular weights of black-hole counting functions from truncated Fourier series.","keywords":["modular forms","automorphic forms","Jacobi forms","mock-modular forms","BPS black holes","machine learning","Fourier coefficients","modular weight"],"falsifier":"Take a negative-weight weakly holomorphic modular form outside the trained families and weight range (for example, an arbitrary eta-quotient with weight below the training minimum), compute its first 30 Fourier coefficients, and pass them to the trained network; if the predicted weight is no better than a random guess, the claim that the network has learned to identify modular weights from truncated expansions is false.","tokens_in":13358,"feed_emoji":"🧠","tokens_out":8534,"duration_ms":83950,"temperature":0.7,"pith_summary":"The paper asks whether a machine can recover the modular symmetry of a counting function from a handful of its Fourier coefficients. The authors train feed-forward neural networks on coefficients of powers of the Dedekind eta function, the Eisenstein series $E_2$, and products of Jacobi theta functions, and find that the networks predict the modular weight accurately for negative-weight modular and quasi-modular forms, including the weakly holomorphic forms that appear in exact BPS black hole counting. Performance degrades sharply for positive weights and for products of several theta functions. The stated payoff is a proof of concept for using machine learning to detect and verify modular symmetries in quantum gravity when only finitely many coefficients are available.","feed_headline":"Neural nets read Fourier coefficients to spot modular weights","feed_subtitle":"Trained on eta, E2, and theta functions, they nail the negative-weight forms used in BPS black hole counting.","key_machinery":"The central object is the pair consisting of a truncated Fourier coefficient vector and the modular weight it maps to, learned by a regression network. The input coefficients come from $\\eta^n$, $E_2\\eta^{2w}$, and forms built from $\\theta_1,\\theta_2,\\theta_3,\\theta_4$, with powers of $\\eta$ chosen to control the leading power of $q$; inputs are $L^2$-normalized or passed through a log-absolute-value encoder. The network is a deep feed-forward net with ReLU and GELU activations, trained with ADAM on mean squared error. The load-bearing identity is the modular transformation law that defines the weight $k$, together with the Rademacher and Kloosterman expansions that generate coefficient data for $E_2$ powers from polar data.","core_discovery":"The paper's central claim is that a fully connected feed-forward neural network, given the first few Fourier coefficients of a modular, quasi-modular, or Jacobi-derived modular form, can predict its modular weight $k$ --- the exponent in the transformation law $\\phi((a\\tau+b)/(c\\tau+d)) = (c\\tau+d)^k \\phi(\\tau)$. On powers of $\\eta$ with negative weights, test errors are below one percent (for example, $0.16\\%$ for half-integer negative powers and $0.21\\%$ for random real negative powers), and the same holds for negative powers of $E_2$; positive-weight examples fail badly, with test errors of $35.6\\%$ and $383\\%$ in the $\\eta$ and $E_2$ experiments. For Jacobi $\\theta$ functions, accuracy is good for positive powers of $\\theta_3,\\theta_4$ and for products of $\\theta$ functions divided by $\\eta$, but worsens for negative powers of $\\theta_1,\\theta_2$ and for mixed-sign products constrained by $\\mathrm{sgn}(k+l+m+n)$. The authors also note that the trained networks perform poorly on weights outside the training range, so the demonstrated ability is interpolation within a known family rather than extrapolation to arbitrary modular forms.","pith_inferences":["The sharp drop in performance outside the trained weight range suggests the networks are interpolating coefficient-growth patterns rather than learning the modular transformation law itself; a control experiment with randomly shuffled coefficient vectors matched to the same statistics would test this directly.","Because the datasets keep only forms with at least $n_c$ nonzero Fourier coefficients, the reported accuracy may overstate performance on sparse expansions, which are common when only a few terms of an unknown counting function are known.","The strong signal for negative weights may reflect the exponential coefficient growth of weakly holomorphic forms, so the method might also detect mock-modular and other rapidly growing families, a transfer that the paper does not test."],"forward_implications":["Given only the first few Fourier coefficients of a negative-weight modular or quasi-modular form, the method identifies its weight with sub-percent accuracy, narrowing the search for the exact counting function.","The method's failure on positive-weight and mixed-sign theta-function products marks a clear boundary: it is currently reliable for the negative-weight families that appear in exact BPS counting, not for arbitrary automorphic forms.","Applying the same protocol to congruence subgroups of $\\mathrm{SL}(2,\\mathbb{Z})$ is a next step the authors identify for detecting automorphic forms in CHL models.","For the $N=2$ STU model, the method is suited to determining the weight of the putative Jacobi form in the approximate counting function, thereby reducing the space of candidate forms.","Success on these families makes automated detection of modular symmetries in gravitational data a concrete possibility, with applications in AdS/CFT comparisons."],"supporting_citations":[{"why":"Supplies the exact half-BPS black-hole microstate counting function whose modular weight is the target of the method.","marker":"[4]"},{"why":"Establishes modular and mock-modular forms as generating functions for single-centred black holes, motivating the $E_2$ and eta data families.","marker":"[7]"},{"why":"Provides the contour-integral extraction of BPS degeneracies from the Igusa cusp form, the source of the Jacobi-form data.","marker":"[16]"},{"why":"Gives the Rademacher expansion of the Siegel modular form used to generate Fourier coefficients for the $E_2$ experiments.","marker":"[18]"},{"why":"Supplies the STU-model counting function where identifying the weight of a putative Jacobi form is the proposed application.","marker":"[23]"},{"why":"Provides the code and datasets that reproduce the reported training and test results.","marker":"[27]"}],"fun_headline_variants":["Neural net predicts modular weights from Fourier data","ML reads modular forms: negative weights easy, positive hard","Black hole counting meets machine learning: weight prediction","Feed-forward nets weigh modular forms from coefficients","From eta to weights: neural networks on modular forms"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that finite Fourier-coefficient data drawn from the same parametric families and weight ranges used in training represent the modular forms that actually occur in black-hole counting; the paper itself reports poor performance outside those weight ranges and filters datasets to forms with at least $n_c$ nonzero coefficients.","fun_headline_variants_meta":{"raw":{"variants":["Neural net predicts modular weights from Fourier data","ML reads modular forms: negative weights easy, positive hard","Black hole counting meets machine learning: weight prediction","Feed-forward nets weigh modular forms from coefficients","From eta to weights: neural networks on modular forms"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000597,"raw_usage":{"total_tokens":2785,"prompt_tokens":930,"completion_tokens":1855,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":546,"completion_tokens_details":{"reasoning_tokens":1782}},"tokens_in":546,"tokens_out":1855,"duration_ms":13325,"temperature":1.0,"reasoning_tokens":1782,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:02:43.589729+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a negative-weight weakly holomorphic modular form outside the trained families and weight range (for example, an arbitrary eta-quotient with weight below the training minimum), compute its first 30 Fourier coefficients, and pass them to the trained network; if the predicted weight is no better than a random guess, the claim that the network has learned to identify modular weights from truncated expansions is false.","supporting_citations":[{"cited_title":"Rademacher expansion of a Siegel modular form for ${\\cal N}= 4$ counting","cited_arxiv_id":"2112.10023","evidence_quote":"Gives the Rademacher expansion of the Siegel modular form used to generate Fourier coefficients for the $E_2$ experiments."},{"cited_title":"An approach to BPS black hole microstate counting in an N=2 STU model","cited_arxiv_id":"1903.07586","evidence_quote":"Supplies the STU-model counting function where identifying the weight of a putative Jacobi form is the proposed application."},{"cited_title":"ml modular forms","cited_arxiv_id":null,"evidence_quote":"Provides the code and datasets that reproduce the reported training and test results."}],"review_version":1}