{"id":"07a7fbca-8e14-4f45-b86d-0ad2988f9488","arxiv_id":"2603.12785","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An upper-bound formula for local learning coefficients at singular points of three-layer networks is derived via blow-ups and matches known exact coefficients when the input dimension is one.","lead":"The paper gives a closed-form upper bound on local learning coefficients of three-layer neural nets at singular realization parameters, for general real-analytic activations including swish. The bound is a counting rule under budget, demand, and supply constraints and matches known exact values when the input dimension is one.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection beyond the independence caveat already flagged by the reader; the Main Theorem's upper-bound claim holds under its stated hypotheses.","rationale":"The central claim is an upper bound under explicit analytic and linear-independence hypotheses, not an unconditional exact formula for every activation. The proof outline (Section 5 + Appendices A–C) is detailed enough that an expert can follow the successive normal-crossing calculations; the N=1 cases recover the known exact values of Aoyagi, and the reduced-rank comparison correctly flags where the inequality is strict. The only genuine soft spot is precisely the independence condition already highlighted by the reader; when it fails the theorem simply does not apply, which the paper states. No deeper algebraic gap or hidden assumption that would invalidate (2.1) under the stated hypotheses was found. Hence the CONDITIONAL verdict with medium correctness risk remains appropriate; no adjustment is warranted.","tokens_in":38937,"tokens_out":609,"duration_ms":5213,"concrete_test":"Independently recompute the local RLCT for the concrete N=1,H=4,M=1,tanh example of §4 (true H*=1) by a second resolution path (e.g., different order of blow-ups or a computer-algebra Gröbner/normal-crossing routine on the ideal generated by the gs). If the resulting min (h_i+1)/k_i is strictly smaller than 11/6, the upper-bound derivation has a gap; if it equals 11/6, the Main Theorem application is confirmed for that instance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest_assumption correctly isolates the load-bearing point: Main Theorem conditions (i)–(iv), especially linear independence of the Zs,n (and the partials of f) in (iii), plus the concrete hypotheses (3.1)/(3.3) for three-layer nets. When those hold, the four-step blow-up sequence (Step 1: m1 blow-ups of (θ,b); Step 2: a↦a' on the open set of condition (ii); Steps 3-k: mk+1-mk blow-ups; Step 4: final blow-up of (θ',a')) produces the normal-crossing form whose RLCT is bounded by the budget–demand–supply expression (2.1). The paper itself records the counter-examples (Remark 3.2(3) for swish-type weights; H*=0 restriction for polynomials) and the cases where the bound is strict (reduced-rank case 1). No further internal inconsistency appears in the chart-by-chart Jacobian calculations or in the combinatorial reading of K,L,n*_s. The abstract's claim of “general settings” for non-polynomial analytic activations is therefore accurate only modulo the independence hypotheses that the body already qualifies.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper derives an upper-bound formula (Main Theorem, Eq. 2.1) for the local learning coefficient (real log canonical threshold) at a class of singular realization parameters of statistical models whose log-likelihood ratio admits a Taylor expansion of a specified form. The bound is expressed combinatorially via quantities (L, n*_s, K) that count the maximum number of items purchasable under budget β, demand α and shelf prices/inventories (m_s, n_s). The formula is obtained by an explicit four-step sequence of blow-ups that produces a normal-crossing form of the Kullback–Leibler divergence. It is then specialized to three-layer neural networks with real-analytic activations, yielding concrete upper bounds (3.2) and (3.4) at two singular strata P1 and P2 for non-polynomial activations (modulo linear-independence hypotheses) and, for polynomial activations, only when the true distribution has no hidden units. When the input dimension is one the numerical values recover previously known exact learning coefficients; for higher input dimension the bounds are consistent with earlier upper bounds of Aoyagi but can be strict (e.g., reduced-rank regression case 1).","tokens_in":39325,"tokens_out":1141,"duration_ms":16891,"significance":"If the Main Theorem and its applications hold, the work supplies the first broadly applicable upper-bound formula for local learning coefficients at singular points of three-layer networks, covering activations such as swish and (under H*=0) polynomials, and thereby extends the exact results of Aoyagi for Vandermonde-type and ReLU singularities as well as the author’s earlier semiregular (nonsingular-point) formula. The budget–demand–supply interpretation and the systematic accounting of how the numbers of weight parameters (r, α, β) and the orders (m_s) enter the coefficient give a transparent geometric picture that is useful for model selection via sBIC and for understanding Bayesian asymptotics of over-parametrized networks. The detailed chart-by-chart blow-up analysis (Appendix C), the genericity lemma for Vandermonde-type Jacobians (Lemma A.1), and the complete worked example (Section 4) constitute solid technical contributions that can be reused for deeper architectures.","major_comments":[{"comment":"Abstract and §3.1 claim that the formula “applies in general settings” for non-polynomial analytic activations, yet the load-bearing linear-independence hypotheses (Main Theorem (iii) and concrete conditions (3.1)/(3.3)) can fail even under Assumption 1, as the paper itself records for swish-type activations with certain true weights (Remark 3.2(3)). The abstract and introduction should state the independence requirement with the same prominence given to the H*=0 restriction for polynomials, so that the scope of the upper bounds (3.2) and (3.4) is not overstated.","section":null},{"comment":"§3.2.1 (reduced-rank regression): after the coordinate change the Main Theorem recovers only three of the four cases of the exact learning coefficient of Aoyagi–Watanabe (2005). The missing case (case 1) shows that the inequality in (2.1) can be strict. Remark 2.2 already notes that equality holds when the Jacobian of condition (ii) is nonsingular for every b\neq0; a short additional paragraph quantifying how often this occurs for the reduced-rank stratum would clarify when the bound is tight versus merely an upper bound.","section":null}],"minor_comments":[{"comment":"Figure 2 caption: “Uppe bound of λ” is missing the letter “r”.","section":null},{"comment":"Notation for multi-indices and the re-indexing of (h,k) into a single index n in Appendix A is dense; a short table summarizing the correspondence between (r,α,β,γ,m_s,n_s) and network dimensions for P1 versus P2 would help the reader.","section":null},{"comment":"In the statement of the Main Theorem the case γ=∞ is handled by a footnote; moving the definition of L into the main text would improve readability.","section":null},{"comment":"Several self-citations to the author’s semiregular papers [19,20] are essential for the nonsingular baseline, but a one-sentence reminder of the precise statement of the earlier formula would make the comparison self-contained.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is technically solid and fills a genuine gap, but it is written at a high algebraic-geometry density for a cs.LG venue. If the journal expects a broader machine-learning audience, the authors might be asked to expand the intuitive budget interpretation and the worked example at the expense of some of the chart-by-chart Jacobian bookkeeping. No concerns about novelty or citation pattern."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The new piece is a Main Theorem that gives an explicit upper bound on the local learning coefficient at a class of singular realization parameters, not just the nonsingular ones covered by the author's earlier semiregular work. The bound is written as a simple budget-demand-supply counting rule in the Taylor degrees and ranks of the log-likelihood ratio, and it recovers Aoyagi's exact N=1 formulas and several of his earlier upper bounds as special cases. That combinatorial reading is clean and useful for anyone who has to compute or compare RLCTs for three-layer analytic nets (including swish).\n\nThe proof is the usual resolution-of-singularities route: four explicit blow-up stages, lemmas that relate f to K, and chart control. Appendices check the hypotheses for the two natural singular strata P1 and P2, and Section 4 walks through a full tanh example so you can see the Jacobians. Citations are appropriate; the self-cites are the nonsingular baseline being improved, not circular.\n\nSoft spots are real but already stated. Linear independence of the activation monomials (condition (iii) and (3.1)/(3.3)) can fail for some swish-type true weights, and the polynomial case is restricted to H*=0. The bound is known to be strict for reduced-rank case 1. None of that breaks the theorem under its hypotheses; it just means the abstract's \"general settings\" claim is accurate only modulo those qualifications.\n\nThis is for people who already work in singular learning theory or need concrete RLCT upper bounds for sBIC-style selection on three-layer nets. It is not a general deep-net result and does not claim to be. The math is careful enough that a serious editor should send it to referees rather than desk-reject. I would read the Main Theorem and the N=1 comparison carefully; the rest is supporting detail.","headline":"Solid upper-bound formula for local RLCT at singular points of three-layer nets; tight when N=1, sometimes loose otherwise, with independence caveats already flagged by the author.","tokens_in":39889,"tokens_out":490,"would_cite":true,"duration_ms":5777,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","14E15","68T07"],"pacs":[],"model":"grok-4.5","headline":"An upper bound for local learning coefficients at singular points of three-layer nets is given by a budget-demand-supply counting rule on the Taylor expansion of the log-likelihood ratio.","keywords":["three-layer neural networks","singular learning theory","real log canonical threshold","local learning coefficient","blow-ups","analytic activation functions","swish","reduced-rank regression"],"falsifier":"Take any three-layer net with N=1 whose exact learning coefficient is already known (e.g., tanh or exponential activations). Compute the right-hand side of the new bound at P1 or P2; if it differs from the known exact value, the Main Theorem is false for that case.","tokens_in":39789,"feed_emoji":"📊","tokens_out":740,"duration_ms":5719,"temperature":0.7,"pith_summary":"Three-layer neural networks are singular statistical models: their Fisher information degenerates, so classical asymptotics and information criteria do not apply. Their Bayesian behavior is controlled by a real number called the local learning coefficient (real log canonical threshold). Previous formulas covered only nonsingular realization parameters and could be far from known exact values. This paper supplies an upper-bound formula that works at a class of singular realization parameters. The bound is assembled from the rank of the Fisher matrix together with the lowest degrees and multiplicities appearing in the Taylor expansion of the log-likelihood ratio; it can be read as the largest number of items one can buy under budget, demand and inventory constraints. The formula applies to general real-analytic activations (including swish and odd analytic functions) and, when the true network has no hidden units, also to polynomial activations. When the input dimension is one the numerical value matches previously known exact coefficients, showing the bound is tight in that case.","feed_headline":"Singular nets get a budget-style upper bound on learning coefficients","feed_subtitle":"A counting rule from the Taylor expansion of the log-likelihood ratio works at singular points and is tight for one-dimensional inputs.","key_machinery":"Main Theorem (equation 2.1): a counting rule under budget-demand-supply constraints obtained by successive blow-ups that produce a normal-crossing form of the Kullback-Leibler divergence. The six integers (r, α, β, γ, (m_s), (n_s)) completely determine the bound.","core_discovery":"Under four explicit conditions on the Taylor expansion of the log-likelihood ratio at a realization parameter P, the local learning coefficient satisfies λ_P ≤ r/2 plus a closed-form expression that counts the maximum number of “items” purchasable under budget β, demand α and successive prices m_s with inventories n*_s. The multiplicity is 2 precisely when the budget is exhausted exactly at a shelf boundary, and 1 otherwise.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Counting-rule upper bounds on local learning coefficients for singular three-layer nets","Budget-demand-supply formula caps learning coefficients at singular realization points","Local learning coefficients of three-layer nets bounded by inventory counting under budget","Upper bounds for singular three-layer net learning coefficients via Taylor expansion count","Formula for local RLCT upper bounds at singular parameters in three-layer networks"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"The random variables built from activation values, first derivatives times inputs, and higher monomials must be linearly independent almost surely; if that independence fails the normal-crossing analysis and the stated upper bound do not apply.","fun_headline_variants_meta":{"raw":{"variants":["Counting-rule upper bounds on local learning coefficients for singular three-layer nets","Budget-demand-supply formula caps learning coefficients at singular realization points","Local learning coefficients of three-layer nets bounded by inventory counting under budget","Upper bounds for singular three-layer net learning coefficients via Taylor expansion counts","Formula for local RLCT upper bounds at singular parameters in three-layer networks"]},"model":"grok-4.5","effort":"low","cost_usd":0.004664,"raw_usage":{"total_tokens":1392,"prompt_tokens":874,"num_sources_used":0,"completion_tokens":99,"cost_in_usd_ticks":46640000,"prompt_tokens_details":{"text_tokens":874,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":419,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":874,"tokens_out":99,"duration_ms":3547,"temperature":1.0,"reasoning_tokens":419,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T22:06:51.262238+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Take any three-layer net with N=1 whose exact learning coefficient is already known (e.g., tanh or exponential activations). Compute the right-hand side of the new bound at P1 or P2; if it differs from the known exact value, the Main Theorem is false for that case.","supporting_citations":[],"review_version":1}