Pith. sign in

REVIEW 2 major objections 4 minor 40 references

Upper Bounds for Local Learning Coefficients of Three-Layer Neural Networks

T0 review · 2 major / 4 minor · reviewed 2026-07-14 · grok-4.5

Pith's one-line read An upper bound for local learning coefficients at singular points of three-layer nets is given by a budget-demand-supply counting rule on the Taylor expansion of the log-likelihood ratio.

desk verdict Solid upper-bound formula for local RLCT at singular points of three-layer nets; tight when N=1, sometimes loose otherwise, with independence caveats already flagged by the author. read the letter →

arxiv 2603.12785 v2 pith:3TXO7I7C submitted 2026-03-13 cs.LG math.STstat.TH

classification cs.LGmath.STstat.TH MSC 62F1514E1568T07
keywords three-layerneuralnetworkssingularlearningtheoryreallogcanonicalthresholdlocalcoefficientblow-upsanalyticactivationfunctionsswishreduced-rankregression
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Three-layer neural networks are singular statistical models: their Fisher information degenerates, so classical asymptotics and information criteria do not apply. Their Bayesian behavior is controlled by a real number called the local learning coefficient (real log canonical threshold). Previous formulas covered only nonsingular realization parameters and could be far from known exact values. This paper supplies an upper-bound formula that works at a class of singular realization parameters. The bound is assembled from the rank of the Fisher matrix together with the lowest degrees and multiplicities appearing in the Taylor expansion of the log-likelihood ratio; it can be read as the largest number of items one can buy under budget, demand and inventory constraints. The formula applies to general real-analytic activations (including swish and odd analytic functions) and, when the true network has no hidden units, also to polynomial activations. When the input dimension is one the numerical value matches previously known exact coefficients, showing the bound is tight in that case.

What carries the argument

Main Theorem (equation 2.1): a counting rule under budget-demand-supply constraints obtained by successive blow-ups that produce a normal-crossing form of the Kullback-Leibler divergence. The six integers (r, α, β, γ, (m_s), (n_s)) completely determine the bound.

What would settle it

Take any three-layer net with N=1 whose exact learning coefficient is already known (e.g., tanh or exponential activations). Compute the right-hand side of the new bound at P1 or P2; if it differs from the known exact value, the Main Theorem is false for that case.

Watch

Extended reading notes

Core claim

Under four explicit conditions on the Taylor expansion of the log-likelihood ratio at a realization parameter P, the local learning coefficient satisfies λ_P ≤ r/2 plus a closed-form expression that counts the maximum number of “items” purchasable under budget β, demand α and successive prices m_s with inventories n*_s. The multiplicity is 2 precisely when the budget is exhausted exactly at a shelf boundary, and 1 otherwise.

Load-bearing premise

The random variables built from activation values, first derivatives times inputs, and higher monomials must be linearly independent almost surely; if that independence fails the normal-crossing analysis and the stated upper bound do not apply.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper derives an upper-bound formula (Main Theorem, Eq. 2.1) for the local learning coefficient (real log canonical threshold) at a class of singular realization parameters of statistical models whose log-likelihood ratio admits a Taylor expansion of a specified form. The bound is expressed combinatorially via quantities (L, n*_s, K) that count the maximum number of items purchasable under budget β, demand α and shelf prices/inventories (m_s, n_s). The formula is obtained by an explicit four-step sequence of blow-ups that produces a normal-crossing form of the Kullback–Leibler divergence. It is then specialized to three-layer neural networks with real-analytic activations, yielding concrete upper bounds (3.2) and (3.4) at two singular strata P1 and P2 for non-polynomial activations (modulo linear-independence hypotheses) and, for polynomial activations, only when the true distribution has no hidden units. When the input dimension is one the numerical values recover previously known exact learning coefficients; for higher input dimension the bounds are consistent with earlier upper bounds of Aoyagi but can be strict (e.g., reduced-rank regression case 1).

Significance. If the Main Theorem and its applications hold, the work supplies the first broadly applicable upper-bound formula for local learning coefficients at singular points of three-layer networks, covering activations such as swish and (under H*=0) polynomials, and thereby extends the exact results of Aoyagi for Vandermonde-type and ReLU singularities as well as the author’s earlier semiregular (nonsingular-point) formula. The budget–demand–supply interpretation and the systematic accounting of how the numbers of weight parameters (r, α, β) and the orders (m_s) enter the coefficient give a transparent geometric picture that is useful for model selection via sBIC and for understanding Bayesian asymptotics of over-parametrized networks. The detailed chart-by-chart blow-up analysis (Appendix C), the genericity lemma for Vandermonde-type Jacobians (Lemma A.1), and the complete worked example (Section 4) constitute solid technical contributions that can be reused for deeper architectures.

major comments (2)
  1. Abstract and §3.1 claim that the formula “applies in general settings” for non-polynomial analytic activations, yet the load-bearing linear-independence hypotheses (Main Theorem (iii) and concrete conditions (3.1)/(3.3)) can fail even under Assumption 1, as the paper itself records for swish-type activations with certain true weights (Remark 3.2(3)). The abstract and introduction should state the independence requirement with the same prominence given to the H*=0 restriction for polynomials, so that the scope of the upper bounds (3.2) and (3.4) is not overstated.
  2. §3.2.1 (reduced-rank regression): after the coordinate change the Main Theorem recovers only three of the four cases of the exact learning coefficient of Aoyagi–Watanabe (2005). The missing case (case 1) shows that the inequality in (2.1) can be strict. Remark 2.2 already notes that equality holds when the Jacobian of condition (ii) is nonsingular for every b eq0; a short additional paragraph quantifying how often this occurs for the reduced-rank stratum would clarify when the bound is tight versus merely an upper bound.
minor comments (4)
  1. Figure 2 caption: “Uppe bound of λ” is missing the letter “r”.
  2. Notation for multi-indices and the re-indexing of (h,k) into a single index n in Appendix A is dense; a short table summarizing the correspondence between (r,α,β,γ,m_s,n_s) and network dimensions for P1 versus P2 would help the reader.
  3. In the statement of the Main Theorem the case γ=∞ is handled by a footnote; moving the definition of L into the main text would improve readability.
  4. Several self-citations to the author’s semiregular papers [19,20] are essential for the nonsingular baseline, but a one-sentence reminder of the precise statement of the earlier formula would make the comparison self-contained.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the Main Theorem upper bound is derived from an explicit four-step blow-up sequence under stated analytic and linear-independence hypotheses, not by construction from the target coefficient or a load-bearing self-citation.

full rationale

The derivation chain is self-contained mathematical analysis. The Main Theorem (eq. 2.1) is obtained by performing coordinate transformations (blow-ups CT1–CT9 and the a↦a' change of Step 2) on the Taylor expansion of the log-likelihood ratio f, producing normal-crossing forms of K whose real log canonical thresholds are bounded by the budget–demand–supply expression; the proof outline (Section 5) and full chart-by-chart calculation (Appendix C) do not presuppose the value of λ_P. Conditions (i)–(iv) and the concrete linear-independence hypotheses (3.1)/(3.3) are assumptions under which the bound holds; the paper itself records the cases in which they fail (Remark 3.2(3), H*=0 restriction for polynomials) and the cases in which the inequality is strict (reduced-rank case 1). Self-citations to the author’s semiregular papers [19,20] supply the nonsingular baseline being improved and some preparatory lemmas (e.g., Lemma B.1), but are not used to force the singular-point formula. Agreement with Aoyagi’s exact N=1 results is an external consistency check, not a circular reduction. No fitted parameters, no uniqueness theorem imported from the same authors, and no renaming of a known empirical pattern appear. The single minor self-citation does not raise the score above 1.

Assumptions & free parameters 0 free parameters · 6 assumptions · 2 invented entities

The result rests on standard resolution of singularities, analyticity of the log-likelihood ratio, realizability of the true distribution, interchange of expectation and derivatives, and several linear-independence / Jacobian-rank conditions that are domain assumptions for the blow-up analysis. No free parameters are fitted. No new physical entities are postulated; the invented objects are purely mathematical (the counting quantities L, n*_s, K and the singular strata P1, P2).

assumptions (6)
  • standard math Hironaka resolution of singularities (Theorem 1.1): real-analytic K admits a proper analytic map to normal-crossing form.
    Used to define local learning coefficients λ_P and to justify the blow-up strategy in the Main Theorem proof.
  • domain assumption f and K are real-analytic near realization parameters; expectation and θ-derivatives may be interchanged.
    Stated in Notation and assumptions; required for Taylor expansions and for relating derivatives of K to moments of derivatives of f.
  • domain assumption Model is realizable: there exists θ* with q = p(·|θ*) a.s.; prior positive at realization parameters.
    Standard singular learning setup; needed so Θ* is nonempty and local coefficients are well-defined.
  • ad hoc to paper Main Theorem conditions (i)–(iv): g_{s,n}(a=0,·)=0; Jacobian rank of lowest-degree terms equals ∑ n*_s for some b≠0; linear independence of (Z_{s,n}) and of Fisher score directions; higher terms h_s lie in the ideal generated by the g_{s,n}.
    These are the precise hypotheses under which the blow-up sequence yields (2.1); verified for three-layer nets in Appendix A under further linear-independence assumptions.
  • ad hoc to paper For polynomial activations, the true distribution has no hidden units (H*=0).
    Explicit restriction in abstract and §3.2; without it the polynomial case is not claimed.
  • domain assumption Linear independence of activation values, first derivatives times inputs, and monomials of degrees m_s (conditions (3.1), (3.3)); Proposition A.1 gives sufficient conditions.
    Needed to identify r and the Z_{s,n} and to apply condition (iii); paper notes failures for some swish-type weight configurations.
invented entities (2)
  • Budget–demand–supply counting quantities (L, n*_s, K) and the associated upper-bound formula (2.1)
    purpose: Package the outcome of the blow-up sequence as an intuitive combinatorial rule and state the upper bound on λ_P.
    Purely mathematical bookkeeping derived from the Taylor degrees m_s and ranks n_s; no independent physical content.
  • Singular realization strata P1 and P2 for three-layer nets independent evidence
    purpose: Identify concrete singular points where the Main Theorem applies and compare upper bounds across activation types.
    Standard constructions in the literature (zero redundant weights vs duplicated true b*); not new particles or forces.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Upper Bounds for Local Learning Coefficients of Three-Layer Neural Networks." pith.science (2026). https://pith.science/paper/3TXO7I7C

@misc{pith2026260312785,
  author       = {Pith},
  title        = {Pith review of: Upper Bounds for Local Learning Coefficients of Three-Layer Neural Networks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3TXO7I7C}},
  note         = {Machine review of arXiv:2603.12785}
}
read the original abstract

Three-layer neural networks are known to form singular learning models, and their Bayesian asymptotic behavior is governed by the learning coefficient, or real log canonical threshold. Although this quantity has been clarified for regular models and for some special singular models, broadly applicable methods for evaluating it in neural networks remain limited. Recently, a formula for the local learning coefficient of semiregular models was proposed, yielding an upper bound on the learning coefficient. However, this formula applies only to nonsingular points in the set of realization parameters and cannot be used at singular points. In particular, for three-layer neural networks, the resulting upper bound has been shown to differ substantially from learning coefficient values already known in some cases. In this paper, we derive a formula for an upper bound on local learning coefficients at a class of singular realization parameters in three-layer neural networks. This formula can be interpreted as a counting rule under budget, demand, and supply constraints. In the non-polynomial real-analytic case, the formula applies in general settings, whereas in the polynomial case it applies under the restriction that the true distribution has no hidden units. In particular, our result covers activation functions such as the swish function and also includes polynomial activation functions under the above restriction, thereby extending previous results to a broader class of activation functions. We further show that, when the input dimension is one, the numerical value given by the right-hand side of our upper-bound formula agrees with the previously known learning coefficient, thereby providing a useful comparison with known exact results. Our result also provides a systematic perspective on how the weight parameters of three-layer neural networks affect the learning coefficient.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

40 extracted references · 3 linked inside Pith

  1. [1]

    Akaike, H. (1974). A new look at the statistical model ide ntification. IEEE Transactions on Automatic Control , 19, 716-723

  2. [2]

    Aoyagi, M. (2006). The zeta function of learning theory a nd generalization error of three layered neural perceptron. RIMS Kokyuroku, Recent Topics on Real and Complex Singularities , 1501, 153-167

  3. [3]

    Aoyagi, M. (2009). Log canonical threshold of Vandermon de matrix type singularities and generalization error of a three-layered neural network in Bayesian estimation. International Journal of Pure and Applied Mathemat- ics, 52, 177-204

  4. [4]

    Aoyagi, M. (2010). A Bayesian learning coefficient of gene ralization error and Vandermonde matrix-type singularities. Communications in Statistics—Theory and Methods , 39, 2667-2687

  5. [5]

    Aoyagi, M. (2013a). Consideration on singularities in l earning theory and the learning coefficient. Entropy, 15, 3714-3733

  6. [6]

    Aoyagi, M. (2013). Learning coefficient in Bayesian estim ation of restricted Boltzmann machine. Journal of Algebraic Statistics , 4, 30-57

  7. [7]

    Aoyagi, M. (2019a). Learning coefficient of Vandermonde m atrix-type sin- gularities in model selection. Entropy, 21, 1-12

  8. [8]

    Aoyagi, M. (2019b). Learning coefficients and informatio n criteria. Fron- tiers in Artificial Intelligence and Applications , 351-362

Show all 40 references
  1. [9]

    Aoyagi, M. (2024). Consideration on the learning efficien cy of multiple- layered neural networks with linear units. Neural Networks, 172, 106132

  2. [10]

    Aoyagi, M. (2025). Singular learning coefficients and effi ciency in learning theory. arXiv:2501.12747

  3. [11]

    Aoyagi, M., & Watanabe, S. (2005). Resolution of singul arities and the generalization error with Bayesian estimation for layered neural network. IEICE Transactions J88-D-II , 10, 2112-2124

  4. [12]

    Aoyagi, M., & Watanabe, S. (2005). Stochastic complexi ties of reduced rank regression in Bayesian estimation. Neural Networks, 18, 924-933

  5. [13]

    W., & Salamon, D

    Robbin, J. W., & Salamon, D. A. (2000). The exponential V andermonde matrix. Linear Algebra and its Applications , 317, 225-226. 23

  6. [14]

    Drton, M., & Plummer, M. (2017). A Bayesian information criterion for singular models. Journal of the Royal Statistical Society Series B: Statistica l Methodology, 79, 323-380

  7. [15]

    Drton, M., Lin, S., Weihs, L., & Zwiernik, P. (2017). Mar ginal likelihood and model selection for Gaussian latent tree and forest mode ls. Bernoulli, 23, 1202-1232

  8. [16]

    Hironaka, H. (1964). Resolution of singularities of an algebraic variety over a field of characteristic zero. Annals of Math , 79, 109-326

  9. [17]

    Imai, T. (2019). Estimating real log canonical thresho lds. arXiv:1906.01341

  10. [18]

    Kashiwara, M. (1976). B-functions and holonomic syste ms. Inventiones Mathematicae, 38, 33-53

  11. [19]

    Kurumadani, Y. (2025). Learning coefficients in semireg ular models I: prop- erties. Japanese Journal of Statistics and Data Science , 8, 1051-1079

  12. [20]

    Kurumadani, Y. (2025). Learning coefficients in semireg ular models II: ex- tensions. Japanese Journal of Statistics and Data Science

  13. [21]

    Lau, E., Furman, Z., Wang, G., Murfet, D., & Wei, S. (2025 ). The lo- cal learning coefficient: A singularity-aware complexity me asure. In Pro- ceedings of the 28th International Conference on Artificial Inte lligence and Statistics (AISTATS 2025) , Proceedings of Machine Learni...

  14. [22]

    Mustata, M. (2002). Singularities of pairs via jet sche mes. Journal of the American Mathematical Society , 15, 599-615

  15. [23]

    Rusakov, D., & Geiger, D. (2002). Asymptotic model sele ction for naive Bayesian networks. In Proceedings of the Eighteenth Conference on Uncer- tainty in Artificial Intelligence , 438-445

  16. [24]

    Rusakov, D., & Geiger, D. (2005). Asymptotic model sele ction for naive Bayesian networks. Journal of Machine Learning Research , 6, 1-35

  17. [25]

    Sato, K., & Watanabe, S. (2019). Bayesian generalizati on error of Poisson mixture and simplex Vandermonde matrix type singul arity. arXiv:1912.13289

  18. [26]

    Schwarz, G. (1978). Estimating the dimension of a model . The Annals of Statistics, 6, 461-464

  19. [27]

    Takeuchi, K. (1976). Distribution of an information st atistic and the crite- rion for the optimal model. Mathematical Science, 153, 12-18

  20. [28]

    Watanabe, S. (2001a). Algebraic analysis for nonident ifiable learning ma- chines. Neural Computation, 13, 899-933. 24

  21. [29]

    Watanabe, S. (2001b). Algebraic geometrical methods f or hierarchical learning machines. Neural Networks, 14, 1049-1060

  22. [30]

    Watanabe, S. (2001c). Algebraic geometry of learning m achines with sin- gularities and their prior distributions. Journal of Japanese Society of Ar- tificial Intelligence , 16, 308-315

  23. [31]

    Watanabe, S., Yamazaki, K., & Aoyagi, M. (2004). Kullba ck information of normal mixture is not an analytic function. Technical Report of IEICE (in Japanese)

  24. [32]

    Watanabe, S. (2009). Algebraic geometry and statistical learning theory, vol. 25 . New York, USA: Cambridge University Press

  25. [33]

    Watanabe, S. (2010). Equations of states in singular st atistical estimation. Neural Networks , 23, 20-34

  26. [34]

    Watanabe, S. (2013). A widely applicable Bayesian info rmation criterion. Journal of Machine Learning Research , 14, 867-897

  27. [35]

    Watanabe, S. (2018). Mathematical theory of Bayesian statistics . Boca Ra- ton, FL, USA: CRC Press

  28. [36]

    Watanabe, T., & Watanabe, S. (2022). Asymptotic behavi or of Bayesian generalization error in multinomial mixtures. arXiv:2203 .06884

  29. [37]

    Zwiernik, P. (2011). An asymptotic behavior of the marg inal likelihood for general Markov models. Journal of Machine Learning Research , 12, 3283- 3310. Appendix A. V erification of the Assumptions of the Main Theor em In this section, we verify that the three-layer neural net...

  30. [38]

    , θ ′′ r , a ′′ 1, 1,

    We now denote the transformed coordinates θ′′ 1 , . . . , θ ′′ r , a ′′ 1, 1, . . . , a ′′ 1,n ∗ 1 again by θ′ 1, . . . , θ ′ r, a ′ 1, 1, . . . , a ′ 1,n ∗ 1 . Applying CT5, {θ′ j → θ′ 1θ′′ j , a ′ 1,n → θ′ 1a′′ 1,n , b 1 → θ′ 1b′ 1 |2 ≤ j ≤ r, 1 ≤ n ≤ n∗ 1}, we obtain f =θ′m...

  31. [39]

    The above argument was given for CT5, but the same conclusion holds when CT6 is applied instead

    In this case, the multiplicity is m = 2; otherwise, m = 1. The above argument was given for CT5, but the same conclusion holds when CT6 is applied instead. 40 Apply CT7 m2 − m1 times Proceeding as above, we apply the coordinate transformatio n π = {θ′ j → bm2− m1 1 θ′′ j , a ′...

  32. [40]

    This case will be considered in the next step. Summarizing, among the normal crossings obtained in Step 3- 1, we have inf Q min j h(Q) j + 1 k(Q) j = min { m1r + β 2m1 , m2r + β + (m2 − m1)n∗ 1 2m2 } , and the multiplicity is m = 2 when β = m1n∗ 1, and m = 1 otherwise. We cont...

Pith tools

Reviewed July 14, 2026 · model on record in the stance chip above.