{"id":"262a17fd-9647-41c3-8254-3937b41d0393","arxiv_id":"2607.23397","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Even for wide three-layer nets, hierarchical learning of both layers yields a smaller asymptotic generalization error than single-layer learning with a fixed random kernel, because of singularities.","lead":"Training the first-layer weights of a wide three-layer net gives a strictly smaller Bayesian generalization error than freezing them as a fixed kernel, even as width grows large. The gap is traced to singularities in the hierarchical parameter space that fixed-kernel (lazy) learning never sees.","discovery_kind":"extension","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"Eq. (28) relies on an unjustified C(h)→C1(M) replacement; the limit is not uniform over the admissible maximizers, so the stated large-H bound is not fully proved.","rationale":"The zero-function noise model is a major scope limitation, as the reader says, but the strongest claim is explicitly conditioned on that setting; it is therefore not by itself an internal correctness objection. The fixed-feature calculation appears comparatively direct, and the singular-learning-theory framework plus the exact M=N=1 check provide meaningful independent support. The more consequential unresolved point is the nonuniform C(h)→C1(M) step in the new general-M bound. The reader’s rationale and proposed condition already gesture toward this finite-H issue, so agreement is partial rather than complete: the reader’s named weakest assumption is the noise-only specification, while the load-bearing proof concern is the unjustified asymptotic replacement. Since the reader’s CONDITIONAL verdict already reflects the need to spell out the finite-H error and clarify the bound, this stress test does not warrant changing the verdict.","tokens_in":9831,"tokens_out":13759,"duration_ms":454548,"concrete_test":"Independently re-derive Eq. (56) from Eq. (55) for arbitrary fixed M without replacing C by C1. With h=s/t^M and q(h) defined by Eqs. (38)–(39), prove a localization for the maximum and show t^M max_{M≤h≤H}{0,h−C(h)h^{(M+1)/M}t}=max_{s≥0}{0,s−C1(M)s^{(M+1)/M}}+o(1), with an explicit error bound. Substitute t=t*(H) and check whether the resulting Λ is ≤α(M)H^{M/(M+1)}N^{1/(M+1)}. If the o(1) term cannot be bounded below the available α slack, Eq. (28) is not established as stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Equation (49) gives Δ(h) ≥ C(h)h^{(M+1)/M}, where C(h)=C0(q(h))C1(M) and C0(q)<1 for finite q. Thus Eq. (55) still contains max{0,s−C(s/t^M)s^{(M+1)/M}}. In passing to Eq. (56), the paper replaces C(s/t^M) by C1(M). Since C≤C1, however, s−Cs^{(M+1)/M} ≥ s−C1s^{(M+1)/M}; the replacement moves the expression in the wrong direction for an upper bound. The asserted uniform convergence over s≥M(t*)^M is also not literally true: at the lower endpoint s=M(t*)^M, h=M remains fixed as t*→0, so q(h) and C0(q(h)) do not approach their large-q limits. Convergence does hold at every fixed interior s because h→∞, but the maximizing h may drift with H, so pointwise convergence alone does not justify exchanging the limit and maximum. The intended limiting calculation is plausible, and α(M) may have enough slack to absorb the finite-H correction, but the paper does not prove that. Because Theorem 2’s comparison with the exact HN/(2n) rate depends on this leading large-H coefficient, this is the load-bearing rigor gap.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper compares two Bayesian learning schemes for a three-layer tanh network (M inputs, H hidden units, N outputs) when the data-generating conditional is the zero function plus isotropic Gaussian noise (eq. 23). In \"single-layer learning\" the input-to-hidden weights b are drawn from the prior and frozen (the kernel/lazy regime); in \"hierarchical learning\" both layers are trained (feature learning). Section 4 computes the single-layer generalization error exactly to leading order, G1(n) = HN/(2n) + o(1/n) (eq. 27), via the Gaussian integral over the last layer and an eigenvalue identity (eq. 35). Section 5 upper-bounds the hierarchical error through the real log canonical threshold, using the bound λ ≤ Λ from the author's earlier work and a new lower bound on the growing index Δ(k) ≥ C(k) k^{(M+1)/M} (eq. 49), obtaining G2(n) ≤ α(M) H^{M/(M+1)} N^{1/(M+1)}/n + o(1/n) with an explicit α(M) (Theorem 2, eq. 28). Since the exponent of H in the second bound is below 1, hierarchical learning is asymptotically superior for generic M, N in the regime M, N ≪ H ≪ n, and the paper concludes that singularities remain operative even at large width.","tokens_in":10022,"tokens_out":7751,"duration_ms":229212,"significance":"If correct, this is a clean, parameter-free asymptotic separation between the fixed-kernel (NNT/GP-like) and feature-learning regimes at large but finite width — a question of current interest given the lazy-vs-feature-learning literature. Strengths: the single-layer calculation is exact to leading order with no fitted constants; the hierarchical bound has a fully explicit coefficient α(M) (with a correct M→∞ expansion, α→e/2, and consistency with the known tight M=N=1 RLCT of [6]); the key inputs from prior singular learning theory (λ ≤ Λ, the Δ(k) machinery) enter as cited theorems with stated hypotheses, not as quantities tuned to the target; and the result yields concrete falsifiable numbers (eqs. 29-30: 500000/n vs ≲13737/n at M=N=100, H=10^4) that simulations could check. The qualitative message — that the infinite-width equivalence to kernel regression is limit-order dependent and that singularity-driven efficiency persists at width — is a genuinely useful theoretical data point.","major_comments":[{"comment":"The passage from eq. (55) to eq. (56) replaces C(s/t*^M) by C1(M), but the justification given is incorrect in two respects, and the step is load-bearing because α(M) — hence the quantitative comparison in eqs. (29)-(30) — depends on it. (i) Direction of the inequality: since C0(q) < 1 for finite q, C(h) < C1(M), so s − C(h)s^{(M+1)/M} ≥ s − C1(M)s^{(M+1)/M}; substituting C1(M) *lowers* the maximand, which is not a valid upper bound for max{0, s − C(h)s^{(M+1)/M}} unless a limit-interchange argument is supplied. (ii) The claimed uniform convergence of C(s/t*^M) to C1(M) on s ≥ M(t*)^M is false at the lower endpoint: at s = M(t*)^M one has h = M fixed as t*→0, so q(h) and C0(q(h)) do not approach their large-q limits. The step is likely repairable — C0 is monotone increasing in q (§5.1), so convergence is uniform on compact s-intervals bounded away from 0; for s ≥ C(M)^{-M} the maximand i","section":"§5.2, eqs. (55)-(56)"},{"comment":"Theorem 2 is stated in the joint regime M, N ≪ H ≪ n, but both ingredient asymptotics are proved only as fixed-model limits. (a) Eq. (36)-(37), G2(n) = λ/n + o(1/n), is the standard RLCT asymptotic for a *fixed* model as n→∞; no uniformity of the o(1/n) remainder in H is established, yet the conclusion is applied with H growing with n. (b) On the single-layer side, eq. (35) requires Σ_h E_b[1/(nν_h(b)+1)] = o(1) for the leading term NH/(2n) to survive; for fixed H this follows from ν_h(b) > 0 a.s. (linear independence of the tanh features) and dominated convergence, but with H → ∞ the smallest eigenvalues of the H×H feature Gram matrix can drift toward 0, and the remainder is not controlled. Please state the precise order of limits (e.g., n→∞ first, then H, or H = H(n) with a rate condition) and bound the remainders accordingly; without this, the comparison G2 ≪ G1 in the stated regime i","section":"§3, Theorem 2; §4 eq. (35); §5 eqs. (36)-(37)"}],"minor_comments":[{"comment":"Statement of Theorem 2: \"G2(b)\" should read \"G2(n)\".","section":"§3, Theorem 2"},{"comment":"First line: \"the growing index δ(k)\" should be Δ(k); two lines below eq. (42), \"If follows that\" → \"It follows that\".","section":"§5.1"},{"comment":"The maximization over 0 ≤ h in the definition of Λ is replaced by a maximum over M ≤ h in eq. (53) without comment. For h < M one has q = 0 and Δ(h) = 2h, so h − Δ(h)t = h(1−2t) ≤ M, contributing only an O(1) additive constant after the N/(2t^M) scaling in eq. (54) — negligible relative to the diverging main term. Please add this one-line justification.","section":"§5.2, eq. (53)"},{"comment":"The symbol I(b) is used for both the empirical Gram matrix (defined after eq. (24)) and its population expectation (\"Let I(b) be the expectation of I(b)\"). Please use distinct symbols and state explicitly the concentration step I(b) → I(b) and the sense in which V1(X)² and V2(X) are lower order than V1(X).","section":"§4, eqs. (31)-(35)"},{"comment":"Please double-check the constants in eq. (31), in particular the term tr(J(b)−1) and the prefactor NH/2; a one-line derivation of the Gaussian integral would help the reader verify them.","section":"§4, eq. (31)"},{"comment":"The abstract and §1 state the conclusion (training the first layer generalizes better) without noting that Theorem 2 is proved only for the noise-only realizable target q(y|x) ∝ exp(−∥y∥²/2). The decomposition (21)-(22) and the assertion G^n_1 ∼ G^n_2 are plausibility arguments, not part of the theorem; please say so explicitly and scope the abstract claim accordingly.","section":"Abstract; §2, eqs. (21)-(22)"},{"comment":"Consistency check worth remarking on: for M = 1 the general bound gives α(1) = √2/2, i.e., Λ ≤ (√2/2)√(HN), whereas [19] and eq. (40) give the tight Λ = √(HN)/2. The general-M bound is therefore loose by a factor √2 in the one case where λ is known exactly; a sentence noting this would calibrate how sharp eq. (28) is likely to be.","section":"§5, Example after eq. (40)"},{"comment":"The AM-GM step in eq. (45) is correct (the M−1 factors 2p,…,2p+M−2 have arithmetic mean 2p + M/2 − 1) but compressed; spelling out the mean computation would save the reader a verification.","section":"§5.1, eq. (45)"},{"comment":"It would strengthen the positioning to relate the H^{M/(M+1)}-vs-H scaling to known sample-complexity separations between kernel and feature-learning regimes (e.g., the works cited as [8, 10, 1]), since the present result is a Bayesian-average-case analog of those worst-case/gradient-based results.","section":"§6, Discussion"}],"recommendation":"major_revision","confidential_remarks":"Single-author manuscript situated almost entirely within the author's own singular-learning-theory literature (five of the technical citations are self-citations, including the two load-bearing inputs λ ≤ Λ and the growing-index machinery from [19, 20]). That framework is established and appropriate here, and the general-M evaluation of the Λ bound is a genuine extension rather than a repackaging. The gap at eqs. (55)-(56) appears repairable with a routine compactness/monotonicity argument, so I do not read it as fatal; but as submitted the proof of Theorem 2's constant is incomplete, and the editor may want the revision to state the order-of-limits convention explicitly. Fit with the journal seems good if the venue welcomes Bayesian/singular-learning theory."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: Watanabe gives an explicit asymptotic gap between fixed-b (single-layer / lazy kernel) and joint (a,b) Bayesian training of a three-layer tanh net on pure noise. Under M,N ≪ H ≪ n you get G1 ∼ HN/(2n) exactly at leading order, versus an upper bound G2 = O(H^{M/(M+1)} N^{1/(M+1)} / n). That is the quantitative claim that singularities still matter for wide nets, and it is new as a width-dependent statistical comparison.\n\nWhat works. Section 4 is standard and clean: Gaussian integral over a, eigenvalue identity, G1 = HN/(2n) + o(1/n). The hierarchical side correctly invokes the author’s prior Λ bound and extends the growing-index estimate Δ(k) to general M via integral comparisons; for M=N=1 it recovers the known tight order. Framing against Barron (approximation) and against NTK/lazy regimes is honest and useful. Circularity is low—this is parameter-free asymptotics inside established singular learning theory, not a fitted story.\n\nSoft spots, in proportion. (1) Only an upper bound on λ for general M,N; tightness is open except M=1. (2) The pure-noise / zero-target setting isolates statistical error and deliberately sets approximation error aside—so the claimed gap is proved only there. (3) The stress-test on eqs. (55)–(56) is right: C(h) ≤ C1(M), so replacing C by C1 shrinks the max and goes the wrong way for an upper bound; uniformity over the maximizer is asserted, not proved. The limiting coefficient α(M) is still plausible (maximizer of the C1 problem sits at fixed s while h→∞), but the write-up does not close the finite-H error. That is a load-bearing gap for the stated leading constant, not a quibble about o(1/n).\n\nWho it is for: people who already read Watanabe / Aoyagi and care about RLCT vs NTK. Not a general DL audience; no experiments, no code. It deserves a serious referee who will force the C→C1 step to be written carefully and the scope (noise-only, upper bound) to stay explicit. I would bring it to an SLT reading group, cite the comparison when discussing lazy vs feature learning from the Bayesian side, and send it to review rather than desk-reject.","headline":"Clean SLT comparison showing hierarchical Bayesian training beats fixed-kernel single-layer even at large width, with an explicit rate gap—but the large-H coefficient proof has a real rigor hole.","tokens_in":11291,"tokens_out":626,"would_cite":true,"duration_ms":32890,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","68T07","14B05"],"pacs":[],"model":"grok-4.5","headline":"Even in wide nets, learning the first-layer weights beats a fixed random kernel, and singularities are why.","keywords":["singular learning theory","real log canonical threshold","wide neural networks","kernel regime","feature learning","generalization error","three-layer networks","Bayesian free energy"],"falsifier":"Compute or tightly bound the exact real log canonical threshold (or finite-n generalization error) for the same three-layer tanh network on a nonzero smooth target; if the hierarchical advantage disappears or reverses while M, N ≪ H ≪ n still holds, the claimed statistical gap fails.","tokens_in":11065,"feed_emoji":"📉","tokens_out":887,"duration_ms":16001,"temperature":0.7,"pith_summary":"The paper compares two ways of training a three-layer neural network with many hidden units: keep the input-to-hidden weights frozen at random initialization and train only the output weights, or train both layers. For learning a pure-noise target (the zero function plus Gaussian noise), it proves that the hierarchical scheme has a strictly smaller asymptotic generalization error than the fixed-kernel scheme once the width is large but still finite. The gap grows with width: fixed kernels pay a linear cost in the number of hidden units, while hierarchical learning pays only a sub-linear cost set by the real log canonical threshold of the singular parameter space. The result is offered as evidence that the two common infinite-width pictures—lazy kernels versus feature learning—are not equivalent, and that singularities remain essential even when networks are wide.","feed_headline":"Wide nets still need singularities: fixed kernels lose","feed_subtitle":"Training first-layer weights beats a frozen random basis by a growing statistical gap","key_machinery":"The real log canonical threshold λ of the hierarchical model, bounded above via the growing index Δ(h) that counts the vanishing order of the tanh basis; the free-energy difference F1(n) − F2(n) then forces G1(n) ≥ G2(n) with the stated rates.","core_discovery":"Under the regime M, N ≪ H ≪ n, Bayesian generalization error for single-layer (fixed-b) learning of the zero map equals HN/(2n) + o(1/n), while hierarchical learning satisfies the tighter upper bound α(M) H^{M/(M+1)} N^{1/(M+1)} / n + o(1/n). Thus training the input-to-hidden weights yields a smaller statistical error than keeping them fixed, even though more parameters are optimized.","pith_inferences":["The same singularity-driven gap should appear in deeper architectures once the first-layer weights are allowed to move, suggesting that depth alone does not eliminate the need for singular learning theory.","Practical kernel methods that freeze early layers may systematically under-perform end-to-end training on pure statistical grounds, independent of optimization or approximation issues.","Measuring posterior mass near singular strata as width grows would give an empirical test of whether the real-log-canonical-threshold scaling is visible in finite networks."],"forward_implications":["Fixed-kernel (lazy) and feature-learning infinite-width limits remain statistically inequivalent even as width tends to infinity.","Generalization error of hierarchical nets scales sub-linearly in width, so wider networks improve the hierarchical advantage rather than erase it.","Phase transitions in the posterior support, driven by singularities, are predicted to persist in wide networks and to improve both approximation and estimation.","The same gap is expected for any odd analytic activation whose Taylor series begins with odd powers."],"fun_headline_variants":["Wide nets: training first layer beats fixed kernels","Hierarchical learning cuts error vs frozen wide nets","Singularities matter: fixed kernels lag in wide nets","Training input weights tops single-layer in wide nets","Wide nets still need singularities for tighter bounds"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The true map being learned is exactly the zero function plus isotropic Gaussian noise, so the comparison isolates only statistical error under perfect realizability and does not yet control approximation error for nonzero targets.","fun_headline_variants_meta":{"raw":{"variants":["Wide nets: training first layer beats fixed kernels","Hierarchical learning cuts error vs frozen wide nets","Singularities matter: fixed kernels lag in wide nets","Training input weights tops single-layer in wide nets","Wide nets still need singularities for tighter bounds"]},"model":"grok-4.5","effort":"low","cost_usd":0.003704,"raw_usage":{"total_tokens":1121,"prompt_tokens":701,"num_sources_used":0,"completion_tokens":58,"cost_in_usd_ticks":37044000,"prompt_tokens_details":{"text_tokens":701,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":362,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":701,"tokens_out":58,"duration_ms":6692,"temperature":1.0,"reasoning_tokens":362,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T23:15:33.527032+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Compute or tightly bound the exact real log canonical threshold (or finite-n generalization error) for the same three-layer tanh network on a nonzero smooth target; if the hierarchical advantage disappears or reverses while M, N ≪ H ≪ n still holds, the claimed statistical gap fails.","supporting_citations":[],"review_version":1}