{"id":"63e493da-5c4c-416a-9cfb-c545ef6287b5","arxiv_id":"2607.26065","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Misspecified kernel ridge regression attains the minimax L2 rate over Hölder-Zygmund classes, but its Hölder-Zygmund norm of the noise component diverges as log n.","lead":"This paper proves that kernel ridge regression with a Sobolev-smooth kernel is minimax optimal for regression over Hölder-Zygmund function classes, and shows that its estimates can be rougher than the target: the Hölder norm of the noise component grows like log n even when the true function is zero.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 2.1 rests on Lemma 2.3 from [7] quoted without hypotheses; until those hypotheses are checked for Sobolev-equivalent kernels, the minimax upper bound is not established.","rationale":"I read the paper in good faith. The properness counterexample (Theorem 3.2) appears internally coherent: the dual representation, Gram concentration, diagonal variance lower bound, and Sudakov packing argument all check out, and the upper bound via Gaussian maxima is sound. The minimax optimality proof is plausible, but it relies on a quoted concentration inequality from [7] whose hypotheses are not stated. The reader's weakest_assumption identifies exactly this point. I also noticed a small norm-weight typo in the approximation-error display, but it is harmless because the corrected weight yields the claimed N-scale bound. Since the main unresolved issue is an omitted verification of an imported lemma, the appropriate disposition is CONDITIONAL, which is the reader's verdict; my stress-test does not move it.","tokens_in":7186,"tokens_out":26146,"duration_ms":214462,"concrete_test":"Locate the concentration lemma in Zhang, Li, and Lin (2024) that is quoted as Lemma 2.3 here, and verify (i) its precise assumptions on kernel eigenvalue decay and eigenbasis, (ii) the source condition required on f*, (iii) whether α→d/(2s+d) and s'→2s/(2s+d) with λ=1/n are admissible, and (iv) whether the conclusion holds for any kernel whose RKHS is equivalent to H^{s+d/2}. Then re-run the estimation-error chain in §2.2 to confirm that the H-norm stochastic term is indeed O(n^{-s/(2s+d)}). If any hypothesis fails, Theorem 2.1's proof must be supplemented.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The minimax upper bound in Theorem 2.1 depends on Lemma 2.3, imported from [7], but the paper gives no statement of the lemma's hypotheses. The lemma is invoked for a kernel whose RKHS is only assumed equivalent to H^{s+d/2}, for a design density bounded above and below, and for target functions in B^s_{∞,∞}, with parameters α,s' chosen at the limit α→d/(2s+d), s'→2s/(2s+d) and λ=1/n. If the original lemma requires, for example, a specific eigenbasis with exact polynomial eigenvalue decay, a source condition of the form f*∈B^s_{2,∞}, or strict inequality conditions on α and s' that forbid the limiting choices, then the displayed stochastic-error bound does not follow. The paper does not verify any of these conditions. This is the load-bearing gap: every other part of the proof of Theorem 2.1 is either standard or fixable. (There is also a harmless typo in §2.2: the Sobolev weight for H^{s+d/2} should be i^{2s/d+1}, not i^{2s/d}; with this correction the displayed '=N' becomes exact, so it does not affect the rate.)","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies kernel ridge regression (KRR) with an RKHS equivalent to a Sobolev space H^{s+d/2}, for target functions in the Hölder–Zygmund class B^s_{∞,∞}. Theorem 2.1 claims that with λ=1/n this misspecified KRR attains the minimax L2 rate n^{-2s/(2s+d)}. The proof uses wavelet characterizations, an approximation-error calculation, and concentration lemmas quoted from a previous paper. The second part constructs an explicit Mercer kernel k(x,y)=∑ i^{-(2s+d)/d} e_i(x)e_i(y) and shows, for f*=0 and Gaussian noise, that the expected squared Hölder–Zygmund norm of the KRR noise component is comparable to σ^2 log n, implying failure of boundedness/properness in the Hölder–Zygmund norm. Theorems 3.1 and 3.2 state this result, and Lemma 3.9 gives a deterministic fixed-design upper bound, while Lemma 3.8 gives a lower bound on a high-probability design event.","tokens_in":7578,"tokens_out":11576,"duration_ms":105544,"significance":"If the results are fully established, they would provide a sharp positive result for spectral algorithms beyond the Sobolev setting and an interesting negative result on properness. The counterexample is explicit, the wavelet framework is natural, and the scale n^{d/(2s+d)} is the right one. The paper also gives a clear separation: the L2 rate is optimal while the Hölder–Zygmund norm of the estimator diverges logarithmically. These are valuable contributions to the misspecified nonparametric regression literature. The main proofs are transparent in structure and use standard tools (wavelet norm equivalences, operator concentration, Sudakov minoration), but two load-bearing gaps currently prevent the claims from being fully accepted.","major_comments":[{"comment":"The estimation-error bound in Theorem 2.1 depends entirely on Lemma 2.3 (and Lemma 2.2), but the manuscript quotes these lemmas without stating their hypotheses. The lemma is invoked for a kernel whose RKHS is only assumed equivalent to H^{s+d/2}, for a design density bounded above and below, and with parameters α and s' chosen at the boundary α→d/(2s+d), s'→2s/(2s+d), λ=1/n. If the original lemma requires, for example, an eigenbasis with exact polynomial eigenvalue decay, a source condition f*∈B^s_{2,∞}, or strict inequalities that forbid the limiting choices, the displayed stochastic error bound does not follow. Please state the hypotheses of Lemmas 2.2–2.3 and verify them for the present kernel and design. In addition, the phrase 'choose α→ ... and s'→ ...' should be replaced by a fixed choice of sufficiently small ε (or an explicit sequence) with a uniform constant, since a limit ins","section":"Section 2.2, Lemmas 2.2–2.3 (quoted from [7])"},{"comment":"The packing argument in the Sudakov lower bound is not fully justified. The claim that 'each such metric ball contains at most (2B/a)^2 points' is asserted after noting that d_X(j,k)<√a σ implies Σ_{jk}≥aσ²/2 and that ∥Σ e_j∥₂≤Bσ². These facts alone do not yield the stated metric-ball bound as written. A complete argument can be made by considering all points in one ball, lower-bounding the quadratic form 1^T Σ_G 1 by O(a m² σ²) and upper-bounding it by m B σ², giving m=O(B/a). But that step is missing. Since the lower bound in Theorem 3.2 depends on this packing bound, please provide the full derivation.","section":"Section 3.1, Lemma 3.8"}],"minor_comments":[{"comment":"The Sobolev weight for the RKHS H ≍ H^{s+d/2} should be i^{(2s+d)/d}, not i^{2s/d}. With the stated i^{2s/d}, the displayed chain leading to ∥g∥²_H ≲ N does not follow from the preceding coefficient bound. This is likely a typo, but it should be corrected for the algebra to be coherent.","section":"Section 2.2, displayed approximation-error calculation"},{"comment":"The constant is stated as C=C(σ,s,d,F), but F is never defined. The theorem statement should also explicitly include the assumption that the RKHS of K is equivalent to H^{s+d/2}([0,1]^d) and the density bounds, rather than leaving them in the surrounding prose.","section":"Theorem 2.1 statement"},{"comment":"The eigenvalues μ_j of the kernel are used before being defined. Please state explicitly that μ_j are the Mercer eigenvalues of k with respect to the wavelet basis, i.e. μ_j = j^{-(2s+d)/d} up to constants, and that the kernel is positive definite.","section":"Section 3.1, Lemma 3.3"},{"comment":"The paper claims minimax optimality but does not state the corresponding lower bound or cite a specific source. Please state the lower bound used (e.g., sup_{∥f*∥≤1} E∥f̂−f*∥² ≥ c n^{-2s/(2s+d)}) and give a precise reference.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the journal's scope and addresses a timely question. The main obstruction is the unverified external lemma in the upper-bound proof; that is load-bearing and cannot be waved away by citing [7] without stating hypotheses. The packing gap in Lemma 3.8 is likely fixable with a short argument, but it must be supplied. I would be willing to review a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper has one genuinely new and solid result: for the explicit kernel k(x,y)=∑ i^{-(2s+d)/d} e_i(x)e_i(y), KRR with f*=0 and Gaussian noise satisfies E||\\hat{f}_\\lambda||^2_{B^s_{\\infty,\\infty}} ≍ σ^2 log n. Properness fails in Hölder–Zygmund norm, with only a logarithmic divergence. The proof is self-contained and mostly checks out. The Sudakov minoration in Lemma 3.8 is terse but valid: the column-norm bound ||Σ e_j||_2 ≤ Bσ^2 limits the number of indices in a small metric ball to a constant, so the packing number is linear in |J_{ℓ_n}|. That is a nice counterexample, and I do not see where it breaks.\n\nThe minimax optimality theorem, by contrast, is not self-contained. The stochastic-error bound is imported from Lemma 2.3 in [7], stated without hypotheses. The kernel is only assumed to have RKHS equivalent to H^{s+d/2}, the design density is bounded, and the parameters take the limiting values α→d/(2s+d), s'→2s/(2s+d), λ=1/n. If [7]'s lemma requires strict inequalities, a particular eigenbasis, or a source condition, the bound does not follow as written. You cannot check any of that from this paper. The approximation-error part is fine once you fix a typo: the Sobolev weight for H^{s+d/2} should be i^{(2s+d)/d}, not i^{2s/d}. With that correction the displayed '=N' is exact, and the rate is unchanged.\n\nSo the properness failure is publishable on its own; the minimax theorem is plausible but currently rests on an unverified import. The paper deserves a serious referee, with the missing lemma hypotheses as the primary issue to chase. I'd cite it if I worked on misspecified spectral algorithms or Hölder/Besov regression; otherwise it's a useful contribution to that literature. Send it to review.","headline":"Properness failure (log n divergence) is new and solid; the minimax optimality claim depends on unstated hypotheses in Lemma 2.3 from [7].","tokens_in":8001,"tokens_out":11422,"would_cite":true,"duration_ms":97005,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G08","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Oversmoothing kernel ridge regression is minimax optimal over Hölder–Zygmund classes, but its noise component's Hölder norm diverges as σ² log n.","keywords":["kernel ridge regression","Hölder–Zygmund class","minimax optimality","properness","spectral algorithms","wavelet methods","nonparametric regression","Sobolev spaces"],"falsifier":"A concrete check: simulate the kernel k(x,y)=∑ i^{-(2s+d)/d} e_i(x)e_i(y) on [0,1]^d for, say, s=1, d=1, with f*=0, Gaussian noise, and n ranging from 10^3 to 10^6. Compute the empirical wavelet coefficients of the KRR solution with λ=1/n and estimate E||S_n||²_{B^s_{∞,∞}} by Monte Carlo. If the growth is not approximately σ² log n but instead saturates, the properness-failure theorem would be false. Conversely, a separate simulation with a smooth kernel (e.g., Matérn with smoothness s+d/2) on a Hölder-smooth function should confirm the L2 rate n^{-2s/(2s+d)} without log factors.","tokens_in":7121,"feed_emoji":"","tokens_out":4797,"duration_ms":38888,"temperature":0.7,"pith_summary":"The paper proves that kernel ridge regression with a kernel whose reproducing kernel Hilbert space is equivalent to a Sobolev space of smoothness s+d/2—one degree 'too smooth' for the target Hölder–Zygmund class—attains the minimax L2 error rate n^{-2s/(2s+d)} for nonparametric regression over the Hölder–Zygmund class, with no extra log factors. It then shows a striking failure: for the zero regression function and Gaussian noise, the expected squared Hölder–Zygmund norm of the KRR noise component grows like σ² log n. This means the estimator is not proper in the Hölder norm, even though it satisfies the same source condition that guarantees properness in Sobolev spaces. Together these results clarify the boundary between spectral algorithms that are minimax optimal for smoothness classes and those that inherit additional norm control.","feed_headline":"KRR matches Hölder minimax rate; noise norm grows as log n","feed_subtitle":"Sobolev-smooth kernels give optimal L2 rates, but the estimator's Hölder norm diverges as σ² log n.","key_machinery":"The key machinery is the wavelet characterization of function spaces: a boundary-adapted orthonormal wavelet basis diagonalizes both the Hölder–Zygmund norm (sup over levels of 2^{ℓ(s+d/2)} max coefficients) and the Sobolev norm (ℓ² weighted). The proof uses the dual representation of the KRR solution, S_n(x)=k_X(x)ᵀ A^{-1} ε, which reduces the noise term to wavelet coefficients G_j=√μ_j ψ_jᵀ A^{-1} ε. Concentration of the empirical Gram matrix at a critical cutoff level ℓ_n (where 2^{-ℓ_n(2s+d)}≍n^{-1}) couples with Sudakov minoration for a lower bound and Gaussian maximum bounds for an upper bound, yielding the log n behavior.","core_discovery":"The central result is that misspecified kernel ridge regression is minimax optimal for Hölder–Zygmund classes: when the RKHS is equivalent to H^{s+d/2} and λ=1/n, the L2 risk is bounded by C n^{-2s/(2s+d)}, matching the minimax lower bound. The proof decomposes the error into approximation and stochastic parts, using wavelet coefficient bounds and a concentration inequality for the empirical covariance. The second result shows that this optimality is fragile: for the kernel k(x,y)=∑ i^{-(2s+d)/d} e_i(x)e_i(y), with f*=0 and Gaussian noise, the expected squared Hölder–Zygmund norm of the estimator's noise component is of order σ² log n, so the estimator's Hölder–Zygmund norm diverges slowly e","pith_inferences":["The log n divergence suggests that any attempt to use the KRR estimator for pointwise or sup-norm inference, where Hölder–Zygmund control matters, would require additional smoothing or truncation even though L2 estimation is already optimal.","A natural testable extension is that replacing kernel ridge with a truncated spectral estimate (hard thresholding at the cutoff level) should restore properness and yield a Hölder–Zygmund norm of order O(1), at the price of an extra log factor in L2 risk; this could be checked numerically.","The mechanism behind the divergence—variance accumulation across wavelet levels near the cutoff—may apply to other spectral algorithms beyond ridge regression, including principal-component regression and gradient methods.","The minimax optimality result may extend to other oversmoothing kernels (e.g., Matérn with smoothness s+d/2) as long as the RKHS is norm-equivalent to the Sobolev space and the covariance concentration inequality holds."],"forward_implications":["For any regression function in the Hölder–Zygmund class B^s_{∞,∞}, choosing λ=1/n and a kernel whose RKHS is equivalent to H^{s+d/2} achieves the minimax L2 rate without any log penalty.","The result extends the known minimax optimality of spectral algorithms from Sobolev classes to Hölder–Zygmund classes for this family of kernels.","The failure of properness means that even though the estimator is L2-optimal, its Hölder–Zygmund norm of the noise component grows as σ√(log n) in expectation, so it does not inherit the target smoothness.","The log n divergence is tight: the conditional upper bound holds for every fixed design, and the lower bound holds on a design event with probability tending to one.","The construction pins down the cutoff frequency ℓ_n as the resolution level where the empirical Gram matrix concentrates while cumulative wavelet variance accumulates logarithmically."],"fun_headline_variants":["KRR hits minimax rate, but Hölder norm blows up as log n","Optimal L2 rate, yet Hölder norm diverges logarithmically","Misspecified KRR: minimax L2, but noise norm grows log n","For Hölder class, KRR optimal in L2 but fails properness"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The minimax upper bound relies on a concentration inequality for the empirical covariance operator, quoted from prior work, that is assumed to hold for the Sobolev-type kernel and the design with bounded density, under parameter limits that approach the boundary of its stated range; if that inequality fails at those boundary values, the proof of the n^{-2s/(2s+d)} rate does not go through.","fun_headline_variants_meta":{"raw":{"variants":["KRR hits minimax rate, but Hölder norm blows up as log n","Optimal L2 rate, yet Hölder norm diverges logarithmically","Misspecified KRR: minimax L2, but noise norm grows log n","For Hölder class, KRR optimal in L2 but fails properness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00017,"raw_usage":{"total_tokens":1063,"prompt_tokens":663,"completion_tokens":400,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":407,"completion_tokens_details":{"reasoning_tokens":316}},"tokens_in":407,"tokens_out":400,"duration_ms":4329,"temperature":1.0,"reasoning_tokens":316,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T10:48:18.267342+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check: simulate the kernel k(x,y)=∑ i^{-(2s+d)/d} e_i(x)e_i(y) on [0,1]^d for, say, s=1, d=1, with f*=0, Gaussian noise, and n ranging from 10^3 to 10^6. Compute the empirical wavelet coefficients of the KRR solution with λ=1/n and estimate E||S_n||²_{B^s_{∞,∞}} by Monte Carlo. If the growth is not approximately σ² log n but instead saturates, the properness-failure theorem would be false. Conversely, a separate simulation with a smooth kernel (e.g., Matérn with smoothness s+d/2) on a Hölder-smooth function should confirm the L2 rate n^{-2s/(2s+d)} without log factors.","supporting_citations":[],"review_version":1}