{"id":"17c56895-b5ed-4f84-883c-0d75bdd108a9","arxiv_id":"2603.28956","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Sharp bias and noise-error bounds for minimum-norm interpolators in 2-uniformly convex Banach spaces, with the first ℓ_p-MNI rates for non-Gaussian sub-Gaussian covariates.","lead":"This paper proves generalization bounds for minimum-norm interpolation in Banach spaces whose norm is only mildly curved (2-uniformly convex), without assuming the covariates are Gaussian. It gives sharp bounds for ℓ_p-norm interpolation with sub-Gaussian covariates, a case that previously could only be handled for Gaussian inputs.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption 2 (Eq. 5) is the load-bearing inductive-bias premise for Theorem 4, but the sub-Gaussian proof never establishes the required lower bound on ∥bf_n(X,ξ)∥; the provided ball inclusion is one-sided.","rationale":"The reader's weakest assumption is Assumption 2, and I agree that it is the most load-bearing point. The paper's main non-Gaussian claim (Theorem 4) is conditional on an inductive-bias property that is used throughout the bias-variance proof but is only asserted to hold in the stated parameter regime. Unlike the omitted net argument (footnote 7) or the eΘ-only-upper-bound issue in Corollary 1, this is a premise of the central theorem itself; if it fails, the upper bound on T2 does not follow even if every combinatorial estimate is correct. The sub-Gaussian verification is especially concerning because the Kashin inclusion is one-sided: containing a Euclidean ball gives an upper bound on the interpolator's norm, while Assumption 2 requires a lower bound. The text in §7.1.5 even states a lower bound α_d ≲ E∥ξ∥_n ≲ 1, which is internally inconsistent with α_d defined earlier unless the notation is swapped; either way, the needed lower bound is not rigorously established. My suggested re-derivation of Step I directly tests whether the available tools plus the theorem's d-range actually force E∥ξ∥_n ≥ C. Since this concern does not change the reader's conditional verdict, I recommend UNCHANGED: the paper should be accepted only after the inductive-bias verification is supplied or made an explicit hypothesis in Theorem 4.","tokens_in":43950,"tokens_out":17531,"duration_ms":174648,"concrete_test":"Re-derive Step I of §7.1.5 for i.i.d. symmetric sub-Gaussian covariates using only the two ingredients actually available there—Kashin's inclusion cB^n ⊂ P eB^d_p and Lemma 7's bound M*(P eB^d_p)≲α_d—and compute the resulting order and explicit constants of E∥ξ∥_n. Then check whether the theorem's condition d≥n exp(C p(d)^{-2}) forces this quantity to be ≥ C with probability at least 1−n^{-2}. If the derived lower bound is O(α_d^{-1}) with α_d^{-1}→0, Assumption 2 is unverified and the proof of Theorem 4 lacks its inductive-bias premise.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 4's eΘ(n/d) noise bound depends on Assumption 2 (Eq. (5)): with probability at least 1−n^{-2}, the pure-noise interpolator satisfies ∥bf_n(X,ξ)∥ ≥ C for an absolute C. This premise is used in the proof of Theorem 6 and in Step II's variational argument (e.g., Eq. (36)): it ensures that the true signal has smaller norm than the noise fit, and it controls the sign/constants in the bias-variance decomposition. If noise can be interpolated with small norm, the upper bound on T2 collapses. The manuscript says that the stated range of n and d is 'essential for Assumption 2 to hold,' but it never supplies the verification. In the sub-Gaussian part (§7.1.5), the only lower-bound ingredient is Kashin's inclusion cB^n ⊂ P eB^d_p together with Lemma 7's M*(P eB^d_p)≲α_d. That inclusion is an upper bound on ∥ξ∥_n, not a lower bound; the lower bound obtained via M·M* ≥ 1 is of order eα_d = α_d^{-1}, which is not shown to be ≥ C in the theorem's regime. Thus the central claim currently rests on an unverified inductive-bias assumption, exactly where the non-Gaussian extension is most fragile.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies minimum-norm interpolation (MNI) in 2-uniformly convex Banach spaces. It proposes a general upper bound on the structural error T1 under UC(2) (Theorem 1), a sharp localized bound in linear models when the unit ball is in isotropic or John's position and covariates are sub-Gaussian (Theorem 2), and a lower bound on the noise term via K-convexity and cotype-2 (Theorem 3). For the ℓ_p-MNI specifically, with p close to 2 and d growing faster than n, it claims sharp bounds T2 = eΘ(n/d) under i.i.d. symmetric sub-Gaussian covariates (Theorem 4 and Corollary 1), extending earlier Gaussian-covariate results of Donhauser et al. The proof combines Dvoretzky/Kashin-type inclusions, entropy estimates, convex concentration, and a variational argument over dyadic blocks. The manuscript is an extended version of the authors' ICML 2024 preliminary work.","tokens_in":44330,"tokens_out":6133,"duration_ms":63509,"significance":"If the main claims are fully established, this would be the first sharp benign-overfitting result for MNI under a non-Hilbert norm with non-Gaussian covariates, matching the known Gaussian rate n/d. The geometric approach — using UC(2), M·M* estimates, K-convexity, and positions of convex bodies — is promising and is a genuine conceptual contribution. The paper also usefully identifies the role of inductive-bias assumptions and positions. However, the advertised non-Gaussian extension currently rests on an unverified inductive-bias assumption (Assumption 2), and there is a mismatch between Corollary 1 and its proof. The significance is therefore conditional on closing these gaps.","major_comments":[{"comment":"Assumption 2 is load-bearing for Theorem 4, but in the sub-Gaussian part of the proof the required lower bound on ∥bfn(X,ξ)∥ = ∥ξ∥_n is never established. Step V (Section 7.1.5) obtains the inclusion cB^n ⊂ P eB^d_p (Eq. (43)) and an upper bound on M*(P eB^d_p) via Lemma 7; the lower bound on ∥ξ∥_n is only an indirect consequence of M·M* ≥ 1, which yields a bound of order α_d^{-1}, and no argument shows this is ≥ C in the theorem's regime. Since Theorem 4's conclusion depends on Assumption 2, the theorem as stated is not proved. Please either supply a direct verification of Assumption 2 for the sub-Gaussian case, or state Theorem 4 with Assumption 2 as an explicit condition and adjust the conclusion accordingly.","section":"§1.1, Eq. (5); §7.1.5"},{"comment":"Corollary 1 states T1 = eΘ(d^{2p-2}/n^p), but its proof in §7.3.3 only establishes the upper bound T1 ≲ Õ(Gn(B_p^d)^{2p}) = Õ(d^{2p-2}/n^p). No matching lower bound is derived, so the eΘ claim is unsupported. The statement should be weakened to an O/Õ upper bound, or the missing lower bound should be proved.","section":"§3.1, Corollary 1; §7.3.3"},{"comment":"The uniform estimate (26) over the sets V_k, which are not finite, is asserted while the footnote says 'the statement requires an entirely elementary net argument which is omitted.' This estimate is used to control all dyadic blocks in Step II and thus supports the main variational argument of Theorem 4. The omitted net argument is load-bearing and should be supplied before the result can be considered proved.","section":"§7.1.1, after Eq. (26), footnote 7"},{"comment":"Several results depend on works that are either not yet published or coauthored by the present authors: Lemma 1 and Remark 3 rely on Bizeul (2025) and Bizeul–Klartag (2025) for the slicing/isotropic-position estimates, Remark 2 defers to Kur et al. (2026), and §4.3 invokes a quantitative Anderson inequality from Bizeul et al. (2026, 'to appear'). These dependencies are not part of the manuscript and are not independently verifiable from the text. For a journal submission, either include the necessary statements and proofs, or clearly mark them as external assumptions. This particularly affects the completeness of Theorem 2.","section":"§3, Lemma 1 and Remark 3; §4.3"}],"minor_comments":[{"comment":"The notation α_d := sqrt(p/(p-1)) eα_d := α_d^{-1} is confusing; α_d and eα_d should be defined separately and unambiguously.","section":"§6, Notation"},{"comment":"There are spelling inconsistencies: 'Blashke–Santalo' and 'Blackhe-Santalo' both appear. Please standardize to 'Blaschke–Santalo'.","section":"§7.3, Lemma 1 proof"},{"comment":"The parameter p(d) is used both as C/log log d and, in Theorem 4, via the expression 1 + C log log log(d)·p(d). This is notationally confusing and should be clarified.","section":"§3.1 and §6"},{"comment":"In the proof of Theorem 3, the sentence 'note that for each realization X and ξ, it holds fLn|X = ξ' appears to contain a typo; the linearized operator should be defined and evaluated precisely.","section":"§7.4.1"},{"comment":"Several references are marked 'to appear' or are not clearly peer-reviewed (Bizeul et al. 2026, Kur et al. 2026, Karhadkar et al. 2026). The reader would benefit from availability information or versions at the time of submission.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper has a plausible high-level argument and identifies an interesting connection between local Banach-space geometry and benign overfitting. However, the main advertised sub-Gaussian ℓ_p result depends on Assumption 2, which is not verified in the text, and Corollary 1's eΘ statement exceeds what is proved. These are load-bearing issues rather than presentation issues. I recommend major revision, with the expectation that the authors either prove Assumption 2 for the stated regime or restructure the main theorems so that all inductive-bias conditions are explicit. The extensive reliance on unpublished/coauthored works should also be addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper introduces a genuinely new convex-geometric framework for MNI, and I think that part is worth taking seriously. The UC(2)-based bias bound, the K-convexity variance lower bound, and the RMM* vocabulary are real additions to a literature that has been stuck on Gaussian tools. The Gaussian-covariate proof of Theorem 4 also contains substantial work: the dyadic decomposition, the variational argument, and the entrywise control are all plausible, and I found no clear contradiction in the central derivation.\n\nThe soft spots are real, and they cluster around the sub-Gaussian extension. Assumption 2 (Eq. 5) is load-bearing for Theorem 4, and I could not find a proof of it anywhere. The stress-test note is right: in the sub-Gaussian Step I, Kashin's inclusion cB^n ⊂ P eB_p^d gives an upper bound on ∥ξ∥_n, not a lower bound. The lower bound obtained from M·M* ≥ 1 has the opposite order and is not shown to be ≥ C in the stated regime. The paper says the upper and lower bounds on n are 'essential for Assumption 2 to hold,' but that is a condition, not a verification. If the noise interpolator can have small norm, the variational argument in Step II collapses, and with it the T2 = eΘ(n/d) claim.\n\nOther issues are smaller but point the same way: Corollary 1 states T1 = eΘ(...) while Section 7.3.3 only proves an upper bound; Footnote 7 admits the net argument in Theorem 5 is omitted; and several discussion-level claims are deferred to author forthcoming papers. Individually minor, but together they mean the sharpness claims run ahead of the written proof.\n\nWho should read this: anyone working on non-Hilbert interpolation or benign overfitting. The geometric framework deserves attention even if Theorem 4 needs another pass. My recommendation is to send it to serious referees rather than desk reject. The right outcome is major revision: prove Assumption 2 in the sub-Gaussian regime or restate the theorem as conditional on it, supply the omitted net argument, and reword Corollary 1 as an upper bound until the matching lower bound exists.","headline":"Serious geometric framework, but the headline sub-Gaussian ℓ_p bounds rest on Assumption 2, which is never actually proved — the non-Gaussian extension is conditional, not established.","tokens_in":44783,"tokens_out":6650,"would_cite":true,"duration_ms":69747,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["46B20","46B09"],"pacs":[],"model":"deepseek-v4-flash","headline":"Minimum-norm interpolation achieves benign overfitting for ℓp norms under sub-Gaussian covariates, provided the norm is 2-uniformly convex.","keywords":["minimum norm interpolation","benign overfitting","2-uniform convexity","ℓp regression","sub-Gaussian covariates","K-convexity","Banach space geometry","overparameterized linear regression"],"falsifier":"Simulate the ℓp-MNI with p = 1 + C/log log d, i.i.d. Rademacher covariates, d ≍ n exp(C p(d)^{-2}), and zero ground truth; compute E||ŵ||_2^2 over many noise draws. If the noise error deviates from Θ(n/d) by more than a polylogarithmic factor, or if ∥ŵ(X,ξ)∥ falls below the constant C with probability larger than n^{-2}, then Theorem 4 or Assumption 2 fails.","tokens_in":43867,"feed_emoji":"📐","tokens_out":5668,"duration_ms":55923,"temperature":0.7,"pith_summary":"The paper aims to show that the minimum-norm interpolator (MNI) generalizes well in overparameterized regression whenever the norm of the hypothesis space is 2-uniformly convex, a curvature condition far weaker than being induced by an inner product. It proves an upper bound on the structural (bias) error under this condition, and shows the bound is sharp for linear models when the unit ball sits in isotropic or John's position and covariates are i.i.d. symmetric sub-Gaussian. The main payoff is for ℓp-MNI with p just above 1: the noise error is eΘ(n/d), matching the Gaussian-covariate rate and establishing benign overfitting for a non-Hilbert norm without Gaussian assumptions. A sympathetic reader should care because this is the first sharp analysis of non-inner-product MNIs beyond Gaussian covariates, and it ties generalization to classical Banach-space geometry.","feed_headline":"ℓp interpolation hits the Gaussian rate under sub-Gaussian data","feed_subtitle":"First sharp proof that minimum-norm interpolation generalizes for non-Hilbert norms, via 2-uniform convexity.","key_machinery":"The load-bearing object is the 2-uniform convexity of the norm, a quantitative parallelogram-type inequality with constant t = p−1 for ℓp when p ∈ (1,2], which replaces the inner product as the source of curvature. Around it, the argument is carried by the M M* ratio R_{MM*} = M_n(F) M*_n(F)/n, the product of the mean norms of the projected unit ball and its polar, together with the K-convexity constant, which together control the bias-variance decomposition T1+T2. A variational interpolation argument plus a finite-volume-ratio covering lemma for sub-Gaussian projections supplies the sharp ℓp rates.","core_discovery":"On the paper's own terms, the central discovery is that 2-uniform convexity of the norm controls the bias of the MNI: the structural error T1 is O(R_{MM*} G_n(F)/t), where t is the uniform-convexity constant and R_{MM*} is the mean-width ratio of the projected unit ball. In linear models with the unit ball in isotropic or John's position and i.i.d. symmetric sub-Gaussian covariates, this bound becomes sharp, T1 = O~((G_n(K)^2 + 1/n)/t^2). For the ℓp-MNI with p near 1 and O(1)-sparse ground truth, the noise error satisfies T2 = eΘ(n/d), the same rate as the Gaussian case, and the paper proves this without any inner-product structure or closed-form solution.","pith_inferences":["The same geometric mechanism suggests that the noise-error rate should hold for any 2-uniformly convex norm placed in the right position, provided sharp covering estimates for that position exist, a step the authors explicitly leave open.","Because the proof of Theorem 4 only needs curvature on average, one can conjecture that cotype 2, together with a tailored position, is the true minimal assumption; the paper itself floats this conjecture.","The techniques could transfer to nonparametric settings such as Sobolev-space MNIs, where a suitable norm combination preserves 2-uniform convexity and may yield minimax-optimal interpolators.","A testable consequence for practitioners: ℓp-MNI with p slightly below 2 should be roughly as harmless as ridge interpolation on sub-Gaussian features, without requiring Gaussian design matrices."],"forward_implications":["If Theorem 4 is correct, the ℓp-MNI with p just above 1 achieves noise error eΘ(n/d) under i.i.d. Rademacher or other symmetric sub-Gaussian covariates, matching the rate previously known only for Gaussian data.","Theorem 2 implies the structural error of the MNI decays at the fast 1/n rate, up to logarithmic factors, for any 2-uniformly convex norm in isotropic or John's position, not just Hilbert norms.","Theorem 3 gives a general lower bound on the noise-driven variance of the MNI in terms of the K-convexity constant, so the n/d rate is intrinsic to the interpolation problem, not an artifact of ℓp geometry.","Corollary 1 shows the bias term T1 = eΘ(d^{2p−2}/n^p), identifying exactly where the exponent p enters the mean squared error."],"fun_headline_variants":["First sharp bounds for MNI with non-Hilbert norms","ℓp interpolation matches Gaussian rate for sub-Gaussian data","Sharp MNI generalization bounds for 2-uniformly convex norms","Non-Hilbert MNI: first sharp results under sub-Gaussian covariates"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"Everything rests on Assumption 2: with high probability the norm of the interpolator of pure noise is bounded below by a universal constant; this inductive-bias condition is verified only for ℓp in a restricted dimension range, and if pure noise can be fit with small norm, the bias-variance bounds collapse.","fun_headline_variants_meta":{"raw":{"variants":["First sharp bounds for MNI with non-Hilbert norms","ℓp interpolation matches Gaussian rate for sub-Gaussian data","Sharp MNI generalization bounds for 2-uniformly convex norms","Non-Hilbert MNI: first sharp results under sub-Gaussian covariates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000245,"raw_usage":{"total_tokens":1415,"prompt_tokens":829,"completion_tokens":586,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":573,"completion_tokens_details":{"reasoning_tokens":513}},"tokens_in":573,"tokens_out":586,"duration_ms":5599,"temperature":1.0,"reasoning_tokens":513,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T17:04:55.405790+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the ℓp-MNI with p = 1 + C/log log d, i.i.d. Rademacher covariates, d ≍ n exp(C p(d)^{-2}), and zero ground truth; compute E||ŵ||_2^2 over many noise draws. If the noise error deviates from Θ(n/d) by more than a polylogarithmic factor, or if ∥ŵ(X,ξ)∥ falls below the constant C with probability larger than n^{-2}, then Theorem 4 or Assumption 2 fails.","supporting_citations":[],"review_version":1}