{"id":"0ea7e35c-284f-4db9-b3b8-e855a47e92ea","arxiv_id":"2607.28846","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":8.0,"correctness_risk":"low","formal_verification":"none","parameter_count":1,"one_line_summary":"Gaussian-refit movement quantiles plus a bias term give a computable, rate-optimal upper confidence bound on the realized prediction error of kernel ridge regression under symmetric noise.","lead":"This paper constructs an upper confidence bound on the prediction error of a single kernel ridge regression fit by repeatedly refitting with Gaussian-perturbed responses and taking an order statistic of the fit movement. The bound is finite-sample valid under symmetric noise without moment assumptions and shrinks at the minimax rate, while cross-validation intervals are shown to stay stuck at a wider n^{-1/2} margin.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 4.1 does not by itself deliver a finite-sample 1−α upper confidence bound: the coverage statement has an uncomputable slack C0 φ(z1−α)^{-1}(δ+δ_N)+C1/√L, so the advertised level guarantee is only asymptotic or conditional on an unverified delocalization condition.","rationale":"The reader's verdict correctly identifies the well-specified RKHS-ball assumption as a major limitation, but my concern is more direct: even within the well-specified regime, Theorem 4.1 does not establish a finite-sample 1−α upper confidence bound because its right-hand side contains an uncomputable delocalization slack. This is not a claim that the theorem is false; rather, the theorem proves a weaker statement than the abstract and introduction advertise. The reader's rationale already notes that 'the useful form of the guarantee requires delocalized noise,' which is close to this concern, but the verdict still ACCEPTs the central claim as stated. I would accept the paper only if the claims are revised to say: the certified statement is a slacked inequality with an explicit but unobservable delocalization term, plus an asymptotic rate result; the exact-level finite-sample guarantee is not proven. The simulations and data-driven envelope are empirical and are not affected by this concern. The mathematical core appears sound; the issue is the interpretation of Theorem 4.1 as a computable finite-sample UCB. Thus CONDITIONAL rather than REJECT.","tokens_in":34181,"tokens_out":14588,"duration_ms":180033,"concrete_test":"Use the §6 synthetic setup with known true f* and noise magnitudes. For n in {500,2000}, L in {199, 2000}, and α=0.05, compute the worst-case bound (19) with B and κ known. Over R replicates, estimate the conditional coverage P(E(f̂) ≤ bU_α | |w|) by resampling signs and refits, and separately compute the slack C0 φ(z)^{-1}(δ+δ_N)+C1/√L from the true |w|. Check whether the empirical miscoverage exceeds α+slack; if the theorem is correct it should not. Then test the stronger practical question: try to bound δ+δ_N from observed data alone, e.g. by maximizing over |w| ≤ a in (18). If no data-only upper bound on the slack can be certified below a small threshold, then the finite-sample level claim, as opposed to the slacked inequality, is not established by Theorem 4.1.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim, as stated in the abstract and Section 1, is that Algorithm 1 with the worst-case envelope (18) yields a computable finite-sample upper confidence bound at level 1−α under only symmetric noise. But Theorem 4.1 proves, for every fixed n and L, the conditional inequality P(E(f̂) > bU_α | |w|) ≤ α + C0 φ(z1−α)^{-1}(δ+δ_N) + C1/√L, where δ, δ_N and ρ are defined in (23) in terms of M = diag(|w|)H^T H diag(|w|). These delocalization functionals depend on the unobservable noise magnitudes and are not computable from the data. The theorem does not assume δ+δ_N is small and provides no data-dependent certificate that it is small. Hence, for a fixed finite sample and finite L, the inequality (24) cannot be converted into a certified statement P(E(f̂) ≤ bU_α) ≥ 1−α at the user-specified α; at best it gives P(E(f̂) ≤ bU_α) ≥ 1−α−slack with an unknown slack. Section 4.1 itself only claims the right-hand side tends to α as δ,ρ→0 and L→∞, which is an asymptotic regime, not a finite-sample level guarantee. The practical data-driven envelope (20) is explicitly empirical, so the finite-sample guarantee rests entirely on Theorem 4.1. The paper is transparent about the need for delocalized noise and about the worst-case envelope being loose, but the slogan that the procedure produces a computable finite-sample UCB at level 1−α under symmetric noise goes beyond what Theorem 4.1 establishes. This is a load-bearing gap because it affects the headline contribution, not merely a proof-technicality; the proven statement is an oracle-type bound with an unmonitored slack whose size controls whether the nominal level is actually reached.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a Gaussian-refit upper confidence bound for the empirical prediction error of kernel ridge regression under fixed design and symmetric noise. The bound is bU_α = (q_{1−α}(a) + b)^2, where q_{1−α}(a) is an order statistic of L refit movements with Gaussian multipliers at an envelope a, and b bounds the smoothing bias. For the theoretical worst-case envelope (18), the paper claims a computable finite-sample upper confidence bound at level 1−α (Theorem 4.1), contraction at the minimax rate (Theorem 4.8), and fundamental limitations of held-out-loss intervals (Propositions 4.10 and 4.12). For a practical data-driven envelope (20), the paper reports full coverage within small multiples of the true error quantile across synthetic and real data, including heavy-tailed noise.","tokens_in":34524,"tokens_out":12146,"duration_ms":129371,"significance":"If the stated guarantees held, the paper would be an important contribution: a computable tail bound for the realized error of a kernel-ridge fit under essentially no moment assumptions, with a matching minimax rate and strong empirical tightness. The paper is transparent about its assumptions and limitations, and it ships explicit proofs, reproducible code, and a thorough experimental comparison. The negative results on cross-validation margins (Propositions 4.10 and 4.12) are interesting and appear sound. However, the central validity theorem as stated does not deliver the advertised finite-sample 1−α bound, and the proof has a gap in the Monte-Carlo calibration step. These issues are load-bearing for the headline contribution and need to be addressed before the paper can be accepted.","major_comments":[{"comment":"The proof treats the sample order statistic q_{1−α}(a) as an upper bound on the population quantile Q0 up to O(L^{−1/2}), then uses the inclusion {E(f̂)>bU_α} ⊆ {T>Q0}. This is not justified. The DKW inequality gives |F_L−F|≤ε with high probability, which implies q_{1−α}(a) ≥ Q0−ε, not q_{1−α}(a) ≥ Q0. In fact P(q_{1−α}(a) < Q0) is of constant order, so the set inclusion fails on a non-vanishing probability set. The O(L^{−1/2}) slack in (24) therefore does not follow from the argument given. A correct treatment requires either a DKW-corrected quantile (e.g., a more conservative order statistic with an explicit margin) or a direct averaged comparison of P(T>q) with P(T>Q0). As written, the proof of Theorem 4.1 is incomplete.","section":"Theorem 4.1 / Appendix A.7, Step 1; Eq. (16)"},{"comment":"The paper claims a computable finite-sample upper confidence bound at level 1−α under symmetric noise. Theorem 4.1, however, proves P(E(f̂)>bU_α | |w|) ≤ α + C0 φ(z_{1−α})^{−1}(δ+δ_N) + C1/√L, where δ, δ_N are defined in (23) from M = diag(|w|)HᵀH diag(|w|) and are not computable from data. For fixed n and L, the slack is unknown, so the user-specified level 1−α is not certified. The discussion after (24) explicitly says the right-hand side tends to α only when δ,ρ→0 and L→∞, which is an asymptotic statement. To claim a finite-sample level, the authors must either (i) provide a conservative calibration using universal bounds on C0, C1 and a more conservative nominal level, or (ii) explicitly relabel the result as an asymptotic or approximate bound. The current abstract and (6) overstate what is proved.","section":"Abstract, §1 Eq. (6), and Theorem 4.1 (24)"},{"comment":"The paper's own statement that the bound is useful only in a delocalized regime is honest, but the functionals δ and ρ cannot be verified from data. This reinforces the previous comment: the finite-sample validity claim is conditional on unverifiable structure. The paper should either give a data-dependent certificate for delocalization (which seems difficult) or clearly state that the theoretical guarantee is asymptotic in the delocalization limit and that the finite-sample guarantee is empirical only.","section":"§4.1, 'Whenever δ→0 and ρ→0...'"}],"minor_comments":[{"comment":"The abstract states the bound 'maintains full coverage within twice the true 95% error quantile'. This is supported for the synthetic design (Table 2, 1.7–1.9×) but not for the real elevation study (Table 4, 2.4–3.4×). Please qualify the claim to 'within a small constant' or specify that the factor-of-two holds on the synthetic design.","section":"Abstract and §6 vs. Table 4"},{"comment":"The data-driven envelope uses a pilot penalty λ/20, the only hand-tuned parameter. Its sensitivity is not discussed. A brief robustness check over this factor would strengthen the empirical claims.","section":"§3.4, Eq. (20); §6.2"}],"recommendation":"major_revision","confidential_remarks":"The main obstacle is the gap in the proof of Theorem 4.1 and the mismatch between the advertised finite-sample guarantee and the proved inequality with uncomputable slack. The empirical work is convincing and the lower-bound results are valuable, but the theoretical headline needs repair or re-scoping before acceptance. I would not reject: the core idea is sound and the issues appear addressable by a more conservative calibration and a more careful statement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The Gaussian-refit construction is the real thing. Replacing Rademacher signs with Gaussian multipliers to get envelope monotonicity via Anderson's inequality is a clever and, as far as I can tell, new mechanism, and the sign-multiplier impossibility result plus the holdout-loss floor are solid contributions on their own. The proof chain is mostly explicit and the authors are honest about the worst-case envelope being loose and about the misspecification regime where coverage collapses to 0.10. The empirical work is careful and the comparisons are fair. This deserves a serious referee.\n\nThe soft spot is exactly where the stress-test note points. Theorem 4.1 proves a conditional inequality with an additive slack of C0/φ(z1−α)(δ+δ_N)+C1/√L, and δ, δ_N depend on the unobservable noise magnitudes. So the procedure does not deliver a certified finite-sample 1−α bound at a user-specified α unless you're willing to assume the delocalization functionals are small, which is an unverified condition. The paper's own scope caveat says validity relies on sufficiently delocalized noise, but the abstract's \"computable tail bound\" and the phrase \"calibrated at any confidence level\" overstate what is proven. This is a gap in the headline claim, not a hidden fatal flaw — the bound is still useful and likely conservative in practice — but it should be fixed in the wording or mitigated with a data-dependent check, even a crude one.\n\nSmaller issues: the worst-case envelope needs known B and κ, and is quite loose for slow spectral decay; the data-driven envelope used in the experiments is explicitly empirical, so the finite-sample guarantee applies only to the worst-case version; and the code isn't archived in the preprint, which hurts reproducibility. The bias input b for the theoretical bound is fine, and the g>0 quantile inversion point is minor as the reader notes.\n\nFor whom: people working on wild refitting, resampling-based uncertainty, or honest inference for kernel methods. I'd send it to review and would engage with it myself; the delocalization-slack issue needs to be addressed, but the core idea is worth publishing and building on.","headline":"Genuinely new mechanism and rate-optimal bound, but the 'finite-sample 1−α UCB' slogan outruns Theorem 4.1, which carries an uncomputable delocalization slack.","tokens_in":35136,"tokens_out":1907,"would_cite":true,"duration_ms":24292,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G08","62G09"],"pacs":[],"model":"deepseek-v4-flash","headline":"Kernel-ridge fits can be certified by a Gaussian-refit upper confidence bound that contracts at the minimax rate and needs no noise moments.","keywords":["upper confidence bound","kernel ridge regression","realized prediction error","Gaussian refitting","order statistic calibration","wild bootstrap","cross-validation floor","minimax rate"],"falsifier":"A single simulation with the true function inside the RKHS ball, symmetric noise, and the certified worst-case envelope, where empirical coverage falls below the nominal level while the paper's delocalization functionals are small, would refute the conditional validity claim. Equivalently, a construction of two problems whose held-out loss distributions are nearly indistinguishable but whose fit errors differ by order n^{-1/2}, together with any measurable held-out-loss bound with a smaller margin, would falsify the holdout floor.","tokens_in":33965,"feed_emoji":"📈","tokens_out":5724,"duration_ms":65855,"temperature":0.7,"pith_summary":"The paper tries to establish that a single kernel-ridge-regression fit can come with a computable, finite-sample upper confidence bound on the distance between the fitted function and the unknown truth, measured as average squared error at the design points. The proposed bound is built by adding Gaussian noise to the residuals, refitting, and recording how far the fit moves; an order statistic over repeated refits supplies the tail quantile, and a bias term is added before squaring. The authors claim this bound is valid under only symmetric noise—no finite moments, no known variance—provided the target lies in a known reproducing-kernel-Hilbert-space ball, and that it contracts at the minimax rate n^{-2s/(2s+1)}, matching the prediction error itself. They also argue cross-validation intervals cannot do this: any interval computed from held-out losses is stuck at an n^{-1/2} margin floor, so its margin-to-error ratio diverges whenever the fit converges faster than n^{-1/2}.","feed_headline":"Gaussian refit puts a bound on kernel-ridge error","feed_subtitle":"Finite-sample bound shrinks at the minimax rate and needs no noise moments—cross-validation cannot keep up.","key_machinery":"The central object is the calibrated Gaussian movement query: draw xi ~ N(0, I_n), refit at y + xi o a, and record m(a) = ||M(y + xi o a) - M(y)||_n, which equals ||H(xi o a)||_n for the linear smoother. The construction rests on a Gaussian comparison inequality: enlarging any coordinate of a cannot lower the upper quantiles of the movement, so a computable envelope a_i >= |w_i| dominates the unknown noise magnitudes. An order statistic of L movements estimates the quantile, and a bias input b bounds ||(H - I) f*||_n; the final bound is (quantile + b)^2. This same mechanism delivers both the finite-sample tail guarantee and the minimax-rate contraction.","core_discovery":"On the paper's own terms, the discovery is that replacing sign-flip refitting with Gaussian refitting makes the wild-refit mechanism computable and rate-sharp for kernel ridge regression. Because Gaussian noise is monotone under coordinatewise enlargement of an envelope vector a >= |w|, the quantiles of the refit movement ||H(xi o a)||_n can be used as a conservative proxy for the unobservable noise term ||H w||_n; the bound is bU_alpha = (q_{1-alpha}(a) + b)^2. The paper proves conditional validity with an explicit remainder depending only on delocalization functionals and the number of refits, and it shows the bound attains the minimax rate. By contrast, sign multipliers degenerate to the","pith_inferences":["Editorial inference: If the conditional-symmetry argument transfers to any firmly non-expansive estimator with a computable envelope—as the paper's constrained-fit experiment suggests—the same refit recipe could certify other black-box predictors, not just linear smoothers.","Editorial inference: The n^{-1/2} holdout floor implies that any conformal or split-based uncertainty method targeting the fitted function's realized error, rather than a future observation, will face the same precision barrier; the Gaussian-refit path is one way to route around it.","Editorial inference: A data-dependent spectral envelope, rather than the worst-case one, might close the gap on slowly decaying kernels where the bound stays valid but loose; this is a natural next test.","Editorial inference: Because validity is conditional on noise magnitudes, an adversarial magnitude pattern that concentrates noise on a few coordinates will make the bound vacuous; stress-testing that regime could map the method's practical limits."],"forward_implications":["A user can report a finite-sample 95% upper bound on the realized prediction error of a kernel-ridge fit using only fair noise signs; no variance estimate or moment condition is needed.","The bound's width tracks the actual error: under polynomial spectral decay it contracts at n^{-2s/(2s+1)}, the minimax rate, rather than being pinned at n^{-1/2}.","Cross-validation intervals, including nested and quantile-repaired versions, inherit an n^{-1/2} margin floor; whenever smoothness s > 1/2, their margin-to-error ratio diverges polynomially.","A single set of L refits calibrates every confidence level at once, because all levels read off the same ordered movements.","The construction extends empirically to nonlinear constrained fits and real spatial data with full coverage in the reported experiments."],"fun_headline_variants":["Kernel ridge error bound shrinks at minimax rate","Gaussian refit beats cross-validation for kernel ridge","New bound for kernel ridge prediction error","Refit-based error bound for kernel ridge regression"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The certified bound presupposes the true function lies inside the known reproducing-kernel-Hilbert-space ball with a known kernel bound; if the target is outside that ball—as in the paper's rigid over-smoothed misspecification experiment, where coverage fell to 0.10—the guarantee collapses.","fun_headline_variants_meta":{"raw":{"variants":["Kernel ridge error bound shrinks at minimax rate","Gaussian refit beats cross-validation for kernel ridge","New bound for kernel ridge prediction error","Refit-based error bound for kernel ridge regression"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000172,"raw_usage":{"total_tokens":1127,"prompt_tokens":775,"completion_tokens":352,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":519,"completion_tokens_details":{"reasoning_tokens":301}},"tokens_in":519,"tokens_out":352,"duration_ms":3687,"temperature":1.0,"reasoning_tokens":301,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T01:31:23.548333+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A single simulation with the true function inside the RKHS ball, symmetric noise, and the certified worst-case envelope, where empirical coverage falls below the nominal level while the paper's delocalization functionals are small, would refute the conditional validity claim. Equivalently, a construction of two problems whose held-out loss distributions are nearly indistinguishable but whose fit errors differ by order n^{-1/2}, together with any measurable held-out-loss bound with a smaller margin, would falsify the holdout floor.","supporting_citations":[],"review_version":1}