{"id":"14dbb3f8-50e6-480f-a0b6-831dbce8a1a7","arxiv_id":"2607.22979","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"LazyVI tests feature importance in binary classifiers by fitting a linearized logistic model on neural tangent features and claims op(n^{-1/2}) asymptotic normality.","lead":"This paper introduces LazyVI, a hypothesis test for which input features matter in binary deep-learning classifiers, built by approximating retraining with a linearized 'lazy' model. It claims asymptotic normality of the variable-importance statistic with a faster error rate than prior work, plus code and simulations.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Condition (R3) is not established: the Donsker class D_n depends on W, and W grows with n, so no fixed P0-Donsker class contains the estimators; the op(n^{-1/2}) expansion in Theorem 25 lacks justification.","rationale":"The reader's weakest_assumption field emphasizes A1, but the rationale also lists R3 as a structural dependency. I agree with the overall REJECT verdict, but I focus on R3 as the single most load-bearing concern because it is a proof gap rather than an unproved assumption. Even if A1 were granted, the Donsker verification is invalid for growing W. The paper's central theorem is an asymptotic linear expansion that depends critically on Williamson et al.'s condition (R3); without a fixed Donsker class, the op(n^{-1/2}) remainder is not justified. The R1 norm issue is less severe because one could define ||·||_F as the L2 norm and potentially satisfy the condition, and A1 is an explicit assumption that the authors acknowledge. R3, by contrast, is asserted in the proof using a class that changes with n, so the theorem as stated is not proven. Therefore the reader's rejection is warranted, and no change to the verdict is needed.","tokens_in":65487,"tokens_out":8967,"duration_ms":93628,"concrete_test":"Analytical check: set W_n = n^{0.4} (which satisfies W = o(√n/log n)) and independently compute the L2(P0)-bracketing entropy of D_n defined in Appendix E.3, using Theorem 36's sup-norm covering bound. If ∫_0^1 sqrt(log N_{[]}(ε, D_n, L2(P0))) dε grows with n — e.g., as n^{0.2} — then the argument cannot yield a fixed Donsker class, and the proof of condition (R3) fails for growing W. A complementary simulation: under the null model of §4.1, for n = 200, 800, 3200 with W_n = n^{0.4}, estimate sup_{g ∈ D_n} |G_n(g)| by Monte Carlo; if the distribution widens with n rather than stabilizing, the empirical-process tightness needed for Theorem 25 is absent.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The proof of Theorem 25 verifies Williamson et al.'s condition (R3) by defining D = {Z ↦ ℓ(f(X),Y) − ℓ(f0(X),Y) − E[ℓ(f(X),Y) − ℓ(f0(X),Y)] : f ∈ F} and then showing the entropy integral ∫ sqrt(log N(ε,F,||·||_∞)) dε is finite via Theorem 36. That theorem bounds log N(ε,F,||·||_∞) by W log(1 + C/ε). This is fine for fixed W, but Theorem 25 assumes W = o(√n/log n), so W = W_n grows with n. Condition (R3) requires a single P0-Donsker class D0 such that P0(g_n ∈ D0) → 1. The proof never supplies such a fixed class: D_n changes with n, and its entropy bound carries a factor √W_n, making the entropy integral diverge as n → ∞. Thus the empirical process indexed by D_n need not converge to a tight Gaussian limit, and the asymptotic linear expansion from Theorem 7 does not follow. This gap is independent of the (also unproved) spectral assumption (A1); even granting A1, the Donsker verification is incomplete. The R1 norm ambiguity is a related but separate issue; R3 is the more direct obstruction to the central claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes LazyVI, a variable-importance testing procedure for binary classification with deep ReLU networks under the lazy-training / neural tangent kernel regime. The predictiveness measure is the expected logistic likelihood, and the estimator uses one fitted network plus a ridge-penalized linearized (NTK) fit for each feature subset, avoiding retraining. The main theoretical claim (Theorem 25) is that, under an eigenvalue decay condition on the NTK at the fitted parameters (A1), a regularization-rate condition (A2), and W = o(sqrt(n)/log n), the variable-importance estimator is asymptotically linear with an op(n^{-1/2}) remainder, yielding a valid z-test. The paper also reports simulations, MNIST experiments, an ADNI gene-association application, and a software implementation. The appendices contain detailed proofs of local Rademacher bounds, NTK-RKHS estimates, and the verification of the variable-importance framework conditions.","tokens_in":65863,"tokens_out":7843,"duration_ms":75940,"significance":"If Theorem 25 were established, it would be a practically useful result: it would provide first-order inference for feature importance in binary ReLU classifiers without retraining for each feature subset, improving on the Op(n^{-1/2}) rate in prior lazy-training work. The paper is also transparently written: it ships code, gives explicit conditions, includes large-scale simulation support, and lays out the local-Rademacher machinery in detail. These are genuine strengths. However, the central theoretical claim is not supported as written. The verification of Williamson et al.'s condition (R3) is incomplete because the relevant function class grows with W_n, and the main rate improvement rests on an unproven spectral assumption on the NTK at the fitted parameters. These are load-bearing gaps, not presentation issues.","major_comments":[{"comment":"The proof defines D = {Z ↦ ℓ(f(X),Y) − ℓ(f0(X),Y) − E_P0[ℓ(f(X),Y) − ℓ(f0(X),Y)] : f ∈ F} and bounds log N(ε,F,||·||_sup) ≤ W log(1 + C/ε) via Theorem 36. For fixed W this would give a finite entropy integral and a Donsker class. But Theorem 25 assumes W = W_n = o(√n/log n), so W_n grows with n. Condition (R3) of Williamson et al. requires a single fixed P0-Donsker class D0 such that P0(g_n ∈ D0) → 1. The proof never supplies such a fixed class: D_n changes with n, and its entropy integral is O(√W_n), which diverges as n → ∞. Thus the empirical process indexed by D_n need not converge to a tight Gaussian limit, and the op(n^{-1/2}) expansion in Theorem 25 does not follow. This gap is independent of Assumption (A1); even granting A1, the Donsker verification is incomplete.","section":"Appendix E.3, Condition (R3)"},{"comment":"The faster rate and Theorem 25 depend critically on Assumption (A1): the eigenvalues of the NTK matrix K_{-S} at the fitted parameters decay as μ_j ≤ c j^{-α} with α > 1. The paper states explicitly in §3.2: 'we are currently unable to provide a theoretical proof that Assumption (A1) holds for the class of deep ReLU neural networks' and extrapolates from NTK results at initialization. Since Lemma 23, Theorem 24, and the resulting o(n^{-1/4}) rates used for condition (R1) all rely on this spectral assumption, the paper's central claim is conditional on an unproved structural conjecture. The empirical check in §4.2 is a single simulation setup and does not establish the assumption for general ReLU networks. The abstract's statement that the method relies on 'only a minimal set of assumptions' overstates the situation when a key assumption is conjectural.","section":"§3.2, Assumption (A1)"},{"comment":"The proof verifies condition (R1) using L2(P) convergence rates from Theorem 15 and Theorem 24. However, Theorem 7's condition (R1) is stated in terms of the function-class norm ||·||_F of F, and the paper never specifies which norm ||·||_F represents. If the authors intend ||·||_F to be the L2(P) norm, this should be stated explicitly and used consistently in conditions (D1) and (D2). If a different norm is intended, the L2(P) bounds do not establish the required function-class-norm convergence. This ambiguity is load-bearing because (R1) is one of the three random conditions needed for the asymptotic linear expansion.","section":"§E.3, Condition (R1)"}],"minor_comments":[{"comment":"The runtime comparison sentence contains a typo: 'LazyVInonlinear required approximately 21,534 seconds in the linear setting and 33,766 seconds in the linear setting.' The second 'linear setting' should presumably read 'nonlinear setting.'","section":"§4.1.3"},{"comment":"The class of functions defined in equation (55) is denoted M, but the lemma statement uses M before its definition; the notation is confusing. Also, the covering-number bound should make explicit the dependence on the Lipschitz constant of the activation, which is stated but not reflected in the constants.","section":"Lemma 34"},{"comment":"Two consecutive paragraphs give conflicting summaries: the first says 'both halves were significant with p<0.001' and identifies region 10 as most important, while the second says the top half had p≈0.388 and the Bottom-Left quadrant was not significant. The text should be reconciled.","section":"§4.3.3"},{"comment":"The line 'Ensure::' contains a double colon typo.","section":"Algorithm 1"}],"recommendation":"reject","confidential_remarks":"The main theorem's proof is not valid as written: the Donsker verification in (R3) fails because the class grows with W_n, and the spectral assumption (A1) is a conjecture for deep ReLU networks at fitted parameters. Even if the authors could supply a triangular-array Donsker argument and prove A1 in a nontrivial setting, the present manuscript would require substantial new work. I would not recommend acceptance in the current form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper extends the lazy-training variable-importance framework from regression (Gao et al. 2022) to binary classification with logistic loss, and claims an op(n^{-1/2}) residual via local Rademacher complexity. That is a natural and useful step, and the simulation work plus the MNIST and ADNI applications give the method a tangible feel. The writing is careful and the authors are transparent about what they cannot prove, especially the spectral decay assumption (A1) at fitted parameters.\n\nThe soft spots are real and load-bearing. The verification of Williamson et al.'s condition (R3) in the proof of Theorem 25 defines a class D = {loss differences indexed by f in F} and shows that, for a fixed width W, its entropy integral is finite using the bound in Theorem 36. But the paper assumes W = o(sqrt(n)/log n), so W grows with n, and the entropy bound carries a factor sqrt(W). There is no fixed P0-Donsker class D_0 containing the g_n's with probability tending to 1. The empirical process argument therefore does not deliver the asymptotic linear expansion. This gap is independent of A1; even granting A1, the Donsker verification is incomplete.\n\nThere is also a norm mismatch in (R1). Williamson et al.'s condition requires ||hat f_n - f_0||_F = o_p(n^{-1/4}), where ||·||_F is the function-class norm. The paper verifies this with the L2(P) norm instead, and never defines a norm on the DNN class F that would make the two comparable. That may be fixable, but as written it is not the same condition.\n\nThe spectral assumption (A1) is explicitly unproved at fitted parameters, and the authors only extrapolate from initialization results and provide an empirical log-log plot. That weakens the advertised sharper rate, though it is at least disclosed. The penalty selection in the simulations is not shown to satisfy A2.\n\nI would not cite the main theorem as it stands. The empirical method and the framing are worth knowing about, but the core theoretical claim needs a genuinely different argument for growing W. I would send this to a serious referee: the problem is important, the extension is sensible, and the gaps are clearly identified rather than hidden. With substantial revision, it could be a solid contribution.","headline":"A plausible and well-motivated extension of lazy-training variable importance to binary classification, but the central theorem is not fully established because the Donsker condition (R3) is verified only for fixed network width W, which grows with n.","tokens_in":66322,"tokens_out":2989,"would_cite":false,"duration_ms":34068,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G10","62G20","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"Lazy training reduces deep-network feature-importance testing to a z-test.","keywords":["variable importance","lazy training","neural tangent kernel","local Rademacher complexity","binary classification","hypothesis testing","deep ReLU networks","asymptotic normality"],"falsifier":"Compute the sorted eigenvalues of the NTK matrix K_{-S} for a deep ReLU network at converged parameters on a moderate-size nonlinear classification problem and fit the log-log slope of the tail. A slope with alpha <= 1, or a heavy tail that persists as network width grows, would invalidate the local-Rademacher rate and the claimed op(n^{-1/2}) result.","tokens_in":65298,"feed_emoji":"🧠","tokens_out":4051,"duration_ms":43564,"temperature":0.7,"pith_summary":"The paper tries to establish that, for binary classification with deep ReLU networks, you can test whether a feature subset matters without retraining the network for every candidate subset. It shows that under a spectral decay condition on the neural tangent kernel and a slow-growth condition on the number of network weights, the estimated variable-importance measure is asymptotically linear with a remainder that is smaller than the usual n^{-1/2} sampling noise. This makes the importance estimate divided by its standard error asymptotically standard normal under the null, so feature significance becomes a classical z-test. The algorithm trains one network, linearizes it around the fitted parameters using the neural tangent kernel, and solves a penalized logistic regression for each feature subset, avoiding expensive retraining. Simulations and an MNIST experiment support the claimed behavior.","feed_headline":"Lazy training turns feature-importance testing into a z-test","feed_subtitle":"Deep ReLU classifiers can get valid feature-importance z-tests from one training run plus kernel ridge regressions.","key_machinery":"The neural tangent kernel (NTK) matrix K_{-S}, formed from the gradients of the trained network evaluated at inputs with the candidate feature subset replaced by its mean. The lazy-regime linearization h_theta approximately equal to h_theta-f plus the gradient inner product with (theta - theta-f) turns each feature-removal fit into a ridge-regularized logistic regression in the RKHS of the NTK, solved by Newton-Raphson. The convergence argument rides on the local Rademacher complexity of the RKHS ball H_B, whose fixed point is controlled by Assumption (A1): the eigenvalues mu_j of K_{-S}/n decay as mu_j <= c j^{-alpha} with alpha > 1. That eigenvalue decay is what converts the estimation err","core_discovery":"The central claim is Theorem 25: for any candidate feature set S, under the null hypothesis and with W = o(sqrt(n)/log n) network weights, the estimated variable-importance difference psi-hat minus psi-0 equals a sample average of influence-function contrasts plus op(n^{-1/2}). If this holds, the lazy variable-importance statistic is asymptotically normal with a variance that can be estimated from the same fitted network, so a valid z-test for feature importance can be run without refitting deep networks for each feature set. The paper's key claimed improvement is the op(n^{-1/2}) remainder rather than the Op(n^{-1/2}) rate achieved in prior work, attributed to the use of local Rademacher co","pith_inferences":["If Assumption (A1) transfers from initialization to fitted parameters, the same framework may extend to overparameterized networks; the paper identifies this as an open bottleneck rather than a proven result.","A natural stress test is to run LazyVI on nonlinear simulations with eigenvalue decay alpha close to 1, where the local-Rademacher rate becomes borderline and the z-test approximation should visibly degrade.","The n-by-n NTK matrix is the computational bottleneck; Nystrom or block approximations to K_{-S} could plausibly scale the method to biobank-sized data, a direction the paper mentions but does not implement."],"forward_implications":["Feature-importance inference for binary deep classifiers requires only one full-network training pass; every candidate feature subset is handled by a cheap linearized logistic regression.","Under the null hypothesis, the LazyVI test statistic is asymptotically standard normal, enabling p-values and confidence intervals without retraining or permutation tests.","The required regime is underparameterized: the number of network weights must grow slower than sqrt(n)/log n.","The likelihood-based setup extends to other exponential-family outcomes whenever the negative log-likelihood satisfies the convexity and Lipschitz conditions used in the proofs."],"fun_headline_variants":["Lazy training yields valid feature-importance z-tests","Single net run, valid z-test for feature importance","Lazy training sharpens feature-importance inference","One training pass, feature-importance z-test ready","Bypass refits: lazy feature-importance z-tests"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise, which the authors state they cannot currently prove for deep ReLU networks, is that the neural tangent kernel matrix at the fitted parameters has eigenvalues decaying as mu_j <= c j^{-alpha} with alpha > 1; without this spectral decay, the local-Rademacher rate and the op(n^{-1/2}) remainder collapse.","fun_headline_variants_meta":{"raw":{"variants":["Lazy training yields valid feature-importance z-tests","Single net run, valid z-test for feature importance","Lazy training sharpens feature-importance inference","One training pass, feature-importance z-test ready","Bypass refits: lazy feature-importance z-tests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1110,"prompt_tokens":645,"completion_tokens":465,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":389,"completion_tokens_details":{"reasoning_tokens":386}},"tokens_in":389,"tokens_out":465,"duration_ms":5156,"temperature":1.0,"reasoning_tokens":386,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T03:57:20.882448+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the sorted eigenvalues of the NTK matrix K_{-S} for a deep ReLU network at converged parameters on a moderate-size nonlinear classification problem and fit the log-log slope of the tail. A slope with alpha <= 1, or a heavy tail that persists as network width grows, would invalidate the local-Rademacher rate and the claimed op(n^{-1/2}) result.","supporting_citations":[],"review_version":1}