{"id":"d85b646b-baa8-48ab-ae2f-0950dfa0284b","arxiv_id":"2607.22493","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"The V-fold jackknife, Studentized with t_{V-1} critical values, yields valid confidence intervals and simultaneous bands for semiparametric estimators and extends to slower-rate estimators by scale invariance.","lead":"This paper develops a V-fold jackknife for confidence intervals and simultaneous bands in semiparametric estimation, needing only V refits instead of the hundreds used by the bootstrap. It shows the Studentized jackknife statistic follows a t-distribution with V-1 degrees of freedom, offering a cheap, influence-function-free route to valid inference.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Slower-rate HAL extension rests on Assumption 1, which Section 7.4 admits is unverified for exact data-adaptive HAL; if fold-level remainder stability fails, the flagship HAL inference lacks asymptotic coverage.","rationale":"The reader's weakest-assumption analysis identifies exactly the same load-bearing point: the advertised extension to HAL and slower-rate estimators depends on Assumption 1, and Section 7.4 explicitly disclaims rigorous verification for exact data-adaptive HAL. The core fixed-V result for regular asymptotically linear estimators is mathematically sound under its stated conditions, so no fatal flaw appears there. However, the paper's central narrative—that the V-fold jackknife gives valid inference where influence-function-based standard errors are anti-conservative or unstable—leans heavily on the HAL application, and that application is only as strong as Assumption 1. Because the authors are transparent about the gap, a conditional verdict is appropriate, not rejection. The concrete test of increasing n in the HAL simulation would either provide empirical support for the assumption or reveal a clear violation, thereby settling whether the concern actually lands. I therefore keep the reader's CONDITIONAL verdict unchanged.","tokens_in":46267,"tokens_out":11172,"duration_ms":117848,"concrete_test":"Use the public code to run the HAL dose-response simulation (Section 8.3, DGP 2, undersmoothed HAL) at increasing sample sizes n=500, 1000, 2000, and 5000 with V fixed at 10 and 20. Compute the empirical distribution of the Studentized statistic √V(Ψhat_Jack − Ψ0)/Shat_Jack and the empirical coverage of the t_{V−1} intervals. If coverage does not approach 1−α and the empirical quantiles do not move toward t_{V−1} as n grows, then Assumption 1's fold-level remainder stability for exact data-adaptive HAL is not supported, and the slower-rate theorems should be presented as conditional on an unverified premise. Additionally, in the same simulation, estimate the fold-level remainder RMS ||d_n||_{V,2} by comparing full-sample and leave-fold-out errors (with Ψ0 known), and check whether it is o_p(σ_n/(√n V)) for V=log n; a clear violation would directly falsify the diverging-V condition in The","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's broadest advertised advantage—valid inference for modern ML estimators without influence-function derivations—depends not on the clean fixed-V result for regular asymptotically linear estimators (Theorem 1), but on the generalized slower-rate extension (Theorems 7–8, Corollaries 3–4). These theorems require Assumption 1, in particular the uniform fold-level remainder condition max_v |R_n(P_{n,-v},P0)| = o_p(σ_n n^{-1/2}) and, for diverging V, ||d_n||_{V,2} = o_p(σ_n/(√n V)). Section 7.4 explicitly says that rigorous verification for exact, data-adaptive HAL is beyond scope and calls Assumption 1 only 'mathematically plausible.' For leave-fold-out HAL refits, the selected basis functions and penalty are themselves data-adaptive and can differ from the full-sample fit; the paper supplies an oracle-model heuristic, not a proof that empirical HAL scores approximate the oracle scores uniformly across all V folds. If this condition fails, then the pointwise t_{V-1} intervals and simultaneous bands for HAL dose-response curves—the paper's flagship application where analytic standard errors fail—have no asymptotic coverage guarantee. The fixed-V Theorem 1 for RAL estimators is not affected, but the central claim that the V-fold jackknife provides theoretically justified inference for semiparametric and ML estimators is weakened because its most practically novel regime rests on an explicitly unverified assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a V-fold (delete-a-group) jackknife procedure for semiparametric inference. For a fixed number of folds V, it proves that, under a standard asymptotically linear expansion with uniformly negligible leave-fold-out remainders, the Studentized jackknife statistic converges to a t-distribution with V−1 degrees of freedom (Theorem 1), yielding confidence intervals even though the jackknife variance estimator does not concentrate. For diverging V, consistency of the variance estimator at rate V^{−1/2} is established under a fold-level remainder-difference condition (Theorem 2). The paper further constructs simultaneous confidence bands from a componentwise-Studentized Wishart-type limit (Theorems 4–5) and extends the framework to generalized asymptotically linear estimators with diverging influence-function variance, where scale invariance removes the need to know the convergence rate (Theorems 7–8). Simulations cover TMLE for the average treatment effect, the Kaplan–Meier survival curve, and HAL-based dose-response curves.","tokens_in":46635,"tokens_out":13261,"duration_ms":143449,"significance":"If the main theorems are correct, the paper makes a genuinely useful contribution: it gives a computationally cheap alternative to the bootstrap that requires only V refits and no analytic influence-function derivation, and it supplies the correct fixed-V t_{V−1} degrees-of-freedom correction for a broad class of regular asymptotically linear estimators. The proof structure for Theorem 1 is clean and the diverging-V variance decomposition is persuasive. The paper is also unusually honest about its own limitations, explicitly stating in Section 7.4 that Assumption 1 is not rigorously verified for exact data-adaptive HAL and in Appendix H that the structural verification is oracle-based. The code is made available. These strengths make the fixed-V result and the diverging-V consistency results valuable even if the slower-rate HAL extension remains conditional rather than fully proven.","major_comments":[{"comment":"The slower-rate extension (Theorems 7–8, Corollaries 3–4) and the Section 8.3 HAL application rest on Assumption 1, in particular the uniform fold-level remainder condition max_v |R_n(P_{n,-v},P0)| = o_p(σ_n n^{-1/2}). Section 7.4 states that rigorous verification for exact, data-adaptive HAL is beyond scope and calls Assumption 1 only 'mathematically plausible.' Because the leave-fold-out HAL refits involve data-adaptive basis selection and penalty, this premise is not established. The abstract and introduction advertise valid inference for modern ML estimators through this slower-rate regime; the manuscript should either supply a proof or explicitly downgrade the HAL inference claims to empirical/conditional status in the abstract and conclusions.","section":"Section 7.4 and Assumption 1 (Eq. 15)"},{"comment":"The main text says Appendix H gives 'sufficient TMLE, one-step, and AIPW-style structural conditions' for the diverging-V remainder condition, and the discussion in Sections 5.2 and 9.1 repeats this with the V=O(log n) conclusion. However, Appendix H's own Assumption 2 explicitly proceeds under an oracle formulation treating D_n as conditionally fixed and notes that fully data-adaptive verification requires additional stochastic equicontinuity conditions. As written, the main text's repeated reference to 'sufficient conditions' will be read by most readers as a theorem for data-adaptive estimators. The status of Appendix H should be flagged in the main text as a heuristic/oracle derivation for data-adaptive nuisance estimators, not a proof.","section":"Appendix H; Section 5.2 and Theorem 2 preamble"},{"comment":"The claim that all simultaneous-band results 'hold verbatim' under Assumption 1 applied componentwise is too quick. Theorem 4's limit T∞ depends on a fixed correlation matrix R0. Under Assumption 1 the correlation matrix R_{0,n} = D_n^{-1} Σ_n D_n^{-1} may itself depend on n; the paper does not prove convergence of R_{0,n} or provide a uniform approximation. Consequently the simultaneous HAL bands reported in Table 6 do not have the same formal support as the fixed-m RAL case. The authors should either state an additional condition under which R_{0,n} converges (or is uniformly approximated) or present the HAL simultaneous bands as heuristic.","section":"Section 7.3"}],"minor_comments":[{"comment":"The quantity T_n is described as a 'cross-fold term' but it is actually a within-fold off-diagonal sum. Rename to avoid confusion with cross-fitting terminology.","section":"Section 5.2"},{"comment":"The Lindeberg condition uses the threshold |φ_n| > ε σ_n √n. For the fixed-V Theorem 7, convergence of the fold means only requires the weaker threshold with √(n/V). If the stronger condition is intended, say so explicitly; otherwise state the weaker condition and note that the √n version is sufficient.","section":"Assumption 1"},{"comment":"The concentration inequality for the sample covariance matrix is asserted without proof and without a precise statement of conditions. Since ThatΣ_Jack is not exactly the sample covariance of i.i.d. oracle fold vectors in finite samples, clarify whether (10) is a heuristic or a formal result for the jackknife covariance, and provide a proof or reference if formal.","section":"Section 6.4, Eq. (10)"},{"comment":"The internal numbering uses G.1–G.5 inside an appendix labeled H. Fix the numbering so the sections correspond to the appendix label.","section":"Appendix H"},{"comment":"Coverage differences of 1–2 percentage points across methods are within Monte Carlo error for 500 replications. Reporting standard errors (e.g., ±2%) or confidence intervals for coverage would help the reader assess which differences are meaningful.","section":"Section 8.1, Table 2"},{"comment":"When citing van der Laan (2023) for 'precise theorems' on oracle-model score approximation, give the specific theorem/proposition numbers so the reader can verify that the conditions match the leave-fold-out HAL setting rather than the full-sample setting.","section":"Section 7.4"}],"recommendation":"major_revision","confidential_remarks":"The paper's core fixed-V result (Theorem 1) is sound and the honest treatment of limitations is commendable. My main concern is that the most practically novel advertised regime—HAL-based inference under slower convergence—relies on an assumption the authors explicitly do not verify for the actual data-adaptive estimator, and the surrounding text sometimes presents this as established. A revision that either supplies the missing verification or clearly recalibrates the HAL claims to 'conditional/empirical' would be appropriate for this journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What should you know: the fixed-V result is the real thing; the HAL extension is not yet proved, and the authors know it.\n\nThe paper's core contribution is Theorem 1: for a regular asymptotically linear estimator, the V-fold jackknife pseudo-values behave like fold-level means of the influence function, and Studentizing with the dispersion across folds gives a t distribution with V-1 degrees of freedom, even though the variance estimator doesn't converge. That extends Brillinger's old result from MLEs to the full RAL class and is useful on its own. The diverging-V consistency at rate 1/sqrt(V) is also handled cleanly. The simultaneous band construction is genuinely new: instead of a standard multivariate t, it uses the correct Wishart-diagonal denominator structure, and that matters for small V. The slower-rate extension is elegant: because the Studentized ratio is scale-invariant, the unknown diverging variance cancels, so the t_{V-1} limit holds without knowing the rate. The simulations are extensive and the code is public; I believe the paper is honest about what it does not prove.\n\nThe soft spot is exactly where the paper is most ambitious. Section 7.4 states outright that Assumption 1—the generalized asymptotic linearity with uniformly small fold-level remainders—is not rigorously verified for exact data-adaptive HAL. Theorems 7 and 8 and the dose-response application rest on that assumption. The stress-test note is right: if the leave-fold-out HAL fits do not satisfy the oracle-like score stability condition, the flagship HAL confidence intervals have no asymptotic coverage guarantee. That is not a fatal flaw in the fixed-V theorem, but it means the paper's broadest advertised advantage—valid inference for ML estimators without influence-function derivations—is not yet supported in the slower-rate regime. The simultaneous-band story is similar: formal plug-in coverage needs V to diverge; the effective-rank argument for small V is heuristic, though the simulations and the KM example suggest it works.\n\nWho should read it: anyone doing inference for TMLE/AIPW-type estimators or thinking about cheap alternatives to the bootstrap, and the HAL people who want a standard error that doesn't fall over. It deserves a serious referee, but the referee should press hard on Assumption 1 and ask for a real verification or a clear narrowing of the claims. I would cite the fixed-V theorem and maybe the diverging-V consistency; I'd be careful about citing the HAL application.","headline":"Clean fixed-V t_{V-1} jackknife theorem for RAL estimators, but the flagship HAL extension rests on an assumption the authors admit they haven't verified.","tokens_in":47099,"tokens_out":2334,"would_cite":true,"duration_ms":24738,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F40","62G20","62F12"],"pacs":[],"model":"deepseek-v4-flash","headline":"The V-fold jackknife provides valid confidence intervals for semiparametric estimators using only V refits, without deriving influence functions.","keywords":["V-fold jackknife","semiparametric inference","asymptotic linearity","Studentization","confidence intervals","simultaneous confidence bands","influence function","highly adaptive lasso"],"falsifier":"A simulation of a known asymptotically linear estimator with V=2 would settle the fixed-V claim: the Studentized jackknife statistic must have t_1 (Cauchy-type) quantiles at large n and the interval Ψ(Pn)±t_{1,0.975}·cSE must cover at the nominal rate. To test the load-bearing premise, use an estimator whose tuning is reselected on each leave-fold-out sample (e.g., lasso with per-fold cross-validated penalty); if the maximum fold-level remainder is not o_p(n^{-1/2}), coverage should drop below 1−α.","tokens_in":46175,"feed_emoji":"📊","tokens_out":9191,"duration_ms":89662,"temperature":0.7,"pith_summary":"This paper claims that the grouped, or V-fold, jackknife can replace both influence-function calculations and the bootstrap for a broad class of modern semiparametric estimators. For any regular asymptotically linear estimator of a pathwise differentiable parameter, the Studentized jackknife statistic converges to a t-distribution with V−1 degrees of freedom when V is fixed, so a confidence interval centered at the full-sample estimate with t_{V−1} critical values has asymptotic coverage 1−α. The method needs only V leave-fold-out refits and no analytic influence function, and it remains valid even though the jackknife variance estimator does not converge in probability. The same construction yields simultaneous confidence bands for vector parameters and extends to estimators with diverging influence-function variance and slower-than-√n convergence, where Studentization cancels the unknown rate. If correct, this gives practitioners a cheap, theoretically grounded inference tool for machine-learning-based semiparametric procedures.","feed_headline":"V-fold jackknife gives valid confidence intervals with just V refits","feed_subtitle":"Studentized jackknife statistics converge to t with V−1 degrees of freedom, skipping analytic influence functions.","key_machinery":"The V-fold jackknife pseudo-value IC_Jack(v)=VΨ(Pn)−(V−1)Ψ(Pn,−v), whose centered values coincide with the fold-level influence-function averages W_v up to o_p(n^{-1/2}) remainders. The Studentized ratio sqrt(V)(ΨJack−Ψ0)/S_Jack then behaves like the usual one-sample t-statistic on V nearly independent Gaussian fold means, giving the t_{V−1} limit. For simultaneous bands, the corresponding vector of componentwise-Studentized statistics converges to an m-dimensional law with V normal fold vectors divided by their componentwise sample standard deviations—not a standard multivariate t—and critical values are simulated from that law.","core_discovery":"The central discovery is an algebraic identity: up to negligible remainders, each jackknife pseudo-value IC(v)=VΨ(Pn)−(V−1)Ψ(Pn,−v) equals the average of the influence function over fold v. Because the folds are disjoint, these V fold-level averages are nearly independent, and their sample variance, scaled by n/V, estimates the asymptotic variance. Studentizing by this dispersion produces a t_{V−1} limit for fixed V, even though the variance estimator itself stays random. The paper extends this from scalar parameters to vector parameters (correlated t-like denominators following a Wishart diagonal), to diverging V with a V^{-1/2}-consistent variance estimator, and to generalized asymptotical","pith_inferences":["If the fixed-V result is as general as claimed, the procedure could become a default inference fallback for any asymptotically linear estimator with an intractable influence function, potentially reducing reliance on the bootstrap in modern semiparametric practice.","The scale-invariance argument suggests the same jackknife recipe should work for other sieve and series estimators whose effective dimension grows with n, provided the generalized asymptotic linearity and fold-stability conditions hold; this is a testable extension beyond the highly adaptive lasso.","The effective-rank heuristic for choosing V (e.g., V≥2 times the effective rank of the correlation matrix) points toward a fully data-driven V-selection rule; verifying whether such a rule preserves simultaneous coverage would be a natural follow-up.","Since fixed-V inference is conservative by design, averaging the variance estimator over several random fold partitions—mentioned but not developed in the paper—could reduce partition-to-partition variability and tighten intervals while retaining the t_{V−1} correction."],"forward_implications":["A practitioner needs only V+1 fits of the estimator—one on the full sample and V on leave-fold-out samples—to get confidence intervals with asymptotically nominal coverage, with no influence-function derivation.","For fixed small V, intervals are conservative because t_{V−1} tails are heavier; as V grows, intervals narrow and the variance estimator becomes consistent at rate V^{-1/2}.","Simultaneous confidence bands for curves (survival, dose-response) can be built from the same V refits by Monte Carlo simulation of the componentwise-Studentized limit, with a bias-corrected version for finite-sample bias.","The method covers estimators with non-standard convergence rates, such as highly adaptive lasso dose-response curves, where the unknown diverging influence-curve variance cancels in the Studentized statistic.","Simulations show the V-fold jackknife maintains coverage in settings where plug-in influence-function standard errors are anti-conservative or fail numerically."],"fun_headline_variants":["V-fold jackknife gives valid CIs without influence functions","Fixed-V jackknife yields t_{V−1} confidence intervals","V refits enough: jackknife pseudo-values quantify uncertainty","Jackknife alternative to bootstrap for semiparametric inference"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that every leave-fold-out refit has the same first-order linear behavior as the full-sample fit, with fold-level remainder terms uniformly negligible (o_p(n^{-1/2}) for fixed V); if refitting on V−1 folds changes the estimator's influence function or leaves non-negligible bias, the t_{V−1} limit and all extensions break.","fun_headline_variants_meta":{"raw":{"variants":["V-fold jackknife gives valid CIs without influence functions","Fixed-V jackknife yields t_{V−1} confidence intervals","V refits enough: jackknife pseudo-values quantify uncertainty","Jackknife alternative to bootstrap for semiparametric inference"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000449,"raw_usage":{"total_tokens":2164,"prompt_tokens":868,"completion_tokens":1296,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":612,"completion_tokens_details":{"reasoning_tokens":1223}},"tokens_in":612,"tokens_out":1296,"duration_ms":12781,"temperature":1.0,"reasoning_tokens":1223,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T04:33:52.544806+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A simulation of a known asymptotically linear estimator with V=2 would settle the fixed-V claim: the Studentized jackknife statistic must have t_1 (Cauchy-type) quantiles at large n and the interval Ψ(Pn)±t_{1,0.975}·cSE must cover at the nominal rate. To test the load-bearing premise, use an estimator whose tuning is reselected on each leave-fold-out sample (e.g., lasso with per-fold cross-validated penalty); if the maximum fold-level remainder is not o_p(n^{-1/2}), coverage should drop below 1−α.","supporting_citations":[],"review_version":1}