Pith. sign in

REVIEW 3 major objections 6 minor 37 references

The V-fold jackknife provides valid confidence intervals for semiparametric estimators using only V refits, without deriving influence functions.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 04:33 UTC pith:WEZWZR6E

load-bearing objection Clean fixed-V t_{V-1} jackknife theorem for RAL estimators, but the flagship HAL extension rests on an assumption the authors admit they haven't verified. the 3 major comments →

arxiv 2607.22493 v1 pith:WEZWZR6E submitted 2026-07-24 stat.ME math.STstat.COstat.MLstat.TH

The V-fold jackknife for semiparametric inference: variance estimation, confidence intervals, and simultaneous confidence bands

classification stat.ME math.STstat.COstat.MLstat.TH MSC 62F4062G2062F12
keywords V-fold jackknifesemiparametric inferenceasymptotic linearityStudentizationconfidence intervalssimultaneous confidence bandsinfluence functionhighly adaptive lasso
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper claims that the grouped, or V-fold, jackknife can replace both influence-function calculations and the bootstrap for a broad class of modern semiparametric estimators. For any regular asymptotically linear estimator of a pathwise differentiable parameter, the Studentized jackknife statistic converges to a t-distribution with V−1 degrees of freedom when V is fixed, so a confidence interval centered at the full-sample estimate with t_{V−1} critical values has asymptotic coverage 1−α. The method needs only V leave-fold-out refits and no analytic influence function, and it remains valid even though the jackknife variance estimator does not converge in probability. The same construction yields simultaneous confidence bands for vector parameters and extends to estimators with diverging influence-function variance and slower-than-√n convergence, where Studentization cancels the unknown rate. If correct, this gives practitioners a cheap, theoretically grounded inference tool for machine-learning-based semiparametric procedures.

Core claim

The central discovery is an algebraic identity: up to negligible remainders, each jackknife pseudo-value IC(v)=VΨ(Pn)−(V−1)Ψ(Pn,−v) equals the average of the influence function over fold v. Because the folds are disjoint, these V fold-level averages are nearly independent, and their sample variance, scaled by n/V, estimates the asymptotic variance. Studentizing by this dispersion produces a t_{V−1} limit for fixed V, even though the variance estimator itself stays random. The paper extends this from scalar parameters to vector parameters (correlated t-like denominators following a Wishart diagonal), to diverging V with a V^{-1/2}-consistent variance estimator, and to generalized asymptotical

What carries the argument

The V-fold jackknife pseudo-value IC_Jack(v)=VΨ(Pn)−(V−1)Ψ(Pn,−v), whose centered values coincide with the fold-level influence-function averages W_v up to o_p(n^{-1/2}) remainders. The Studentized ratio sqrt(V)(ΨJack−Ψ0)/S_Jack then behaves like the usual one-sample t-statistic on V nearly independent Gaussian fold means, giving the t_{V−1} limit. For simultaneous bands, the corresponding vector of componentwise-Studentized statistics converges to an m-dimensional law with V normal fold vectors divided by their componentwise sample standard deviations—not a standard multivariate t—and critical values are simulated from that law.

Load-bearing premise

The load-bearing premise is that every leave-fold-out refit has the same first-order linear behavior as the full-sample fit, with fold-level remainder terms uniformly negligible (o_p(n^{-1/2}) for fixed V); if refitting on V−1 folds changes the estimator's influence function or leaves non-negligible bias, the t_{V−1} limit and all extensions break.

What would settle it

A simulation of a known asymptotically linear estimator with V=2 would settle the fixed-V claim: the Studentized jackknife statistic must have t_1 (Cauchy-type) quantiles at large n and the interval Ψ(Pn)±t_{1,0.975}·cSE must cover at the nominal rate. To test the load-bearing premise, use an estimator whose tuning is reselected on each leave-fold-out sample (e.g., lasso with per-fold cross-validated penalty); if the maximum fold-level remainder is not o_p(n^{-1/2}), coverage should drop below 1−α.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A practitioner needs only V+1 fits of the estimator—one on the full sample and V on leave-fold-out samples—to get confidence intervals with asymptotically nominal coverage, with no influence-function derivation.
  • For fixed small V, intervals are conservative because t_{V−1} tails are heavier; as V grows, intervals narrow and the variance estimator becomes consistent at rate V^{-1/2}.
  • Simultaneous confidence bands for curves (survival, dose-response) can be built from the same V refits by Monte Carlo simulation of the componentwise-Studentized limit, with a bias-corrected version for finite-sample bias.
  • The method covers estimators with non-standard convergence rates, such as highly adaptive lasso dose-response curves, where the unknown diverging influence-curve variance cancels in the Studentized statistic.
  • Simulations show the V-fold jackknife maintains coverage in settings where plug-in influence-function standard errors are anti-conservative or fail numerically.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the fixed-V result is as general as claimed, the procedure could become a default inference fallback for any asymptotically linear estimator with an intractable influence function, potentially reducing reliance on the bootstrap in modern semiparametric practice.
  • The scale-invariance argument suggests the same jackknife recipe should work for other sieve and series estimators whose effective dimension grows with n, provided the generalized asymptotic linearity and fold-stability conditions hold; this is a testable extension beyond the highly adaptive lasso.
  • The effective-rank heuristic for choosing V (e.g., V≥2 times the effective rank of the correlation matrix) points toward a fully data-driven V-selection rule; verifying whether such a rule preserves simultaneous coverage would be a natural follow-up.
  • Since fixed-V inference is conservative by design, averaging the variance estimator over several random fold partitions—mentioned but not developed in the paper—could reduce partition-to-partition variability and tighten intervals while retaining the t_{V−1} correction.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper develops a V-fold (delete-a-group) jackknife procedure for semiparametric inference. For a fixed number of folds V, it proves that, under a standard asymptotically linear expansion with uniformly negligible leave-fold-out remainders, the Studentized jackknife statistic converges to a t-distribution with V−1 degrees of freedom (Theorem 1), yielding confidence intervals even though the jackknife variance estimator does not concentrate. For diverging V, consistency of the variance estimator at rate V^{−1/2} is established under a fold-level remainder-difference condition (Theorem 2). The paper further constructs simultaneous confidence bands from a componentwise-Studentized Wishart-type limit (Theorems 4–5) and extends the framework to generalized asymptotically linear estimators with diverging influence-function variance, where scale invariance removes the need to know the convergence rate (Theorems 7–8). Simulations cover TMLE for the average treatment effect, the Kaplan–Meier survival curve, and HAL-based dose-response curves.

Significance. If the main theorems are correct, the paper makes a genuinely useful contribution: it gives a computationally cheap alternative to the bootstrap that requires only V refits and no analytic influence-function derivation, and it supplies the correct fixed-V t_{V−1} degrees-of-freedom correction for a broad class of regular asymptotically linear estimators. The proof structure for Theorem 1 is clean and the diverging-V variance decomposition is persuasive. The paper is also unusually honest about its own limitations, explicitly stating in Section 7.4 that Assumption 1 is not rigorously verified for exact data-adaptive HAL and in Appendix H that the structural verification is oracle-based. The code is made available. These strengths make the fixed-V result and the diverging-V consistency results valuable even if the slower-rate HAL extension remains conditional rather than fully proven.

major comments (3)
  1. [Section 7.4 and Assumption 1 (Eq. 15)] The slower-rate extension (Theorems 7–8, Corollaries 3–4) and the Section 8.3 HAL application rest on Assumption 1, in particular the uniform fold-level remainder condition max_v |R_n(P_{n,-v},P0)| = o_p(σ_n n^{-1/2}). Section 7.4 states that rigorous verification for exact, data-adaptive HAL is beyond scope and calls Assumption 1 only 'mathematically plausible.' Because the leave-fold-out HAL refits involve data-adaptive basis selection and penalty, this premise is not established. The abstract and introduction advertise valid inference for modern ML estimators through this slower-rate regime; the manuscript should either supply a proof or explicitly downgrade the HAL inference claims to empirical/conditional status in the abstract and conclusions.
  2. [Appendix H; Section 5.2 and Theorem 2 preamble] The main text says Appendix H gives 'sufficient TMLE, one-step, and AIPW-style structural conditions' for the diverging-V remainder condition, and the discussion in Sections 5.2 and 9.1 repeats this with the V=O(log n) conclusion. However, Appendix H's own Assumption 2 explicitly proceeds under an oracle formulation treating D_n as conditionally fixed and notes that fully data-adaptive verification requires additional stochastic equicontinuity conditions. As written, the main text's repeated reference to 'sufficient conditions' will be read by most readers as a theorem for data-adaptive estimators. The status of Appendix H should be flagged in the main text as a heuristic/oracle derivation for data-adaptive nuisance estimators, not a proof.
  3. [Section 7.3] The claim that all simultaneous-band results 'hold verbatim' under Assumption 1 applied componentwise is too quick. Theorem 4's limit T∞ depends on a fixed correlation matrix R0. Under Assumption 1 the correlation matrix R_{0,n} = D_n^{-1} Σ_n D_n^{-1} may itself depend on n; the paper does not prove convergence of R_{0,n} or provide a uniform approximation. Consequently the simultaneous HAL bands reported in Table 6 do not have the same formal support as the fixed-m RAL case. The authors should either state an additional condition under which R_{0,n} converges (or is uniformly approximated) or present the HAL simultaneous bands as heuristic.
minor comments (6)
  1. [Section 5.2] The quantity T_n is described as a 'cross-fold term' but it is actually a within-fold off-diagonal sum. Rename to avoid confusion with cross-fitting terminology.
  2. [Assumption 1] The Lindeberg condition uses the threshold |φ_n| > ε σ_n √n. For the fixed-V Theorem 7, convergence of the fold means only requires the weaker threshold with √(n/V). If the stronger condition is intended, say so explicitly; otherwise state the weaker condition and note that the √n version is sufficient.
  3. [Section 6.4, Eq. (10)] The concentration inequality for the sample covariance matrix is asserted without proof and without a precise statement of conditions. Since ThatΣ_Jack is not exactly the sample covariance of i.i.d. oracle fold vectors in finite samples, clarify whether (10) is a heuristic or a formal result for the jackknife covariance, and provide a proof or reference if formal.
  4. [Appendix H] The internal numbering uses G.1–G.5 inside an appendix labeled H. Fix the numbering so the sections correspond to the appendix label.
  5. [Section 8.1, Table 2] Coverage differences of 1–2 percentage points across methods are within Monte Carlo error for 500 replications. Reporting standard errors (e.g., ±2%) or confidence intervals for coverage would help the reader assess which differences are meaningful.
  6. [Section 7.4] When citing van der Laan (2023) for 'precise theorems' on oracle-model score approximation, give the specific theorem/proposition numbers so the reader can verify that the conditions match the leave-fold-out HAL setting rather than the full-sample setting.

Circularity Check

0 steps flagged

No significant circularity: core V-fold jackknife theorems are self-contained conditional on stated RAL/remainder assumptions; HAL extension is explicitly conditional and not a definitional reduction.

full rationale

The paper's central derivation is not circular. Theorem 1 and Corollary 1 are proved from an assumed asymptotically linear expansion, exact algebraic cancellation of the full-sample remainder in the centered pseudo-values, the CLT for independent fold means, and the convergence of \hat S_Jack to the oracle fold variance S_W up to o_p(n^{-1}). Lemma 1 and Lemma 2 establish the two key equivalences by direct calculation, not by assuming the t_{V-1} conclusion. Theorem 2's diverging-V variance consistency is likewise an exact decomposition plus a remainder bound, and Theorem 4's simultaneous limit is a multivariate CLT statement. The slower-rate results (Theorems 7-8) are explicit conditionals: 'Under Assumption 1, for any fixed V, the Studentized V-fold jackknife statistic adapts to the unknown convergence rate...' The proof rescales by the diverging sigma_n and uses the Lindeberg condition; this is an assumption-driven theorem, not a self-justifying loop. The HAL-specific application is where the paper relies on prior work by the same authors and is candid about its limits. Section 7.4 states: 'A rigorous theoretical verification for exact, data-adaptive HAL implementations is beyond the scope of this section, because the selected basis functions are themselves data-adaptive,' and calls Assumption 1 only 'mathematically plausible.' This is a real gap/limitation and a correctness risk for the HAL inference claim, but it is not circularity: the V-fold jackknife procedure does not build the t_{V-1} result into its definition, and the HAL asymptotically linear representation is imported as an external input from van der Laan (2023) and Shi et al. (2024), not derived from the jackknife. Those citations are load-bearing for the HAL application, but they are external theorems with stated assumptions rather than a renaming of the present paper's output. No equation is defined in terms of the target result, and no fitted parameter is relabeled as a prediction. The fixed-V result against an external benchmark (Kaplan-Meier Greenwood comparison) is also self-contained, supporting a non-circularity finding.

Axiom & Free-Parameter Ledger

1 free parameters · 6 axioms · 0 invented entities

The central theoretical machinery uses standard asymptotic linearity and Studentization; the only user-chosen quantity is the number of folds V. No new physical or statistical entities are introduced. The main external dependence is the oracle HAL working-model theory from the authors' own prior work, which is invoked rather than re-derived.

free parameters (1)
  • number of folds V = 5, 10, 20 in simulations
    User-selected tuning parameter, not fitted to data. Affects interval width and, for bands, correlation matrix estimation; the theorem holds for any fixed V≥2, while formal plug-in band validity requires V→∞.
axioms (6)
  • domain assumption Asymptotic linear representation: Ψhat(Pn)−Ψ(P0) = Pnφ + R(Pn,P0) with E[φ]=0, E[φ²]<∞, R = o_p(n^{−1/2})
    Invoked in Section 3 and underpins Theorems 1-2; not proven for any estimator in the paper, and for machine-learning estimators it is exactly the unverified condition.
  • domain assumption Fold-wise remainder stability: max_v |R(Pn,−v,P0)| = o_p(n^{−1/2}) for fixed V, and ∥dn∥V,2 = o_p((√n V)^{-1}) for diverging V
    Theorems 1, 2, and 5 rely on these bounds; the paper does not derive them from primitive conditions except through Appendix H sufficient (and partly heuristic) bounds.
  • ad hoc to paper Exactly balanced folds with n divisible by V
    Section 3 assumes exact balance; approximate balance is dismissed as negligible but not formally treated.
  • domain assumption Assumption 1: generalized asymptotic linearity with diverging σn², Lindeberg condition, and scaled remainders
    Theorems 7-8 rest on this; Section 7.4 says rigorous verification for exact data-adaptive HAL is beyond scope.
  • domain assumption Weak law of large numbers for n^{-1}σn^{-2}Σφn² → 1
    Assumed in Theorem 8 and implied by Eφ^4 = o(n σn^4); needed for diverging-V consistency in the slower-rate setting.
  • domain assumption Fixed m for simultaneous bands, with V→∞ needed for consistent estimation of the correlation matrix
    Theorems 4-5; at fixed V, R0 is not identifiable from data, as acknowledged in Section 6.4.

pith-pipeline@v1.3.0-alltime-deepseek · 45972 in / 20318 out tokens · 201779 ms · 2026-08-01T04:33:52.544806+00:00 · methodology

0 comments
read the original abstract

For decades, the bootstrap has been a default tool for statistical inference because of its broad applicability and minimal analytic requirements. Although its validity is well understood for smooth parametric estimators, its theoretical properties for many modern semiparametric and machine-learning estimators remain largely unstudied. Nevertheless, bootstrap procedures are often used routinely in such settings, even when their validity is unknown and their computational cost is substantial. We develop the $V$-fold jackknife as a computationally efficient and theoretically justified alternative for semiparametric inference. It requires only $V$ leave-fold-out refits and uses the empirical dispersion of jackknife pseudo-values to quantify uncertainty, without deriving or evaluating an influence function. For regular asymptotically linear estimators of pathwise differentiable parameters, we show that, for fixed $V$, the Studentized $V$-fold jackknife statistic converges to a $t$-distribution with $V-1$ degrees of freedom, giving valid confidence intervals even though the jackknife variance estimator does not converge in probability. When $V\to\infty$, we establish consistency of the variance estimator at rate $V^{-1/2}$, allowing $V$ to diverge slowly, for example at rate $\log n$. We also develop simultaneous confidence bands based on the correct componentwise-Studentized limiting distribution. Finally, we extend the theory to generalized asymptotically linear estimators with diverging influence-function variance and slower-than-$\sqrt n$ convergence; scale invariance of Studentization eliminates the need to know the effective convergence rate. Simulations on the average treatment effect, Kaplan--Meier survival curve, and highly adaptive lasso dose-response curves confirm reliable inference, including where influence-function-based standard errors are anti-conservative or unstable.

Figures

Figures reproduced from arXiv: 2607.22493 by Ashkan Ertefaie, Mark van der Laan, Yi Li.

Figure 1
Figure 1. Figure 1: Coverage (top) and CI width (bottom) of V -fold jackknife confi￾dence intervals as a function of the number of folds V , for four scenarios and two sample sizes. Horizontal dashed line in the top panels: nominal 95% level. Faded horizontal lines: coverage of the EIC-based and bootstrap methods for each scenario. The V -fold jackknife maintains coverage above the nominal level for moderate positivity and su… view at source ↗
Figure 2
Figure 2. Figure 2: Pointwise coverage of 95% confidence intervals for the Kaplan– Meier survival function across evaluation times (n = 200, 500 replications). The V -fold jackknife with t9 critical values (V = 10) closely tracks the Green￾wood formula, and the bootstrap Wald method performs comparably [PITH_FULL_IMAGE:figures/full_fig_p024_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Pointwise coverage of 95% confidence intervals across the dose grid (20 equally spaced points on [0, 5], i.e. A ∈ {0, 0.263, 0.526, . . . , 5}) for HAL dose-response curve estimation (n = 500, 500 replications). Each point on the horizontal axis represents a dose level at which the dose-response curve E[Y | A = a, W] is evaluated and coverage is computed separately. Rows: DGP 1 (smooth), DGP 2 (oscillating… view at source ↗
Figure 4
Figure 4. Figure 4: Bias diagnosis for HAL dose-response curve estimation (n = 500, 500 replications). Solid black line: absolute bias | ¯ ψˆ(a)−ψ0(a)| at each of the 20 evaluation points, where ¯ ψˆ(a) is the simulation average of the point estimate. Dashed blue line: mean estimated standard error from the V -fold jackknife (V = 20). When the bias exceeds the standard error, the asymptotic linearity condition is violated and… view at source ↗
Figure 5
Figure 5. Figure 5: Eigenvalue spectrum of the estimated correlation matrix R0 for the Kaplan–Meier survival curve (m = 91 time points) and the HAL dose-response curve (m = 20 evaluation points, four DGPs, CV-HAL and undersmoothed variants). The correlation matrix is estimated from the 500×m matrix of point estimates across simulation replications. (a) Normalized eigenvalue spectrum (fraction of trace m), truncated at index 2… view at source ↗
Figure 6
Figure 6. Figure 6: True dose-response curves ψ0(a) = E[Y | do(A = a)] for the four data generating processes used in the HAL simulation study. DGP 1: smooth monotone increase; DGP 2: oscillating; DGP 3: smooth for a ≤ 2, oscillating for a > 2; DGP 4: jump discontinuities at a = 2 and a = 4. True values computed by Monte Carlo integration (106 samples per evaluation point). qα is substantially larger than the pointwise tV −1 … view at source ↗
Figure 7
Figure 7. Figure 7: Pointwise coverage (%) of simultaneous 95% confidence bands for the Kaplan–Meier survival function across evaluation times (n = 200, 500 replications). Methods: V -fold simultaneous band with V ∈ {5, 10, 20} and the studentized bootstrap simultaneous band. Horizontal dashed line: nominal 95% level. All methods achieve 98–100% pointwise coverage uniformly across time points, with no systematic differences a… view at source ↗
Figure 8
Figure 8. Figure 8: Pointwise coverage (%) of simultaneous 95% confidence bands for the HAL dose-response curve across the dose grid (n = 500, 500 replications). Rows: DGP 1 (smooth), DGP 2 (oscillating), DGP 3 (piecewise), DGP 4 (dis￾continuous). Columns: CV-HAL and undersmoothed HAL. Solid lines: stan￾dard simultaneous bands; dashed lines: bias-corrected simultaneous bands. Colors indicate V ∈ {5, 10, 20}. Horizontal dashed… view at source ↗
Figure 9
Figure 9. Figure 9: Pointwise coverage (%) of 95% confidence intervals for the HAL dose-response curve, comparing three centering strategies (V = 20, n = 500, 500 replications). Standard: centered at Ψˆ j (Pn); Jack-centered (JC): centered at Ψˆ (j) Jack(Pn) = Ψˆ j (Pn) − ˆbj ; Bias-corrected (BC): centered at Ψˆ j (Pn) with in￾terval widened by | ˆbj |. The JC intervals have lower coverage than the standard intervals in all … view at source ↗
Figure 10
Figure 10. Figure 10: Pointwise coverage (%) of simultaneous 95% confidence bands for the HAL dose-response curve, comparing three centering strategies (V = 20, n = 500, 500 replications). Same conventions as fig. 9 but using the simultaneous critical value ˆqα in place of tV −1, 0.975. The coverage degradation from JC centering is even more pronounced for simultaneous bands [PITH_FULL_IMAGE:figures/full_fig_p053_10.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

37 extracted references · 2 linked inside Pith

  1. [1]

    The Annals of Statistics , year =

    Efron, Bradley , title =. The Annals of Statistics , year =

  2. [2]

    , title =

    Efron, Bradley and Tibshirani, Robert J. , title =

  3. [3]

    Quenouille, M. H. , title =. Journal of the Royal Statistical Society, Series B , year =

  4. [4]

    , title =

    Tukey, John W. , title =. The Annals of Mathematical Statistics , year =

  5. [5]

    van der Vaart, A. W. , title =. 1998 , series =

  6. [6]

    , title =

    Arvesen, James N. , title =. The Annals of Mathematical Statistics , year =

  7. [7]

    , title =

    Brillinger, David R. , title =. Review of the International Statistical Institute , year =

  8. [8]

    , title =

    Miller, Rupert G. , title =. Biometrika , year =

  9. [9]

    Wu, C. F. Jeff , title =. The Annals of Statistics , year =

  10. [10]

    Shao, Jun and Wu, C. F. Jeff , title =. The Annals of Statistics , year =

  11. [11]

    , title =

    Kott, Phillip S. , title =. Proceedings of the Section on Survey Research Methods , organization =. 1998 , pages =

  12. [12]

    , title =

    Kott, Phillip S. , title =. Journal of Official Statistics , year =

  13. [13]

    and Garren, Steven T

    Kott, Phillip S. and Garren, Steven T. , title =. Journal of Official Statistics , year =

  14. [14]

    Biometrika , year =

    Yang, Shu and Pieper, Karen and Cools, Frank , title =. Biometrika , year =

  15. [15]

    , title =

    van der Laan, Mark J. , title =. 2019 , note =

  16. [16]

    , title =

    Cai, Weixin and van der Laan, Mark J. , title =. Biometrics , year =

  17. [17]

    Assessing the Causal Effect of Policies: An Example Using Stochastic Interventions , journal =

    D. Assessing the Causal Effect of Policies: An Example Using Stochastic Interventions , journal =. 2013 , volume =

  18. [18]

    Simultaneous Confidence Bands: Theory, Implementation, and an Application to

    Montiel Olea, Jos. Simultaneous Confidence Bands: Theory, Implementation, and an Application to. Journal of Applied Econometrics , year =

  19. [19]

    2022 , note =

    Lam, Henry , title =. 2022 , note =

  20. [20]

    Cheap Subsampling Bootstrap Confidence Intervals for Fast and Robust Inference , year =

    Ohlendorff, Johan Sebastian and Munch, Anders and S. Cheap Subsampling Bootstrap Confidence Intervals for Fast and Robust Inference , year =

  21. [21]

    and Romano, Joseph P

    Politis, Dimitris N. and Romano, Joseph P. and Wolf, Michael , title =

  22. [22]

    , title =

    Gruber, Susan and van der Laan, Mark J. , title =. Journal of Statistical Software , year =

  23. [23]

    , title =

    Tran, Linh and Petersen, Maya and Schwab, Joshua and van der Laan, Mark J. , title =. Journal of Causal Inference , year =

  24. [24]

    Davison, A. C. and Hinkley, D. V. , title =. 1997 , series =

  25. [25]

    and van der Laan, Mark J

    Shi, Junming and Zhang, Wenxin and Hubbard, Alan E. and van der Laan, Mark J. , title =. arXiv preprint arXiv:2406.05607 , year =

  26. [26]

    , title =

    van der Laan, Mark J. , title =. arXiv preprint arXiv:2301.13354 , year =

  27. [27]

    The Econometrics Journal , year =

    Chernozhukov, Victor and Chetverikov, Denis and Demirer, Mert and Duflo, Esther and Hansen, Christian and Newey, Whitney and Robins, James , title =. The Econometrics Journal , year =

  28. [28]

    , title =

    Zheng, Wenjing and van der Laan, Mark J. , title =. Targeted Learning: Causal Inference for Observational and Experimental Data , editor =. 2011 , pages =

  29. [29]

    Anderson, T. W. , title =

  30. [30]

    Journal of Theoretical Probability , year =

    Vershynin, Roman , title =. Journal of Theoretical Probability , year =

  31. [31]

    Bernoulli , year =

    Koltchinskii, Vladimir and Lounici, Karim , title =. Bernoulli , year =

  32. [32]

    The Annals of Statistics , year =

    Chernozhukov, Victor and Chetverikov, Denis and Kato, Kengo , title =. The Annals of Statistics , year =

  33. [33]

    Probability Theory and Related Fields , year =

    Chen, Xiaohui and Kato, Kengo , title =. Probability Theory and Related Fields , year =

  34. [34]

    Predictive Inference with the Jackknife+ , journal =

    Barber, Rina Foygel and Cand\`. Predictive Inference with the Jackknife+ , journal =. 2021 , volume =

  35. [35]

    2004 , isbn =

    Kotz, Samuel and Nadarajah, Saralees , title =. 2004 , isbn =

  36. [36]

    Scandinavian Journal of Statistics , volume =

    Non-and semi-parametric maximum likelihood estimators and the von mises method (part 1)[with discussion and reply] , author =. Scandinavian Journal of Statistics , volume =. 1989 , publisher =

  37. [37]

    , title =

    Praestgaard, Jens and Wellner, Jon A. , title =. The Annals of Probability , year =