Pith. sign in

REVIEW 2 major objections 5 minor 102 references

Same-sample higher-order influence functions stay √n-consistent and more numerically stable than sample-split versions for bilinear functionals when k = o(n).

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 14:13 UTC pith:KJSK7B2V

load-bearing objection Solid theory for same-sample HOIFs of bilinear forms: the combinatorial analysis is real, the bias rate is weaker, and the practical claim rests on one toy simulation. the 2 major comments →

arxiv 2607.04743 v2 pith:KJSK7B2V submitted 2026-07-06 math.ST econ.EMstat.MEstat.MLstat.TH

Stabilized Higher-Order Influence Functions: Statistical Theory of a Class of Bilinear Forms

classification math.ST econ.EMstat.MEstat.MLstat.TH MSC 62G0562G2062E20
keywords higher-order influence functionsbilinear formsU-statisticsMöbius inversionGram matrixfunctional estimationcausal inferenceenumerative combinatorics
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper targets bilinear functionals of the form ψ = μ⊤Ωη, where Ω is the inverse Gram matrix of a k-dimensional basis of covariates. These forms cover common targets such as quadratic density functionals, signal-to-noise ratios, treatment-specific means, and generalized covariance measures. Prior empirical higher-order influence-function estimators achieve √n rates when k is smaller than n, but they invert a sample Gram matrix on a held-out split and become numerically unstable as k/n grows. The authors replace that split with a same-sample inverse and prove that the resulting estimator remains √n-consistent and asymptotically normal under the same regime, while simulations show far better finite-sample stability. The technical workhorse is a Möbius inversion that rewrites each higher-order U-statistic kernel as a sum of lower-order multiplicative kernels whose moments are controlled by the first Betti numbers of associated graphs.

Core claim

Under standard moment, eigenvalue, and L∞-stability assumptions, the same-sample stabilized HOIF estimator ψ̂_{m,k}(Ω̂) of the bilinear form ψ = μ⊤Ωη is √n-consistent and asymptotically normal whenever k ≲ n/log^{3}n and the correction order m grows like log n; its bias is of order (mk/n)⌈(m−1)/4⌉ and its variance is of order 1/n + k/n^{2}, matching the guarantees previously known only for the sample-split empirical HOIF.

What carries the argument

The Möbius inversion decomposition of each order-j HOIF kernel on the partition lattice of the interior indices, which rewrites the same-sample U-statistic as a finite sum of lower-order multiplicative kernels whose moments are bounded by graph-counting of first Betti numbers after a leave-out Neumann expansion of Ω̂.

Load-bearing premise

The projection onto the span of the k basis functions must stay uniformly bounded in the supremum norm, independently of both dimension and sample size, and the two outcome variables must be almost-surely bounded.

What would settle it

Simulate the same-sample estimator at m ≈ log n with k/n approaching 1 under bounded outcomes and check whether bias falls below n−1/2 and whether the Monte-Carlo variance tracks 1/n + k/n^{2}; a systematic excess of either quantity would falsify Theorem 1.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Practitioners can invert the Gram matrix on the full sample rather than a held-out split and still retain √n asymptotic normality for bilinear targets when k = o(n).
  • The same-sample construction removes the leading source of numerical breakdown that previously limited uptake of empirical HOIFs as ρ = k/n grows.
  • Any smooth functional that admits a bilinear approximation of the form μ⊤Ωη inherits these guarantees once the approximation bias is controlled separately.
  • The Möbius-plus-graph-counting analysis supplies a reusable template for variance bounds of other higher-order U-statistics whose kernels depend on the whole sample through an inverse Gram matrix.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the uniform L∞-stability of the projection can be relaxed to high-probability or average bounds, the same theory would cover many unbounded or heavy-tailed nuisance estimators used in practice.
  • The combinatorial skeleton (partition lattices and Betti numbers) is likely portable to higher-order bias corrections for functionals that are not bilinear, such as those arising from Z-estimation or multi-index models.
  • Once ridge or nonlinear-shrinkage estimators of Ω are substituted for the plain inverse, the same leave-out expansion may yield rates in the proportional regime k ≃ n where the unregularized inverse fails.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper constructs a same-sample (no sample-splitting) higher-order influence function estimator ˆψ_{m,k}(ˆΩ) for bilinear functionals ψ=μ⊤Ωη, with Ω estimated by the inverse sample Gram matrix from the same observations used in the U-statistic. Under Assumptions 1–3 and k=o(n), Theorem 1 gives a bias bound of order (Cmk/n)^{⌈(m−1)/4⌉}, a variance bound of order 1/n+k/n² when k≲n/log³n and m≍log n, and √n-CAN. The analysis relies on a Möbius-inversion decomposition of the HOIF kernels (Lemma 2), Neumann expansion of ˆΩ−I, cancellation of low-order bias terms (Lemma 6), and a graph-counting lemma that bounds moments of multiplicative kernels by the first Betti number (Lemma 3), together with leave-*-out expansions and Efron–Stein for the variance. A small simulation (m=3) illustrates improved numerical stability relative to the sample-split empirical HOIF of Liu et al. (2017).

Significance. The work addresses a genuine practical obstacle to HOIF methods—numerical instability of large sample Gram inverses under sample splitting—while retaining √n-CAN theory for bilinear forms without density estimation. The combinatorial toolkit (Möbius inversion on partition lattices plus Betti-number moment bounds for dependent U-statistic kernels) is a substantive technical contribution and is carefully developed in the appendix. Finite-sample stability gains are demonstrated, albeit in a limited design. If the proofs hold as written, the paper supplies the first rigorous guarantees for same-sample empirical HOIFs in the k=o(n) regime and should be of interest to researchers in causal inference and functional estimation.

major comments (2)
  1. Abstract and §1 claim the new estimators enjoy “similar statistical guarantees” to Liu et al. (2017). Theorem 1(1) gives bias of order (Cmk/n)^{⌈(m−1)/4⌉}, while Proposition 1 gives (k/n)^{m/2}. The weaker exponent is load-bearing for the comparison: under the same (m,k) regime the bias decay is slower, even though √n-CAN still holds when m≍log n and k≲n/log³n. The abstract, introduction, and discussion of Theorem 1 should state the rate difference explicitly and clarify the regimes in which both estimators are √n-CAN, rather than describing the guarantees as similar without qualification.
  2. Assumption 2 (uniform L∞-stability of Π with C_Π independent of k and n) enters every application of the graph-counting lemma (Lemma 3) and thus every bias and variance bound in Theorem 1. Remark 1 labels it technical convenience, but the manuscript never indicates for which standard bases (e.g., truncated Fourier, polynomials, wavelets, or random features) the constant remains free of k. A short discussion or sufficient condition in §2 would make the scope of Theorem 1 clearer and is needed for the result to be usable beyond the abstract bilinear setting.
minor comments (5)
  1. Figure 1 and Appendix A: the simulation is restricted to m=3, n=300, X∼N(0,I), and a single linear signal. The stability claim is plausible but rests on a narrow design; either expand the design slightly or phrase the finite-sample claims more cautiously pending the promised follow-up.
  2. Notation: ˆIF vs IF and the double-index convention ˆIF_{j,j,k} are inherited from prior HOIF papers but are dense for new readers; a short notational table in §1.2 would help.
  3. Lemma 2 / Remark 6: the explicit expansions for j=3,4 are useful; consider moving one fully expanded example into the main text near the statement of Theorem 1 to aid intuition before the proof sketch.
  4. Typos and polish: “of order o(n²)” spacing; occasional missing spaces after commas in displays; “enumerative combinatorics” is listed in keywords and used well—ensure Stanley (2011) and Lauritzen (1996) page or theorem references are precise where Möbius inversion is invoked.
  5. Section 5(1): the conjecture on shrinkage for k≳n is interesting; a one-sentence pointer to which Ledoit–Wolf or ridge results would be the natural starting point would strengthen the outlook.

Circularity Check

0 steps flagged

No significant circularity: Theorem 1 is a self-contained bias/variance/CAN proof for a defined estimator of an external bilinear functional under stated assumptions.

full rationale

The paper’s central claim (Theorem 1) is that the same-sample HOIF estimator ˆψ_{m,k}(ˆΩ) of the external bilinear form ψ=μ⊤Ωη is √n-CAN under Assumptions 1–3 when k=o(n) and m≍log n. The derivation chain is definitional then analytic: define ˆψ via U-statistic kernels with ˆΩ from the same sample; apply Möbius inversion on partition lattices (Lemma 2) to rewrite higher-order terms; expand ˆΩ−I by Neumann series; control moments of multiplicative kernels by a graph-counting bound on the first Betti number (Lemma 3); obtain bias o(n^{−1/2}) and variance ≲1/n+k/n². None of these steps define the target in terms of the estimator, fit free parameters to data and re-label them as predictions, or import a uniqueness theorem that forces the result. Self-citations (Robins et al.; Liu et al. 2017, 2020) supply the HOIF framework, the sample-split baseline (Proposition 1), and the observation that low-order same-sample versions appeared without theory; the new same-sample guarantees are proved from scratch with enumerative combinatorics and leave-*-out analysis. Simulation (Figure 1) is illustrative, not a fitted “prediction.” Score 0 is appropriate: the derivation is independent of its inputs by construction.

Axiom & Free-Parameter Ledger

2 free parameters · 6 axioms · 3 invented entities

The central √n-CAN claim rests on standard U-statistic and matrix concentration tools plus three domain regularity assumptions (Gram eigenvalues, L∞ projection stability, bounded A,Y) and the regime k=o(n) with m of order log n. No free parameters are fitted to establish the theorem; the simulation uses a fixed Gaussian design with true ψ=1. Invented entities are proof devices (graph association for kernels, Möbius collections B_ι), not physical postulates.

free parameters (2)
  • truncation level J = ⌈C0 log n⌉ in Neumann expansion of ˆΩ−I
    Chosen large enough so remainder RJ is negligible under m≍log n and k≲n/log³n; C0 is a universal constant, not fit to data, but the proof depends on this hand-chosen growth.
  • HOIF order m and basis dimension k (regime choices)
    Theory requires m≍log n and k≲n/log³n for the variance and CLT statements; practitioners must choose these. Not fitted to prove the theorem, but free in applications.
axioms (6)
  • domain assumption Assumption 1: E(X⊤X)=O(k), ∥X⊤X∥∞=O(k), eigenvalues of Σ bounded away from 0 and ∞
    Used throughout to reduce to Σ=I w.l.o.g. and to control operator norms in graph counting and matrix Bernstein bounds.
  • domain assumption Assumption 2: projection operator Π is uniformly bounded on L∞ with C_Π independent of k,n
    Standard in prior HOIF work; essential for leaf-integration in the graph-counting lemma (Lemma 15).
  • domain assumption Assumption 3: A and Y bounded almost surely
    Technical convenience (Remark 1); avoids exponential inequalities for unbounded U-statistic kernels.
  • domain assumption k=o(n) with refined rates k≲n/log³n and m≍log n for variance/CLT
    Required so sample Gram inverse is consistent and geometric series in leave-*-out expansions converge.
  • standard math Möbius inversion on partition lattices and matrix Neumann series / Bernstein inequality
    Standard combinatorial and matrix concentration tools (Stanley; Tropp); used as black boxes in Lemmas 2, 26, 27, 28.
  • ad hoc to paper Target is exactly the bilinear form ψ=μ⊤Ωη (approximation bias of true functional by bilinear form ignored)
    Section 2: authors take ˜ψ_k as the target without analyzing approximation bias from the true smooth functional; scope is deliberately restricted to bilinear forms.
invented entities (3)
  • Stabilized same-sample HOIF estimator ˆψ_{m,k}(ˆΩ) independent evidence
    purpose: Estimate bilinear ψ without sample splitting while improving numerical stability via self-normalization of ˆΣ
    Defined in (7); low-order versions appeared in Liu et al. (2020); theory is new.
  • Graph-counting association of multiplicative U-statistic kernels with undirected graphs and first Betti number r(G) no independent evidence
    purpose: Convert moment bounds on products of bilinear forms into combinatorial counts (Lemma 3)
    Proof device built from standard graph theory; not a physical entity. Independent of the statistical claim once the association is fixed.
  • Möbius inversion decomposition of ˆIF_{j,j,k}(ˆΩ) into lower-order U-statistics (Lemma 2) independent evidence
    purpose: Handle dependence of the kernel on ˆΩ and enable bias/variance analysis without sample splitting
    Algebraic identity using ∑_i H_i=0; standard Möbius machinery applied to this estimator.

pith-pipeline@v1.1.0-grok45 · 56429 in / 3928 out tokens · 36033 ms · 2026-07-11T14:13:50.179662+00:00 · methodology

0 comments
read the original abstract

Higher-order influence functions, introduced in a series of articles (Robins et al., 2008, 2009a; van der Vaart, 2014; Robins et al., 2016, 2023; Liu et al., 2017), are a unified framework for constructing rate-optimal point estimates of a class of statistical functionals under various complexity-reducing assumptions on the posited statistical model that generates the observed data. Although higher-order (influence functions) estimators are theoretically appealing, they have very limited practical uptake compared to their first-order counterparts. The original higher-order estimators proposed in Robins et al. (2008) and Robins et al. (2017) involve nonparametric density estimation of multi-dimensional covariates, a highly nontrivial statistical and computational problem on its own. The density estimator is, in turn, used in the evaluation of the inverse population Gram matrix $\Omega$ of a set of $k$-dimensional basis transformations of covariates. There, $k$ is allowed to be as large as $o (n^2)$. To partially address this potential shortcoming, Liu et al. (2017) restrict $k$ to $o (n)$ and instead estimate $\Omega$ directly using the inverse sample Gram matrix estimator, but computed from an independent sample often obtained by sample-splitting. Liu et al. (2017) refer to this alternative estimator as the empirical higher-order estimator. Although the empirical higher-order estimator bypasses density estimation, it suffers from numerical instability due to inverting a large-dimensional sample Gram matrix. In this article, for a class of bilinear forms/functionals that often appear in substantive fields, we propose a new stabilized higher-order estimator without sample splitting, which exhibits more stable finite-sample performance compared to the empirical higher-order estimator. We also prove that this new class of higher-order estimators enjoys similar statistical guarantees.

Figures

Figures reproduced from arXiv: 2607.04743 by Chang Li, Lin Liu, Na Liu, Yujia Gu.

Figure 1
Figure 1. Figure 1: Finite-sample comparison between the sample-split empirical HOIF estimator [PITH_FULL_IMAGE:figures/full_fig_p011_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Schematic overview of the variance analysis. The diagram displays the main algebraic [PITH_FULL_IMAGE:figures/full_fig_p019_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Three graph structures in the covariance decomposition. In each panel, the digits on [PITH_FULL_IMAGE:figures/full_fig_p023_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

102 extracted references · 16 linked inside Pith

  1. [1]

    Econometric methods for program evaluation

    Alberto Abadie and Matias D Cattaneo. Econometric methods for program evaluation. Annual Review of Economics, 10 0 (1): 0 465--503, 2018

  2. [2]

    Computational Complexity: A Modern Approach

    Sanjeev Arora and Boaz Barak. Computational Complexity: A Modern Approach. Cambridge University Press, 2009

  3. [3]

    The fundamental limits of structure-agnostic functional estimation

    Sivaraman Balakrishnan, Edward H Kennedy, and Larry Wasserman. The fundamental limits of structure-agnostic functional estimation. Statistical Science, 2026

  4. [4]

    A class of U -statistics and asymptotic normality of the number of k -clusters

    Rabi N Bhattacharya and Jayanta K Ghosh. A class of U -statistics and asymptotic normality of the number of k -clusters. Journal of Multivariate Analysis, 43 0 (2): 0 300--330, 1992

  5. [5]

    Estimating integrated squared density derivatives: Sharp best order of convergence estimates

    Peter J Bickel and Ya'acov Ritov. Estimating integrated squared density derivatives: Sharp best order of convergence estimates. Sankhy \=a : The Indian Journal of Statistics, Series A , 50 0 (3): 0 381--393, 1988

  6. [6]

    Efficient and Adaptive Estimation for Semiparametric Models

    Peter J Bickel, Chris A J Klaassen, Ya'acov Ritov, and Jon A Wellner. Efficient and Adaptive Estimation for Semiparametric Models. Johns Hopkins Series in the Mathematical Sciences. Springer New York, 1998

  7. [7]

    Fisher-type information involving higher order derivatives

    Sergey G Bobkov. Fisher-type information involving higher order derivatives. arXiv preprint arXiv:2412.10200, 2024

  8. [8]

    Higher order concentration of measure

    Sergey G Bobkov, Friedrich G \"o tze, and Holger Sambale. Higher order concentration of measure. Communications in Contemporary Mathematics, 21 0 (03): 0 1850043, 2019

  9. [9]

    Higher-order N eyman orthogonality in moment-condition models

    St \'e phane Bonhomme, Koen Jochmans, Whitney K Newey, and Martin Weidner. Higher-order N eyman orthogonality in moment-condition models. arXiv preprint arXiv:2605.10842, 2026

  10. [10]

    Fast convergence rates for dose-response estimation

    Matteo Bonvini and Edward H Kennedy. Fast convergence rates for dose-response estimation. arXiv preprint arXiv:2207.11825, 2022

  11. [11]

    Doubly-robust inference and optimality in structure-agnostic models with smoothness

    Matteo Bonvini, Edward H Kennedy, Oliver Dukes, and Sivaraman Balakrishnan. Doubly-robust inference and optimality in structure-agnostic models with smoothness. arXiv preprint arXiv:2405.08525, 2024

  12. [12]

    Adaptive, rate-optimal hypothesis testing in nonparametric IV models

    Christoph Breunig and Xiaohong Chen. Adaptive, rate-optimal hypothesis testing in nonparametric IV models. Econometrica, 92 0 (6): 0 2027--2067, 2024

  13. [13]

    Double robust B ayesian inference on average treatment effects

    Christoph Breunig, Ruixuan Liu, and Zhengfei Yu. Double robust B ayesian inference on average treatment effects. Econometrica, 93 0 (2): 0 539--568, 2025

  14. [14]

    Augmented balancing weights as linear regression

    David Bruns-Smith, Oliver Dukes, Avi Feller, and Elizabeth L Ogburn. Augmented balancing weights as linear regression. Journal of the Royal Statistical Society Series B: Statistical Methodology, 2026

  15. [15]

    Kernel-based semiparametric estimators: Small bandwidth asymptotics and bootstrap consistency

    Matias D Cattaneo and Michael Jansson. Kernel-based semiparametric estimators: Small bandwidth asymptotics and bootstrap consistency. Econometrica, 86 0 (3): 0 955--995, 2018

  16. [16]

    Inference in linear regression models with many covariates and heteroscedasticity

    Matias D Cattaneo, Michael Jansson, and Whitney K Newey. Inference in linear regression models with many covariates and heteroscedasticity. Journal of the American Statistical Association, 113 0 (523): 0 1350--1361, 2018

  17. [17]

    Two-step estimation and inference with possibly many included covariates

    Matias D Cattaneo, Michael Jansson, and Xinwei Ma. Two-step estimation and inference with possibly many included covariates. The Review of Economic Studies, 86 0 (3): 0 1095--1122, 2019

  18. [18]

    Bootstrap inference in the presence of bias

    Giuseppe Cavaliere, S \' lvia Gon c alves, Morten rregaard Nielsen, and Edoardo Zanelli. Bootstrap inference in the presence of bias. Journal of the American Statistical Association, 119 0 (548): 0 2908--2918, 2024

  19. [19]

    Tail bounds for canonical U -statistics and U -processes with unbounded kernels

    Abhishek Chakrabortty and Arun K Kuchibhotla. Tail bounds for canonical U -statistics and U -processes with unbounded kernels. arXiv preprint arXiv:2504.01318, 2025

  20. [20]

    M \"o bius Inversion in Physics

    Nanxian Chen. M \"o bius Inversion in Physics . World Scientific, 2010

  21. [21]

    Method-of-moments inference for GLM s and doubly-robust functionals under proportional asymptotics

    Xingyu Chen, Lin Liu, and Rajarshi Mukherjee. Method-of-moments inference for GLM s and doubly-robust functionals under proportional asymptotics. arXiv preprint arXiv:2408.06103, 2024

  22. [22]

    On computing and the complexity of computing higher-order U -statistics, exactly

    Xingyu Chen, Lin Liu, and Ruiqi Zhang. On computing and the complexity of computing higher-order U -statistics, exactly. arXiv preprint arXiv:2508.12627, 2025

  23. [23]

    Dimension free ridge regression

    Chen Cheng and Andrea Montanari. Dimension free ridge regression. The Annals of Statistics, 52 0 (6): 0 2879--2912, 2024

  24. [24]

    Double/debiased machine learning for treatment and structural parameters

    Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins. Double/debiased machine learning for treatment and structural parameters. The Econometrics Journal, 21 0 (1): 0 C1--C68, 2018

  25. [25]

    Locally robust semiparametric estimation

    Victor Chernozhukov, Juan Carlos Escanciano, Hidehiko Ichimura, Whitney K Newey, and James M Robins. Locally robust semiparametric estimation. Econometrica, 90 0 (4): 0 1501--1535, 2022

  26. [26]

    The generative leap: Tight sample complexity for efficiently learning G aussian multi-index models

    Alex Damian, Jason D Lee, and Joan Bruna. The generative leap: Tight sample complexity for efficiently learning G aussian multi-index models. In Proceedings of The Thirty-ninth Annual Conference on Neural Information Processing Systems, pages 28276--28311, 2025

  27. [27]

    Second-order inference for the mean of a variable missing at random

    Ivan Diaz, Marco Carone, and Mark J van der Laan. Second-order inference for the mean of a variable missing at random. The International Journal of Biostatistics, 12 0 (1): 0 333--349, 2016

  28. [28]

    Functional convergence of sequential U -processes with size-dependent kernels

    Christian D \"o bler, Miko aj J Kasprzak, and Giovanni Peccati. Functional convergence of sequential U -processes with size-dependent kernels. The Annals of Applied Probability, 32 0 (1): 0 551--601, 2022

  29. [29]

    The jackknife estimate of variance

    Bradley Efron and Charles Stein. The jackknife estimate of variance. The Annals of Statistics, 9 0 (3): 0 586--596, 1981

  30. [30]

    Visually communicating and teaching intuition for influence functions

    Aaron Fisher and Edward H Kennedy. Visually communicating and teaching intuition for influence functions. The American Statistician, 75 0 (2): 0 162--172, 2021

  31. [31]

    o tze. Expansions for von M ises functionals. Zeitschrift f \

    Friedrich G \"o tze. Expansions for von M ises functionals. Zeitschrift f \"u r Wahrscheinlichkeitstheorie und verwandte Gebiete , 65: 0 599--625, 1984

  32. [32]

    On the role of the propensity score in efficient semiparametric estimation of average treatment effects

    Jinyong Hahn. On the role of the propensity score in efficient semiparametric estimation of average treatment effects. Econometrica, 66 0 (2): 0 315--331, 1998

  33. [33]

    Functional restriction and efficiency in causal inference

    Jinyong Hahn. Functional restriction and efficiency in causal inference. The Review of Economics and Statistics, 86 0 (1): 0 73--76, 2004

  34. [34]

    Demystifying statistical learning based on efficient influence functions

    Oliver Hines, Oliver Dukes, Karla Diaz-Ordaz, and Stijn Vansteelandt. Demystifying statistical learning based on efficient influence functions. The American Statistician, 76 0 (3): 0 292--304, 2022

  35. [35]

    Learning single index models via harmonic decomposition

    Nirmit Joshi, Hugo Koubbi, Theodor Misiakiewicz, and Nati Srebro. Learning single index models via harmonic decomposition. In Proceedings of the Thirty-ninth Annual Conference on Neural Information Processing Systems, pages 45052--45127, 2026

  36. [36]

    Nonparametric von M ises estimators for entropies, divergences and mutual informations

    Kirthevasan Kandasamy, Akshay Krishnamurthy, Barnab \"y s P \'o czos, Larry Wasserman, and James M Robins. Nonparametric von M ises estimators for entropies, divergences and mutual informations. In Proceedings of the 29th International Conference on Neural Information Processing Systems-Volume 1, pages 397--405, 2015

  37. [37]

    Towards optimal doubly robust estimation of heterogeneous causal effects

    Edward H Kennedy. Towards optimal doubly robust estimation of heterogeneous causal effects. Electronic Journal of Statistics, 17 0 (2): 0 3008--3049, 2023

  38. [38]

    Minimax rates for heterogeneous causal effect estimation

    Edward H Kennedy, Sivaraman Balakrishnan, James M Robins, and Larry Wasserman. Minimax rates for heterogeneous causal effect estimation. The Annals of Statistics, 52 0 (2): 0 793--816, 2024

  39. [39]

    Estimation of smooth functionals in high-dimensional models: Bootstrap chains and G aussian approximation

    Vladimir Koltchinskii. Estimation of smooth functionals in high-dimensional models: Bootstrap chains and G aussian approximation. The Annals of Statistics, 50 0 (4): 0 2386--2415, 2022

  40. [40]

    Estimation of smooth functionals of covariance operators: Jackknife bias reduction and bounds in terms of effective rank

    Vladimir Koltchinskii. Estimation of smooth functionals of covariance operators: Jackknife bias reduction and bounds in terms of effective rank. Annales de l'Institut Henri Poincare (B) Probabilites et statistiques, 61 0 (1): 0 665--712, 2025

  41. [41]

    Estimating learnability in the sublinear data regime

    Weihao Kong and Gregory Valiant. Estimating learnability in the sublinear data regime. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, pages 5460--5469, 2018

  42. [42]

    The M oment- SOS hierarchy: Applications and related topics

    Jean B Lasserre. The M oment- SOS hierarchy: Applications and related topics. Acta Numerica, 33: 0 841--908, 2024

  43. [43]

    Graphical Models, volume 17

    Steffen L Lauritzen. Graphical Models, volume 17. Clarendon Press, 1996

  44. [44]

    Nonlinear shrinkage estimation of large-dimensional covariance matrices

    Olivier Ledoit and Michael Wolf. Nonlinear shrinkage estimation of large-dimensional covariance matrices. The Annals of Statistics, 40 0 (2): 0 1024--1060, 2012

  45. [45]

    Analytical nonlinear shrinkage of large-dimensional covariance matrices

    Olivier Ledoit and Michael Wolf. Analytical nonlinear shrinkage of large-dimensional covariance matrices. The Annals of Statistics, 48 0 (5): 0 3043--3065, 2020

  46. [46]

    When is it worthwhile to jackknife? B reaking the quadratic barrier for Z -estimators

    Licong Lin, Fangzhou Su, Wenlong Mou, Peng Ding, and Martin Wainwright. When is it worthwhile to jackknife? B reaking the quadratic barrier for Z -estimators. arXiv preprint arXiv:2411.02909, 2024

  47. [47]

    New n -consistent, numerically stable empirical higher-order influence function estimators

    Lin Liu and Chang Li. New n -consistent, numerically stable empirical higher-order influence function estimators. arXiv preprint arXiv:2302.08097, 2023

  48. [48]

    Semiparametric efficient empirical higher order influence function estimators

    Lin Liu, Rajarshi Mukherjee, Whitney K Newey, and James M Robins. Semiparametric efficient empirical higher order influence function estimators. arXiv preprint arXiv:1705.07577, 2017

  49. [49]

    On nearly assumption-free tests of nominal confidence interval coverage for causal parameters estimated by machine learning

    Lin Liu, Rajarshi Mukherjee, and James M Robins. On nearly assumption-free tests of nominal confidence interval coverage for causal parameters estimated by machine learning. Statistical Science, 35 0 (3): 0 518--539, 2020

  50. [50]

    Root-n consistent semiparametric learning with high-dimensional nuisance functions under minimal sparsity

    Lin Liu, Xinbo Wang, and Yuhao Wang. Root-n consistent semiparametric learning with high-dimensional nuisance functions under minimal sparsity. arXiv preprint arXiv:2305.04174, 2023

  51. [51]

    Assumption-lean falsification tests of rate double-robustness of double-machine-learning estimators

    Lin Liu, Rajarshi Mukherjee, and James M Robins. Assumption-lean falsification tests of rate double-robustness of double-machine-learning estimators. Journal of Econometrics, 240 0 (2): 0 105500, 2024

  52. [52]

    On the asymptotic inadmissibility of double machine learning estimators under structure-agnostic models

    Lin Liu, Rajarshi Mukherjee, and James M Robins. On the asymptotic inadmissibility of double machine learning estimators under structure-agnostic models. arXiv preprint arXiv:2606.22391, 2026

  53. [53]

    Quantum thermodynamics and semi-definite optimization

    Nana Liu, Michele Minervini, Dhrumil Patel, and Mark M Wilde. Quantum thermodynamics and semi-definite optimization. arXiv preprint arXiv:2505.04514, 2025

  54. [54]

    Double cross-fit doubly robust estimators: Beyond series regression

    Alec McClean, Sivaraman Balakrishnan, Edward H Kennedy, and Larry Wasserman. Double cross-fit doubly robust estimators: Beyond series regression. Journal of the Royal Statistical Society Series B: Statistical Methodology, 2026

  55. [55]

    Tensor Methods in Statistics

    Peter McCullagh. Tensor Methods in Statistics. Monographs on Statistics and Applied Probability. Chapman and Hall/CRC, 2018

  56. [56]

    Nuisance function tuning and sample splitting for optimally estimating a doubly robust functional

    Sean McGrath and Rajarshi Mukherjee. Nuisance function tuning and sample splitting for optimally estimating a doubly robust functional. The Annals of Statistics, 2026

  57. [57]

    Cross-fitting and fast remainder rates for semiparametric estimation

    Whitney K Newey and James M Robins. Cross-fitting and fast remainder rates for semiparametric estimation. arXiv preprint arXiv:1801.09138, 2018

  58. [58]

    Twicing kernels and a small bias property of semiparametric estimators

    Whitney K Newey, Fushing Hsieh, and James M Robins. Twicing kernels and a small bias property of semiparametric estimators. Econometrica, 72 0 (3): 0 947--962, 2004

  59. [59]

    Reconciling model- X and doubly robust approaches to conditional independence testing

    Ziang Niu, Abhinav Chakraborty, Oliver Dukes, and Eugene Katsevich. Reconciling model- X and doubly robust approaches to conditional independence testing. The Annals of Statistics, 52 0 (3): 0 895--921, 2024

  60. [60]

    Asymptotic Expansions for General Statistical Models, volume 31 of Lecture Notes in Statistics

    Johann Pfanzagl. Asymptotic Expansions for General Statistical Models, volume 31 of Lecture Notes in Statistics. Springer Science & Business Media, 1983

  61. [61]

    Estimation in Semiparametric Models: Some Recent Developments, volume 63 of Lecture Notes in Statistics

    Johann Pfanzagl. Estimation in Semiparametric Models: Some Recent Developments, volume 63 of Lecture Notes in Statistics. Springer Science & Business Media, 1990

  62. [62]

    Parametric Statistical Theory

    Johann Pfanzagl. Parametric Statistical Theory. Walter de Gruyter, 2011

  63. [63]

    Concentration of polynomial random matrices via E fron-- S tein inequalities

    Goutham Rajendran and Madhur Tulsiani. Concentration of polynomial random matrices via E fron-- S tein inequalities. In Proceedings of the 2023 Annual ACM-SIAM Symposium on Discrete Algorithms (SODA), pages 3614--3653. SIAM, 2023

  64. [64]

    Semiparametric B ayesian causal inference

    Kolyan Ray and Aad van der Vaart. Semiparametric B ayesian causal inference. The Annals of Statistics, 48 0 (5): 0 2999--3020, 2020

  65. [65]

    Nested M arkov properties for acyclic directed mixed graphs

    Thomas S Richardson, Robin J Evans, James M Robins, and Ilya Shpitser. Nested M arkov properties for acyclic directed mixed graphs. The Annals of Statistics, 51 0 (1): 0 334--361, 2023

  66. [66]

    Achieving information bounds in non and semiparametric models

    Ya'acov Ritov and Peter J Bickel. Achieving information bounds in non and semiparametric models. The Annals of Statistics, 18 0 (2): 0 925--938, 1990

  67. [67]

    Comment: Performance of double-robust estimators when ``inverse probability'' weights are highly variable

    James Robins, Mariela Sued, Quanhong Lei-Gomez, and Andrea Rotnitzky. Comment: Performance of double-robust estimators when ``inverse probability'' weights are highly variable. Statistical Science, 22 0 (4): 0 544--559, 2007

  68. [68]

    Higher order influence functions and minimax estimation of nonlinear functionals

    James Robins, Lingling Li, Eric Tchetgen Tchetgen, and Aad van der Vaart. Higher order influence functions and minimax estimation of nonlinear functionals. In Probability and Statistics: Essays in Honor of David A. Freedman, pages 335--421. Institute of Mathematical Statistics, 2008

  69. [69]

    Quadratic semiparametric von M ises calculus

    James Robins, Lingling Li, Eric Tchetgen Tchetgen, and Aad W van der Vaart. Quadratic semiparametric von M ises calculus. Metrika, 69: 0 227--247, 2009 a

  70. [70]

    Semiparametric minimax rates

    James Robins, Eric Tchetgen Tchetgen, Lingling Li, and Aad van der Vaart. Semiparametric minimax rates. Electronic Journal of Statistics, 3: 0 1305--1321, 2009 b

  71. [71]

    Technical report: Higher order influence functions and minimax estimation of nonlinear functionals

    James Robins, Lingling Li, Eric Tchetgen Tchetgen, and Aad van der Vaart. Technical report: Higher order influence functions and minimax estimation of nonlinear functionals. arXiv preprint arXiv:1601.05820, 2016

  72. [72]

    Estimation of regression coefficients when some regressors are not always observed

    James M Robins, Andrea Rotnitzky, and Lue Ping Zhao. Estimation of regression coefficients when some regressors are not always observed. Journal of the American Statistical Association, 89 0 (427): 0 846--866, 1994

  73. [73]

    Minimax estimation of a functional on a structured high-dimensional model

    James M Robins, Lingling Li, Rajarshi Mukherjee, Eric Tchetgen Tchetgen, and Aad van der Vaart. Minimax estimation of a functional on a structured high-dimensional model. The Annals of Statistics, 45 0 (5): 0 1951--1987, 2017

  74. [74]

    Minimax estimation of a functional on a structured high-dimensional model ( C orrected version)

    James M Robins, Lingling Li, Lin Liu, Rajarshi Mukherjee, Eric Tchetgen Tchetgen, and Aad van der Vaart. Minimax estimation of a functional on a structured high-dimensional model ( C orrected version). arXiv preprint arXiv:1512.02174, 2023

  75. [75]

    Characterization of parameters with a mixed bias property

    Andrea Rotnitzky, Ezequiel Smucler, and James M Robins. Characterization of parameters with a mixed bias property. Biometrika, 108 0 (1): 0 231--238, 2021

  76. [76]

    A note on the relation between one–step, outcome regression and IPW –type estimators of parameters with the mixed bias property

    Andrea Rotnitzky, Ezequiel Smucler, and James M Robins. A note on the relation between one–step, outcome regression and IPW –type estimators of parameters with the mixed bias property. Statistics & Probability Letters, 236 0 (110796), 2026

  77. [77]

    a fer. M \

    Florian Sch \"a fer. M \"o bius inversion and the iterated bootstrap. SIAM Journal on Mathematics of Data Science, 8 0 (2): 0 362--381, 2026

  78. [78]

    Adjusting for nonignorable drop-out using semiparametric nonresponse models

    Daniel O Scharfstein, Andrea Rotnitzky, and James M Robins. Adjusting for nonignorable drop-out using semiparametric nonresponse models. Journal of the American Statistical Association, 94 0 (448): 0 1096--1120, 1999

  79. [79]

    The hardness of conditional independence testing and the generalised covariance measure

    Rajen D Shah and Jonas Peters. The hardness of conditional independence testing and the generalised covariance measure. The Annals of Statistics, 48 0 (3): 0 1514--1538, 2020

  80. [80]

    An efficient algorithm for computing interventional distributions in latent variable causal models

    Ilya Shpitser, Thomas S Richardson, and James M Robins. An efficient algorithm for computing interventional distributions in latent variable causal models. In Proceedings of the Twenty-Seventh Conference on Uncertainty in Artificial Intelligence, pages 661--670, 2011

Showing first 80 references.