Pith. sign in

REVIEW 4 major objections 7 minor 40 references

A Machine-Learning-Compatible Omnibus Test for Treatment Effect Heterogeneity

T0 review · 4 major / 7 minor · reviewed 2026-07-08 · glm-5.2

Pith's one-line read One Test Detects Whether Treatment Effects Vary — With ML Built In

desk verdict Solid omnibus heterogeneity test with a real gap in verifiability for the LATE extension read the letter →

arxiv 2607.06412 v1 pith:VLHHMSDG submitted 2026-07-07 econ.EM

classification econ.EM
keywords testheterogeneityefficientempiricalincludingnullomnibusprocedure
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes a single nonparametric test that asks whether a treatment effect varies systematically across observed covariates, while allowing the nuisance functions needed for identification to be estimated by arbitrary machine-learning methods. The core construction reformulates the null hypothesis of constant conditional average treatment effects as a continuum of unconditional moment restrictions, following integrated conditional moment principles. These moments are built from cross-fitted doubly robust scores, so that first-stage ML estimation error does not distort the test's asymptotic distribution provided the nuisance estimators converge faster than n^{-1/4} in L2 norm. The test statistic reduces to a quadratic form involving a kernel matrix and the estimated score vector, and critical values are obtained via a weighted bootstrap that reuses the single set of fitted nuisance functions across all iterations, avoiding repeated ML estimation. The authors prove valid asymptotic size and consistency for average treatment effects under unconfoundedness, and extend the same construction to difference-in-differences designs (panel and repeated cross-sections) and local average treatment effects under instrumental variables, deriving the corresponding null distributions and bootstrap procedures for each.

What carries the argument

The test statistic S_n = (1/n) * U-hat^T * F_rho * U-hat, where U-hat is the vector of cross-fitted doubly robust scores and F_rho is an n x n kernel matrix derived from the Fourier transform of the weighting measure rho. Bootstrap critical values are obtained by drawing iid weights xi with mean 1 and variance 1, recomputing only the low-dimensional quantities (score vector and kernel matrix), and taking quantiles of the bootstrap statistic S_n*.

What would settle it

In a Monte Carlo simulation where the ML nuisance estimators fail to achieve n^{-1/4} convergence (e.g., due to high-dimensional covariates with complex nonlinear relationships and default-tuned methods), the test's rejection rate under the null should deviate substantially from the nominal significance level, indicating that the bootstrap critical values are invalid.

Watch

Extended reading notes

Core claim

The central object is the test statistic S_n, defined as the integral over a continuum of index parameters t of the squared empirical moments of the cross-fitted doubly robust score against analytic transformations of the heterogeneity covariates. The discovery is that, under the n^{-1/4} convergence rate condition on nuisance estimators, Neyman orthogonality insulates this statistic from first-stage ML error, its null distribution converges to a well-defined Gaussian process limit, and a computationally lightweight weighted bootstrap that never re-estimates nuisance functions provides asymptotically correct critical values. The same architecture extends across four major causal designs — un

Load-bearing premise

The load-bearing condition is that the ML estimators of propensity scores and outcome regressions converge to their true functions faster than n^{-1/4} in L2 norm. If this rate is not achieved in finite samples — and the paper does not verify it for specific ML methods on the empirical datasets used — the bootstrap distribution may not approximate the null correctly, potentially inflating the test's false positive rate.

Editorial extensions

If this is right

  • Researchers can formally test for treatment-effect heterogeneity with respect to policy-relevant covariates while using flexible ML methods (random forests, neural networks, lasso) for high-dimensional confounding adjustment, without the test's validity depending on correct specification of the nuisance models.
  • The computational efficiency of the bootstrap — which requires a single pass of ML estimation and then only matrix operations — makes the test feasible for large datasets where re-estimating nuisance functions hundreds of times would be prohibitive.
  • The modular extension to DiD and LATE settings means the same diagnostic tool can be applied across the most common empirical designs in policy evaluation, providing a unified first-step test before proceeding to estimate the structure of heterogeneity.
  • The test fills a gap between estimating full conditional average treatment effect functions (which is ambitious and inference-heavy) and running simple subgroup analyses (which may miss nonlinear or interaction-driven heterogeneity).
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. This paper proposes a nonparametric omnibus test for treatment-effect heterogeneity that is compatible with machine-learning nuisance estimators. The test reformulates the null of constant conditional average treatment effects (with respect to a low-dimensional subvector X_c) as a continuum of unconditional moment restrictions using Bierens-type integrated conditional moment (ICM) principles, constructs a doubly-robust cross-fitted score, and develops a computationally efficient weighted bootstrap that avoids re-estimating nuisance functions across iterations. The framework is extended from ATE under unconfoundedness to DiD (panel and repeated cross-sections) and LATE settings. Six propositions establish asymptotic null distributions and bootstrap validity; Monte Carlo simulations and two empirical applications illustrate finite-sample performance.

Significance. The paper addresses a well-motivated gap: a global, omnibus test for systematic heterogeneity that works with ML nuisance estimators and avoids re-estimation in the bootstrap loop. The computationally efficient bootstrap (Eqs. 6–7 and analogues) is a practical contribution. The breadth of designs (ATE, DiD panel, DiD repeated cross-sections, LATE) is a strength. The simulation evidence in Tables 1 and B.1, showing near-nominal size and good power, and the comparison with the BLP method of Chernozhukov et al. (2025) are valuable. The ICM construction is standard but its combination with doubly-robust cross-fitting and the efficient bootstrap across multiple designs is novel and useful.

major comments (4)
  1. All six propositions (3.1, 3.2, 4.1–4.6) are stated with proofs referenced as being in supplementary material that is not included in the arXiv submission. Since the propositions are the central theoretical contributions, the proofs must be provided for review. This is especially important for Propositions 4.3–4.4 (repeated cross-sections) and 4.5–4.6 (LATE), where the influence functions contain non-trivial delta-method terms and the bootstrap schemes differ structurally from the ATE case.
  2. The LATE bootstrap (Section 4.3) uses a different centering scheme than the ATE bootstrap (Section 3.3). In the ATE case, the bootstrap subtracts Ũ*_i = V̂_i − δ̂*, where δ̂* = E_n{ξ_i V̂_i} depends on bootstrap weights, providing a centering mechanism that replicates the ι_t term in the influence function. In the LATE case, the bootstrap subtracts the original (non-bootstrapped) Û_LATE,i = V̂^Y_i − δ̂_LATE V̂^D_i, which is fixed across bootstrap draws. The delta-method centering in the LATE influence function (Proposition 4.5) contains the term E{V^D ϕ_t(X_c)}/E{V^D} multiplied by (V^Y − E{V^Y} − δ_0(V^D − E{V^D})). Whether the bootstrapped ratio δ̂*_LATE = E_n{ξ_i V̂^Y_i}/E_n{ξ_i V̂^D_i} inside Û*_LATE,i correctly replicates this full structure is non-trivial. The supplementary proof should explicitly verify that the bootstrap conditional law converges to the same Gaussian process as f
  3. There are no simulation results for the LATE setting (Propositions 4.5–4.6). The simulations in Section 5 cover only the repeated-cross-section DiD design, and Appendix B covers ATE under unconfoundedness. Given that the LATE bootstrap involves the ratio estimator and the structurally different centering scheme noted above, simulation evidence for size and power in the LATE case would substantially strengthen the paper's claims.
  4. Assumption 3.4(i) (and its analogues, Assumptions 4.5(i), 4.8(i), 4.14(i)) requires n^{-1/4} convergence of nuisance estimators in L2 norm. The paper cites Farrell et al. (2021) for neural networks and Scornet et al. (2015) for random forests. However, the simulations and empirical applications use lasso and random forests with default-style tuning. Whether these specific implementations achieve the required rate in finite samples is not verified. While this is a standard assumption in the double-ML literature, a brief discussion of which specific ML methods and tuning choices are known to satisfy this rate, and any practical diagnostics, would strengthen the practical guidance.
minor comments (7)
  1. Section 5.2 and Appendix B.2: the simulations use a fixed bandwidth h=1 for the Gaussian kernel. No sensitivity analysis to the bandwidth choice is provided. Given that h is a free parameter, a brief discussion of how the test performs under alternative bandwidths, or guidance on practical selection, would be helpful.
  2. The notation switches between U, V, Ũ, Û, Ũ*, Û* across sections. A consolidated notation table or more consistent naming convention would improve readability.
  3. In Section 4.2, the bootstrap scheme for repeated cross-sections involves both bλ* and eλ* (and similarly for π and δ). The distinction between the 'star' and 'e-star' quantities is explained but somewhat terse. A brief intuitive explanation of why two different recentering quantities are needed would help readers.
  4. Figure 2 (401(k) application): the x-axis labels show 'Number of covariates' from 1 to 9, but the specific covariates added at each step are not listed. Including the covariate names (or providing them in a table or caption) would make the empirical results more interpretable.
  5. The paper does not discuss the choice of number of cross-fitting folds N. The simulations use N=5, but no guidance is given on how N affects performance or whether the choice matters in practice.
  6. Table 1 and Table B.1: it would be useful to report Monte Carlo standard errors for the rejection rates to allow readers to assess whether deviations from nominal size are statistically significant.
  7. The reference to Escanciano (2026) is to a working paper presented at a conference. The authors should update the citation if a written version becomes available and clarify the relationship to that work's results more precisely.

Simulated Author's Rebuttal

4 responses · 0 unresolved

We thank the referee for a careful and constructive report. The referee raises four major points: (1) proofs referenced as supplementary material are not included in the arXiv submission; (2) the LATE bootstrap centering scheme differs structurally from the ATE case and requires explicit verification; (3) no simulation evidence is provided for the LATE setting; and (4) practical guidance on which ML methods satisfy the n^{-1/4} convergence rate is needed. We address each point below. We agree with all four comments and will incorporate the requested revisions in the next version of the manuscript.

read point-by-point responses
  1. Referee: All six propositions are stated with proofs referenced as being in supplementary material that is not included in the arXiv submission. Since the propositions are the central theoretical contributions, the proofs must be provided for review, especially for Propositions 4.3–4.4 (repeated cross-sections) and 4.5–4.6 (LATE), where the influence functions contain non-trivial delta-method terms and the bootstrap schemes differ structurally from the ATE case.

    Authors: We agree completely. The proofs were omitted from the arXiv submission due to an oversight in assembling the submission files. We will include the full supplementary material containing all proofs in the revised arXiv version. We confirm that detailed proofs exist for all six propositions, including the repeated-cross-section propositions (4.3–4.4) and the LATE propositions (4.5–4.6). The proofs for the LATE case in particular contain a delta-method expansion that handles the ratio structure of the estimator, and the bootstrap validity proof explicitly verifies that the conditional law of the bootstrap statistic converges to the same Gaussian process as the original statistic. We will ensure these are available for the referee's assessment. revision: yes

  2. Referee: The LATE bootstrap uses a different centering scheme than the ATE bootstrap. In the ATE case, the bootstrap subtracts Ũ*_i = V̂_i − δ̂*, where δ̂* depends on bootstrap weights. In the LATE case, the bootstrap subtracts the original (non-bootstrapped) Û_LATE,i, which is fixed across bootstrap draws. Whether the bootstrapped ratio δ̂*_LATE = E_n{ξ_i V̂^Y_i}/E_n{ξ_i V̂^D_i} inside Û*_LATE,i correctly replicates the full delta-method structure is non-trivial. The supplementary proof should explicitly verify that the bootstrap conditional law converges to the same Gaussian process as the original statistic.

    Authors: The referee correctly identifies a structural difference between the ATE and LATE bootstrap schemes. In the ATE case, the bootstrap centering subtracts a bootstrap-weight-dependent quantity δ̂* = E_n{ξ_i V̂_i}, which replicates the ι_t centering term in the influence function. In the LATE case, the bootstrap recomputes the ratio δ̂*_LATE = E_n{ξ_i V̂^Y_i}/E_n{ξ_i V̂^D_i} using bootstrap weights, and the centering subtracts the original Û_LATE,i = V̂^Y_i − δ̂_LATE V̂^D_i, which is fixed across bootstrap draws. The key step in the proof is to show that the linearization of the bootstrap ratio around the original ratio produces the delta-method correction term E{V^D ϕ_t(X_c)}/E{V^D} multiplied by (V^Y − E{V^Y} − δ_0(V^D − E{V^D})), which matches the influence function in Proposition 4.5. Specifically, a first-order expansion of δ̂*_LATE around δ̂_LATE yields δ̂*_LATE − δ̂_LATE ≈ (E_n{ξ_i V̂^Y_i} − δ̂_LATE E_n{ξ_i V̂^D_i})/E_n{V̂^D_i}, and the resulting bootstrap centered statistic Û*_LATE,i − Û_LATE,i decomposes into terms that replicate the influence function structure. The full verification is in the supplementary proof, which we will include in the revised submission. We will also add a brief discussion in the main text explaining why the LATE bootstrap scheme correctly replicates the delta-method structure, to make this point more transparent to the reader. revision: yes

  3. Referee: There are no simulation results for the LATE setting (Propositions 4.5–4.6). The simulations in Section 5 cover only the repeated-cross-section DiD design, and Appendix B covers ATE under unconfoundedness. Given that the LATE bootstrap involves the ratio estimator and the structurally different centering scheme noted above, simulation evidence for size and power in the LATE case would substantially strengthen the paper's claims.

    Authors: We agree that simulation evidence for the LATE case would strengthen the paper, particularly given the ratio structure and the different bootstrap centering scheme. We will add LATE simulations to the revised manuscript. The DGP will feature a binary instrument, endogenous treatment take-up, and compliers whose treatment effects vary with X_c under the alternative. We will report rejection rates under the null and power curves under the alternative, using the same ML methods (lasso and random forests) and sample sizes as in the existing simulations. This will allow direct comparison of finite-sample performance across the three designs (ATE, DiD, LATE) and provide empirical evidence that the LATE bootstrap, despite its structural differences from the ATE case, achieves near-nominal size and good power. revision: yes

  4. Referee: Assumption 3.4(i) and its analogues require n^{-1/4} convergence of nuisance estimators in L2 norm. The paper cites Farrell et al. (2021) for neural networks and Scornet et al. (2015) for random forests. However, the simulations and empirical applications use lasso and random forests with default-style tuning. Whether these specific implementations achieve the required rate in finite samples is not verified. A brief discussion of which specific ML methods and tuning choices are known to satisfy this rate, and any practical diagnostics, would strengthen the practical guidance.

    Authors: The referee raises a valid point about the gap between the asymptotic rate condition and the specific implementations used in practice. We will add a discussion to the revised manuscript addressing which ML methods are known to satisfy the n^{-1/4} L2 convergence rate and under what conditions. Specifically: (1) For lasso, the rate is established under approximate sparsity conditions on the nuisance functions, with the rate depending on the sparsity index and regularization parameter choice (e.g., Belloni et al., 2014; Chernozhukov et al., 2018). Cross-validation-based tuning, as used in our simulations, is standard in practice though the theoretical results typically require specific penalty choices. (2) For random forests, Scornet et al. (2015) establish consistency under specific conditions; Wager and Athey (2018) provide rates for honest forests. The n^{-1/4} rate requires honesty and subsampling, which may not hold for all default implementations. (3) For neural networks, Farrell et al. (2021) establish the rate under specific architecture and regularization conditions. We will also note that the n^{-1/4} rate is a sufficient condition for the asymptotic theory, and the Neyman orthogonality of the score provides additional robustness: the first-order expansion holds as long as the product of nuisance estimation errors converges faster than n^{-1/2}, which can be achieved even when individual nuisance estimators converge at rates slower than n^{-1/4} if the errors are sufficiently uncorrelated. We will add a brief remark on practical diagnostics, such as comparing cross-fitted nuisance estimates across folds or using sample splitting to assess stability. revision: yes

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity found; derivation chain uses external results at each load-bearing step

full rationale

The paper's derivation chain proceeds through well-separated stages, each relying on external results rather than self-citation or definitional equivalence. (1) The null H0: δ_c(X_c) = δ_0 is reformulated as a conditional moment restriction E{μ₁(X) − μ₀(X) − δ₀ | X_c} = 0 (Eq. 2) via the Law of Iterated Expectations — a standard identity, not circular. (2) The conditional-to-unconditional moment conversion (Eq. 3–4) cites Bierens (2016, Theorem 2.2), an external result. (3) The doubly robust score construction follows Chernozhukov et al. (2018), another external citation. (4) The test statistic S_n is a standard ICM-type statistic built from cross-fitted estimated scores; no parameter is fitted to the target quantity and then 'predicted.' (5) The weighted bootstrap recomputes only low-dimensional quantities (δ̂, π̂, λ̂) under reweighting while keeping nuisance estimates fixed — this is a genuine resampling procedure, not a construction that defines the bootstrap distribution to equal the null distribution by fiat. (6) The DiD and LATE extensions follow the same pattern with adapted scores; the LATE bootstrap's delta-method replication concern (raised in the skeptic headline) is a correctness/verifiability issue (proofs in unavailable supplementary material), not a circularity issue — the bootstrap is not defined in terms of the asymptotic distribution it aims to approximate. (7) No load-bearing self-citations were identified: the key mathematical ingredients (Bierens' ICM theorem, DML orthogonality, convergence rate results from Farrell et al. 2021 and Scornet et al. 2015) are all external. The score of 1 rather than 0 reflects only the minor concern that proofs are unavailable for independent verification, but this is a verifiability limitation, not circularity.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new particles, forces, dimensions, or postulated entities. All objects (propensity scores, outcome regressions, CATEs, LATEs) are standard in the causal inference literature. The test statistic and bootstrap procedure are constructed from existing mathematical objects (doubly robust scores, ICM principles, kernel matrices). The free parameters (bandwidth, folds, bootstrap iterations) are implementation choices, not theoretical constructs.

free parameters (3)
  • Kernel bandwidth h = 1
    The Gaussian kernel bandwidth h=1 is used in all simulations and empirical applications (Sections 5, 6, Appendix B). It is fixed a priori, not fitted to data, but the choice is not justified or sensitivity-tested.
  • Number of cross-fitting folds N = 5
    Five folds are used in all simulations and applications. Standard choice but not theoretically derived.
  • Number of bootstrap iterations B = 499
    Fixed across all simulations. Standard choice.
assumptions (5)
  • standard math Bierens (2016, Theorem 2.2): the conditional moment restriction E{g(X)|X_c}=0 is equivalent to E{g(X)phi(t'X_c)}=0 for all t in T, when phi is analytic non-polynomial with non-vanishing derivatives at 0.
    Invoked in Section 3.1 (Equation 3) and Assumption 3.3(iv). This is the load-bearing equivalence that converts the conditional null into a testable continuum of unconditional moments.
  • domain assumption Neyman orthogonality of the doubly robust score V with respect to nuisance functions mu_0, mu_1, m_D.
    Invoked in Section 3.1 and 3.2. Ensures that first-stage estimation errors vanish at first order. Standard in the double ML literature (Chernozhukov et al. 2018).
  • domain assumption Conditional unconfoundedness: (Y(1), Y(0)) _||_ D | X (Assumption 3.1).
    Standard identifying assumption for ATE under unconfoundedness. Extended to conditional parallel trends (Assumption 4.1) for DiD and instrument exogeneity (Assumption 4.9) for LATE.
  • domain assumption Nuisance estimators converge at rate o_P(n^{-1/4}) in L2(P) (Assumption 3.4(i)).
    Required for the asymptotic distribution in Proposition 3.1. The paper cites Farrell et al. (2021) and Scornet et al. (2015) for specific ML methods but does not verify this for the empirical applications.
  • standard math iid sampling (Assumption 3.3(i), 4.4(i), 4.7(i), 4.13(i)).
    Standard regularity condition invoked throughout. May be restrictive in panel or clustered settings.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Machine-Learning-Compatible Omnibus Test for Treatment Effect Heterogeneity." pith.science (2026). https://pith.science/paper/VLHHMSDG

@misc{pith2026260706412,
  author       = {Pith},
  title        = {Pith review of: A Machine-Learning-Compatible Omnibus Test for Treatment Effect Heterogeneity},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VLHHMSDG}},
  note         = {Machine review of arXiv:2607.06412}
}
read the original abstract

This study proposes a formal, computationally efficient nonparametric omnibus test for treatment-effect heterogeneity that is compatible with a broad class of estimators, including modern machine-learning methods. The test is designed for settings in which identification can rely on high-dimensional controls while heterogeneity is assessed with respect to a low-dimensional subset of covariates. We derive the test statistic's asymptotic null distribution and develop a bootstrap procedure that is efficient because it avoids re-estimating nuisance parameters in each iteration. The testing approach applies to multiple empirical designs, including randomized experiments, selection-on-observables, difference-in-differences, and instrumental-variables settings. Monte Carlo simulations show that the test attains near-nominal size under the null and exhibits good power against heterogeneous alternatives. We further illustrate the procedure using two empirical applications on retirement savings and trade liberalization.

Figures

Figures reproduced from arXiv: 2607.06412 by the authors.

Figure 3
Figure 3. Heterogeneity Test: Trade Liberalization and Corruption [PITH_FULL_IMAGE:figures/full_fig_p033_3.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

40 extracted references · 40 canonical work pages

  1. [1]

    Abadie, A. (2005). Semiparametric difference-in-differences estimators. The review of economic studies , 72(1):1--19

  2. [2]

    and Imbens, G

    Athey, S. and Imbens, G. (2016). Recursive partitioning for heterogeneous causal effects. Proceedings of the National Academy of Sciences , 113(27):7353--7360

  3. [3]

    and Imbens, G

    Athey, S. and Imbens, G. W. (2019). Machine learning methods that economists should know about. Annual Review of Economics , 11(1):685--725

  4. [4]

    Athey, S., Tibshirani, J., and Wager, S. (2019). Generalized random forests. Annals of Statistics , 47(2):1148--1178

  5. [5]

    Bierens, H. J. (2016). Econometric Model Specification . World Scientific

  6. [6]

    Chang, N.-C. (2020). Double/debiased machine learning for difference-in-differences models. The Econometrics Journal , 23(2):177--191

  7. [7]

    Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., and Robins, J. (2018). Double/debiased machine learning for treatment and structural parameters

  8. [8]

    Chernozhukov, V., Demirer, M., Duflo, E., and Fernández-Val, I. (2025). Fisher-Schultz Lecture: Generic machine learning inference on heterogenous treatment effects in randomized experiments, with an application to immunization in India . Econometrica , 93(4):1121--1164

Show all 40 references
  1. [9]

    Chernozhukov, V., Hansen, C., Kallus, N., Spindler, M., and Syrgkanis, V. (2024). Applied causal inference powered by ML and AI . arXiv:2403.02467

  2. [10]

    K., and Singh, R

    Chernozhukov, V., Newey, W. K., and Singh, R. (2022). De-biased machine learning of global and local parameters using regularized riesz representers. The Econometrics Journal , 25(3):576--601

  3. [11]

    K., and Singh, R

    Chernozhukov, V., Newey, W. K., and Singh, R. (2023). A simple and general debiased machine learning theorem with finite-sample guarantees. Biometrika , 110(1):257--264

  4. [12]

    K., Hotz, V

    Crump, R. K., Hotz, V. J., Imbens, G. W., and Mitnik, O. A. (2008). Nonparametric tests for treatment effect heterogeneity. The Review of Economics and Statistics , 90(3):389--405

  5. [13]

    Ding, P., Feller, A., and Miratrix, L. (2016). Randomization inference for treatment effect variation. Journal of the Royal Statistical Society Series B: Statistical Methodology , 78(3):655--671

  6. [14]

    Ding, P., Feller, A., and Miratrix, L. (2019). Decomposing treatment effect variation. Journal of the American Statistical Association , 114(525):304--317

  7. [15]

    Escanciano, J. C. (2026). Kernel-based specification testing with high-dimensional nuisance parameters. Working paper, presented at the ISNPS Conference, June 2026

  8. [16]

    P., and Zhang, Y

    Fan, Q., Hsu, Y.-C., Lieli, R. P., and Zhang, Y. (2022). Estimation of conditional average treatment effects with high-dimensional data. Journal of Business & Economic Statistics , 40(1):313--327

  9. [17]

    H., Liang, T., and Misra, S

    Farrell, M. H., Liang, T., and Misra, S. (2021). Deep neural networks for estimation and inference. Econometrica , 89(1):181--213

  10. [18]

    Foster, D. J. and Syrgkanis, V. (2023). Orthogonal statistical learning. Annals of Statistics , 51(3):879--908

  11. [19]

    Friedberg, R., Tibshirani, J., Athey, S., and Wager, S. (2021). Local linear forests. Journal of Computational and Graphical Statistics , 30(2):503--517

  12. [20]

    Green, D. P. and Kern, H. L. (2012). Modeling heterogeneous treatment effects in survey experiments with bayesian additive regression trees. Public Opinion Quarterly , 76(3):491--511

  13. [21]

    R., Murray, J

    Hahn, P. R., Murray, J. S., and Carvalho, C. (2020). Bayesian regression tree models for causal inference: Regularization, confounding, and heterogeneous effects. Bayesian Analysis , 15(3):965--1056

  14. [22]

    Hansen, C., Kozbur, D., and Misra, S. (2023). Targeted undersmoothing: Sensitivity analysis for sparse estimators. Review of Economics and Statistics , 105(1):101--118

  15. [23]

    Heckman, J., Smith, J., and Clements, N. (1997). Making the Most Out of Programme Evaluations and Social Experiments: Accounting for Heterogeneity in Programme Impacts . Review of Economic Studies , 64(4):487--535

  16. [24]

    Fisher–Schultz Lecture: Generic machine learning inference on heterogeneous treatment effects in randomized experiments, with an application to immunization in India

    Imai, K. and Li, M. L. (2025a). A comment on: “ Fisher–Schultz Lecture: Generic machine learning inference on heterogeneous treatment effects in randomized experiments, with an application to immunization in India” by Victor Chernozhukov, Mert Demirer, Esther Duflo, and Iván F...

  17. [25]

    and Li, M

    Imai, K. and Li, M. L. (2025b). Statistical inference for heterogeneous treatment effects discovered by generic machine learning in randomized experiments. Journal of Business & Economic Statistics , 43(1):256--268

  18. [26]

    Kennedy, E. H. (2023). Towards optimal doubly robust estimation of heterogeneous causal effects. Electronic Journal of Statistics , 17(2):3008--3049

  19. [27]

    and Mareckova, J

    Lechner, M. and Mareckova, J. (2022). Modified causal forest. arXiv:2209.03744

  20. [28]

    and Zimmert, M

    Lechner, M. and Zimmert, M. (2019). Nonparametric estimation of causal heterogeneity under high-dimensional confounding. arXiv: 1908.08779

  21. [29]

    A., Muir, I., and Sun, G

    List, J. A., Muir, I., and Sun, G. (2025). Using machine learning for efficient flexible regression adjustment in economic experiments. Econometric Reviews , 44(1):2--40

  22. [30]

    Nekipelov, D., Semenova, V., and Syrgkanis, V. (2022). Regularised orthogonal machine learning for nonlinear semiparametric models. The Econometrics Journal , 25(1):233--255

  23. [31]

    and Wager, S

    Nie, X. and Wager, S. (2020). Quasi-oracle estimation of heterogeneous treatment effects. Biometrika , 108(2):299--–319

  24. [32]

    M., Venti, S

    Poterba, J. M., Venti, S. F., and Wise, D. A. (1995). Do 401(k) contributions crowd out other personal saving? Journal of Public Economics , 58(1):1--32

  25. [33]

    Scheidegger, C., Guo, Z., and B \"u hlmann, P. (2026). Inference for heterogeneous treatment effects with efficient instruments and machine learning. Electronic Journal of Statistics , 20(1):718--770

  26. [34]

    Scornet, E., Biau, G., and Vert, J.-P. (2015). Consistency of random forests. Annals of Statistics , 43(4)

  27. [35]

    and Chernozhukov, V

    Semenova, V. and Chernozhukov, V. (2021). Debiased machine learning of conditional average treatment effects and other causal functions. The Econometrics Journal , 24(2):264--289

  28. [36]

    Semenova, V., Goldman, M., Chernozhukov, V., and Taddy, M. (2023). Inference on heterogeneous treatment effects in high-dimensional dynamic panels under weak dependence. Quantitative Economics , 14(2):471--510

  29. [37]

    Sequeira, S. (2016). Corruption, Trade Costs , and Gains from Tariff Liberalization : Evidence from Southern Africa . American Economic Review , 106(10):3029--3063

  30. [38]

    Tabord-Meehan, M. (2023). Stratification trees for adaptive randomisation in randomised controlled trials. Review of Economic Studies , 90(5):2646--2673

  31. [39]

    Taddy, M., Gardner, M., Chen, L., and Draper, D. (2016). A nonparametric bayesian analysis of heterogeneous treatment effects in digital experimentation. Journal of Business & Economic Statistics , 34(4):661--672

  32. [40]

    and Athey, S

    Wager, S. and Athey, S. (2018). Estimation and inference of heterogeneous treatment effects using random forests. Journal of the American Statistical Association , 113(523):1228--1242

Pith tools

Reviewed July 8, 2026 · model on record in the stance chip above.