REVIEW 3 major objections 3 minor 48 references
Nonparametric "rich covariates" without saturation
T0 review · 3 major / 3 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Two nonparametric fixes restore a causal interpretation for linear IV with covariates.
desk verdict Worth a serious referee, but the variance formulas in Theorems 2 and 4 need correction before the inference claims are usable. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the conditional expectation of the instrument, $\zeta_0(c)=E[z\mid c]$. The theory estimates it with a higher-order Nadaraya-Watson kernel; the simulations use a shallow ReLU neural network. The two constructions then force rich covariates exactly: $z-\hat\zeta(c)$ is mean-independent of every function of $c$, so its conditional expectation is the constant zero, and adding $\hat\zeta(c)$ as a regressor makes the instrument's conditional expectation linear in the regressors by definition. The asymptotic mechanism is undersmoothing—choosing the bandwidth so that $nh^{2d}/(\ln n)^2\to\infty$ and $nh^{2m}\to 0$, which requires $m>d$—so that first-step bias disappears faster than $\sqrt{n}$ while variance remains controlled, in the standard two-step semiparametric framework.
What would settle it
Run Monte Carlo replications of the kernel-first-step estimator under the paper's assumptions, form 95% confidence intervals using the variance estimator of Theorem 2, and check empirical coverage at a sample size like $n=8000$. If coverage is systematically below 95%, the $\sqrt{n}$-consistency claim fails; if the same experiment with a neural-network first step on a high-dimensional, non-smooth $\zeta_0$ shows bias that does not shrink at the $\sqrt{n}$ rate, the practical recommendation for neural networks is also in doubt.
Extended reading notes
Core claim
The paper's central claim is that two feasible two-step linear IV estimators—one that instruments the endogenous binary treatment with $z-\hat\zeta(c)$, and one that adds $\hat\zeta(c)$ to the regressor list—satisfy the rich-covariates condition asymptotically without parametric assumptions on $E[z\mid c]$. Theorems 1 and 3 show that, under the stated kernel, bandwidth, smoothness, integrability, and identification assumptions, both estimators are $\sqrt{n}$-consistent and asymptotically normal, with the variance estimators of Theorems 2 and 4 consistent. The probability limit of the treatment coefficient is $\alpha_{\mathrm{rich}} = E[\omega(c)\,EC(y(1)-y(0)\mid c)]$, with overlap weights $\omega(c)=\mathrm{Cov}(z,t\mid c)/E[\mathrm{Cov}(z,t\mid c)]$, so the estimand is a positively weighted average of conditional complier effects and is weakly causal under monotonicity. The formal results cover only the kernel first step; the neural-network implementation is supported by simulation evidence, with its asymptotic theory explicitly left for future work.
Load-bearing premise
The $\sqrt{n}$-consistency theorems require an undersmoothed higher-order kernel bandwidth and a true conditional expectation $\zeta_0$ that is $m$-times continuously differentiable with $m$ exceeding the number of covariates $d$; the paper's recommended neural-network first step is not covered by those theorems, so its guarantees in that implementation rest on simulations.
Editorial extensions
If this is right
- Applied researchers can enforce rich covariates without saturation by running one nonparametric first-step regression and then using either $z-\hat\zeta(c)$ as the instrument or $\hat\zeta(c)$ as an extra regressor.
- Under the kernel conditions, the resulting treatment-coefficient estimates are $\sqrt{n}$-consistent and asymptotically normal, so the usual t-statistics and confidence intervals are available.
- Both estimators identify the same positively weighted complier average, $\alpha_{\mathrm{rich}}$, so the causal interpretation does not depend on which construction is used.
- In the simulations the instrument-residual version is less sensitive to the number of controls and to the first-step estimator than the control-function version, and both improve on a parametric instrument-residual alternative when the parametric model is misspecified.
- In the empirical application the proposed estimators reverse the sign of the LIVE estimate and line up with the saturated-model estimate, suggesting that enforcing rich covariates can materially change conclusions.
Reading between the lines
- Extending the asymptotic theory to the neural-network first step would turn the paper's practical recommendation into a theorem; that is the natural next step and would close the gap between what is proved and what is recommended.
- Because the control-function estimator's asymptotic variance depends on the coefficient on $\hat\zeta(c)$, the choice of other regressors changes efficiency through an additional channel; this may explain the instrument-residual estimator's better performance in the simulations.
- The sign reversal in the empirical application implies that published linear-IV results with many sparse binary controls may be sensitive to the rich-covariates assumption; re-examining such applications with these estimators is a direct way to test how much the assumption matters.
- The framework should extend to multi-valued or continuous instruments and to overidentified models using existing estimand results, which would widen the applicability beyond binary treatment and binary instrument.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes two two-step linear IV estimators that enforce the rich-covariates condition without a saturated model or unconditional randomization: the instrument-residual estimator, which uses z − ζ̂(c) as the instrument, and the control-function estimator, which adds ζ̂(c) as a regressor. It shows that both estimators target αrich = E[ω(c)EC(y(1)−y(0)|c)], proves √n-consistency and asymptotic normality for a kernel first step under Assumptions 1–6, and reports simulations with kernel and neural-network first steps and an application to Dube and Harish (2020). The point-estimator proofs are plausibly derived from Newey and McFadden (1994), but the variance estimators in Theorems 2 and 4 are not valid as written.
Significance. The paper addresses an important practical gap: it offers alternatives to saturated specifications for ensuring rich covariates, with an estimand that has a clear causal interpretation under strong monotonicity, and it connects to existing Lee/Kim-Lee and DML estimators. The simulation design is thoughtful, and the honest discussion of the neural-network first step (explicitly not covered by the theorems) is a strength. The proof strategy via Newey and McFadden (1994) is appropriate and the identification argument appears sound. The blocking issue is inferential: the proposed variance estimators contain a sign error and an infeasible or inconsistent plug-in for the adjustment term, so the paper's claim of a consistent variance estimator is not yet supported.
major comments (3)
- [Section 4.2, Theorem 2] The displayed estimator is Ω̂ := −(1/n)∑_{i=1}^n τ̂_i τ̂_i′, with a leading minus sign. This makes Ω̂ negative semidefinite, so it cannot be a consistent estimator of the covariance matrix Ω, which is positive semidefinite. The same minus sign appears in the control-function variance estimator in Section 4.3. If this is a typographical error, it must still be corrected before the formula is usable.
- [Section 4.2–4.3, Theorems 2 and 4] The plug-in for the first-step adjustment term is inconsistent with the influence function. Theorem 1 defines φ=(−z*(ζ0)E[ε|c],0_k')′, so the first component of q(ζ0)ε+φ is z*(ζ0)(ε−E[ε|c]). The estimator in Theorem 2 sets φ̂_i=(z*_i(ζ̂(c_i))ε̂_i,0_k')′, which uses the regression residual ε̂_i in place of E[ε_i|c_i]. Since E[ε|c] is not assumed to be zero (the whole setting allows misspecification), τ̂_i has first component 2z*_iε̂_i rather than z*_i(ε̂_i−E[ε_i|c_i]); this changes the variance being estimated. No estimator of E[ε|c] is supplied, and the proof of Theorem 2 verifies Lipschitz conditions for the infeasible quantity −E[ε(β)|c]z*(ζ), not for the estimator actually implemented. Theorem 4 has the same problem, with ε̂_i^* used in place of E[ε*|c]. The consistency claim for the variance estimator is therefore not established.
- [Section 6, footnote 26] The empirical illustration reports cluster-robust IV standard errors that 'do not account for the fact that the estimator relies on a preliminary first step.' This means the application does not use the Theorem 2/Theorem 4 variance estimators, so the paper provides no practical demonstration of the proposed inference procedure. The simulations in Section 5 also do not report coverage rates or rejection rates. Unless a feasible first-step-robust variance estimator is supplied and implemented, the paper's asymptotic-inference claim remains non-operational.
minor comments (3)
- [Section 5] The bandwidth formula for the Nadaraya-Watson first step is chosen by experimentation, and the simulation conclusions for the kernel version may depend on this calibration; a sentence on sensitivity would help.
- [Abstract and Section 7] The abstract and conclusions recommend neural networks for the first step, while Section 3.3 and Section 7 correctly state that the theorems cover only the kernel first step; this scope caveat should be made more prominent in the abstract so that readers do not take the neural-network version as theorem-backed.
- [Table 2 and surrounding text] The row labels are informative, but the text should clarify that the 'DML neural net no cross-fitting' estimate is a comparison method rather than one of the proposed estimators, and the footnote on standard errors should state explicitly that none of the reported standard errors correspond to the variance formulas in Theorems 2 and 4.
Circularity Check
No significant circularity: the rich-covariates condition is enforced by construction, the estimand is benchmarked externally, and the kernel asymptotics rely on standard Newey–McFadden theorems with no self-citation chain.
full rationale
The paper's two estimators are constructed so that the rich-covariates condition holds algebraically at the population level: the instrument residual z − ζ0(c) is mean-independent of any function of c, and adding ζ0(c) as a regressor makes L[z | r, ζ0(c)] equal to ζ0(c). This is an explicit construction, not a fitted parameter later relabeled as a prediction. The target estimand αrich is taken from Lee (2021) and Blandhol et al. (2025), with the equivalence shown in Appendix A; it is not derived from the paper's own fitted values. In the application, the proposed estimators are compared against the externally computed saturated-model estimate (−0.509) reported by Blandhol et al. (2022), and no parameter is fit to that benchmark. Theorems 1 and 3 are proven by verifying the conditions of Newey and McFadden (1994), an independent reference theorem, with explicit derivative and bound calculations; no uniqueness theorem from the authors' own prior work is invoked. The reference list contains no self-citations that are load-bearing. The skeptical observation about the negative sign and plug-in error in the variance estimator in Theorems 2 and 4 is a substantive mathematical correctness concern, not a circularity concern: even if the variance estimator is inconsistent as written, the √n-consistency claim does not reduce to its own inputs. Similarly, footnote 26's use of standard errors that ignore the first step is an implementation disclosure, not a circular derivation. Finally, the paper explicitly limits its asymptotic theory to the kernel first step and states that extension to neural networks is left for future work, so no unproven claim is smuggled in through self-citation or ansatz. Overall, the derivation chain is self-contained against external benchmarks and contains no circular step.
Assumptions & free parameters
free parameters (2)
- Kernel bandwidth constants in simulations =
(1.1 + 0.725d) σ̂ n^{-1/(2d+1)}
- Neural network architecture =
Single hidden layer, 100 ReLU nodes, Adam optimizer, default pystacked options
assumptions (6)
- domain assumption Conditional instrument exogeneity: (y(0), y(1), t(0), t(1)) ⊥⊥ z | c
- domain assumption Strong monotonicity: sign of Cov(z,t|c) is the same for all c
- domain assumption Compact support of c with density bounded away from zero and infinity (Assumption 3)
- domain assumption ζ0 is m-times continuously differentiable with m > d (Assumption 4)
- domain assumption Higher-order kernel with vanishing moments up to m-1 and bandwidth conditions (Assumptions 1 and 2)
- standard math Moment existence and nonsingularity (Assumptions 5, 6, and 6*)
Cite this review
Pith. "Pith review of Nonparametric "rich covariates" without saturation." pith.science (2026). https://pith.science/paper/7GZP63DS
@misc{pith2026250521213,
author = {Pith},
title = {Pith review of: Nonparametric "rich covariates" without saturation},
year = {2026},
howpublished = {\url{https://pith.science/paper/7GZP63DS}},
note = {Machine review of arXiv:2505.21213}
}
read the original abstract
We consider two nonparametric approaches to ensure that linear instrumental variables estimators satisfy the rich-covariates condition emphasized by Blandhol et al. (2025), even when the instrument is not unconditionally randomly assigned and the model is not saturated. Both approaches start with a nonparametric estimate of the expectation of the instrument conditional on the covariates, and ensure that the rich-covariates condition is satisfied either by using as the instrument the difference between the original instrument and its estimated conditional expectation, or by adding the estimated conditional expectation to the set of regressors. We derive asymptotic properties when the first step uses kernel regression, and assess finite-sample performance in simulations where we also use neural networks in the first step. Finally, we present an empirical illustration that highlights some significant advantages of the proposed methods.
Reference graph
Works this paper leans on
-
[1]
Abadie, A. (2003): Semiparametric Instrumental Variable Estimation of Treatment Response Models, Journal of Econometrics, 113, 231--263
work page 2003
-
[2]
Ahrens, A., C. B. Hansen, and M. E. Schaffer (2023): pystacked: Stacking Generalization and Machine Learning in Stata , The Stata Journal, 23, 909--931
work page 2023
-
[3]
Ahrens, A., C. B. Hansen, M. E. Schaffer, and T. Wiemann (2024): ddml: Double/Debiased Machine Learning in Stata , The Stata Journal, 24, 3--45
work page 2024
-
[4]
Angrist, J. D., K. Graddy, and G. W. Imbens (2000): The Interpretation of Instrumental Variables Estimators in Simultaneous Equations Models with an Application to the Demand for Fish, The Review of Economic Studies, 67, 499--527
work page 2000
-
[5]
Angrist, J. D. and G. W. Imbens (1995): Two-Stage Least Squares Estimation of Average Causal Effects in Models with Variable Treatment Intensity, Journal of the American Statistical Association, 90, 431--442
work page 1995
-
[6]
Angrist, J. D. and J.-S. Pischke (2009): Mostly Harmless Econometrics: An Empiricist's Companion, Princeton (NJ): Princeton University Press
work page 2009
-
[7]
--- -.1pt --- -.1pt --- (2010): The Credibility Revolution in Empirical Economics: How Better Research Design is Taking the Con Out of Econometrics, Journal of Economic Perspectives, 24, 3--30
work page 2010
-
[8]
Bach, F. (2017): Breaking the Curse of Dimensionality with Convex Neural Networks, Journal of Machine Learning Research, 18, 1--53
work page 2017
Show all 48 references
-
[9]
(1994): Approximation and Estimation Bounds for Artificial Neural Networks, Machine Learning, 14, 115--133
Barron, A. (1994): Approximation and Estimation Bounds for Artificial Neural Networks, Machine Learning, 14, 115--133
1994
-
[10]
Bauer, B. and M. Kohler (2019): On Deep Learning as a Remedy for the Curse of Dimensionality in Nonparametric Regression, Annals of Statistics, 47, 2261--2285
2019
-
[11]
Bonney, M
Blandhol, C., J. Bonney, M. Mogstad, and A. Torgovitsky (2022): When is TSLS Actually LATE ? University of Chicago, Becker Friedman Institute for Economics Working Paper No. 2022-16
2022
-
[12]
--- -.1pt --- -.1pt --- (2025): When is TSLS Actually LATE ? NBER Working Paper No. w29709
2025
-
[13]
Borusyak, K. and P. Hull (2021): Non-Random Exposure to Exogenous Shocks: Theory and Applications, mimeo
2021
-
[14]
--- -.1pt --- -.1pt --- (2023): Nonrandom Exposure to Exogenous Shocks, Econometrica, 91, 2155--2185
2023
-
[15]
Kohler, S
Braun, A., M. Kohler, S. Langer, and H. Walk (2024): Convergence Rates for Shallow Neural Networks Learned by Gradient Descent, Bernoulli, 30, 475--502
2024
-
[16]
Chen, X. and H. White (1999): Improved Rates and Asymptotic Normality for Non- parametric Neural Network Estimators, IEEE Transactions on Information Theory, 45, 682--691
1999
-
[17]
Chetverikov, M
Chernozhukov, V., D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W. Newey, and J. Robins (2018): Double/Debiased Machine Learning for Treatment and Structural Parameters, The Econometrics Journal, 21, C1--C68
2018
-
[18]
Dube, O. and S. Harish (2020): Queens, Journal of Political Economy, 128, 2579--2652
2020
-
[19]
Farrell, M. H., T. Liang, and S. Misra (2021): Deep Neural Networks for Estimation and Inference, Econometrica, 89, 181--213
2021
-
[20]
Liu, and J
Gao, Q., L. Liu, and J. S. Racine (2015): A Partially Linear Kernel Estimator for Categorical Data, Econometric Reviews, 34, 959--978
2015
-
[21]
Hansen, B. E. (2005): Exact Mean Integrated Squared Error of Higher Order Kernel Estimators, Econometric Theory, 21, 1031–1057
2005
-
[22]
Stinchcombe, and H
Hornik, K., M. Stinchcombe, and H. White (1989): Multilayer Feedforward Networks are Universal Aapproximators, Neural Networks, 2, 359--366
1989
-
[23]
Imbens, G. W. and J. D. Angrist (1994): Identification and Estimation of Local Average Treatment Effects, Econometrica, 62, 467--475
1994
-
[24]
and M.-J
Kim, B. and M.-J. Lee (2024): Instrument-Residual Estimator for Multi-valued Instruments Under Full Monotonicity, Statistics & Probability Letters, 213, 110187
2024
-
[25]
(1980): Existence of Moments of k-Class Estimators, Econometrica, 48, 241--249
Kinal, T. (1980): Existence of Moments of k-Class Estimators, Econometrica, 48, 241--249
1980
-
[26]
Kingma, D. P. and J. Ba (2015): Adam: A Method for Stochastic Optimization, arXiv:1412.6980v9
2015 arXiv
-
[27]
Kohler, M. and S. Langer (2021): On the Rate of Convergence of Fully Connected Deep Neural Network Regression Estimates, The Annals of Statistics, 49, 2231--2249
2021
-
[28]
(2013): Estimation in an Instrumental Variables Model With Treatment Effect Heterogeneity, Working Paper 2013-2, Princeton University
Kolesár, M. (2013): Estimation in an Instrumental Variables Model With Treatment Effect Heterogeneity, Working Paper 2013-2, Princeton University
2013
-
[29]
and M.-J
Lee, G. and M.-J. Lee (2025): Double-Debiasing Instrumental Variable Estimator with Machine-Learned Plug-In for Any Outcome and Heterogeneity, mimeo
2025
-
[30]
Lee, M.-J. (2021): Instrument Residual Estimator for any Response Variable with Endogenous Binary Treatment, Journal of the Royal Statistical Society Series B: Statistical Methodology, 83, 612--635
2021
-
[31]
Lee, M.-J. and C. Han (2024): Ordinary Least Squares and Instrumental-Variables Estimators for any Outcome and Heterogeneity, The Stata Journal, 24, 72--92
2024
-
[32]
Mogstad, M. and A. Torgovitsky (2024): Instrumental Variables with Unobserved Heterogeneity in Treatment Effects, in Handbook of Labor Economics, ed. by C. Dustmann and T. Lemieux, Amsterdam: Elsevier, vol. 5, 1--114
2024
-
[33]
Torgovitsky, and C
Mogstad, M., A. Torgovitsky, and C. Walters (2021): The Causal Interpretation of Two-Stage Least Squares with Multiple Instrumental Variables, American Economic Review, 111, 3663--3698
2021
-
[34]
Murphy, K. M. and R. H. Topel (1985): Estimation and Inference in Two-Step Econometric Models, Journal of Business & Economic Statistics, 3, 370--379
1985
-
[35]
Nadaraya, E. A. (1964): On Estimating Regression, Theory of Probability and Its Applications, 9, 141--142
1964
-
[36]
Newey, W. K. (1994): The Asymptotic Variance of Semiparametric Estimators, Econometrica, 62, 1349--1382
1994
-
[37]
--- -.1pt --- -.1pt --- (1997): Convergence Rates and Asymptotic Normality for Series Estimators, Journal of Econometrics, 79, 147--168
1997
-
[38]
Newey, W. K., F. Hsieh, and J. M. Robins (2004): Twicing Kernels and a Small Bias Property of Semiparametric Estimators, Econometrica, 72, 947--962
2004
-
[39]
Newey, W. K. and D. L. McFadden (1994): Large Sample Estimation and Hypothesis Testing, in Handbook of Econometrics, ed. by R. F. Engle and D. L. McFadden, Amsterdam: Elsevier, vol. 4, 2111--2245
1994
-
[40]
(1984): Econometric Issues in the Analysis of Regressions with Generated Regressors, International Economic Review, 25, 221--247
Pagan, A. (1984): Econometric Issues in the Analysis of Regressions with Generated Regressors, International Economic Review, 25, 221--247
1984
-
[41]
Varoquaux, A
Pedregosa, F., G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay (2011): Scikit-learn: Machine Learning in P ython, Journal of Machine...
2011
-
[42]
Ramsey, J. B. (1969): Tests for Specification Errors in Classical Linear Least-Squares Regression Analysis, Journal of the Royal Statistical Society Series B: Statistical Methodology, 31, 350--371
1969
-
[43]
Robins, J. M., S. D. Mark, and W. K. Newey (1992): Estimating Exposure Effects by Modelling the Expectation of Exposure Conditional on Confounders, Biometrics, 48, 479--495
1992
-
[44]
(1988): Root-N-Consistent Semiparametric Regression, Econometrica, 56, 931--954
Robinson, P. (1988): Root-N-Consistent Semiparametric Regression, Econometrica, 56, 931--954
1988
-
[45]
(2020): Nonparametric Regression Using Deep Neural Networks with ReLU Activation Function, The Annals of Statistics, 48, 1875--1897
Schmidt-Hieber, J. (2020): Nonparametric Regression Using Deep Neural Networks with ReLU Activation Function, The Annals of Statistics, 48, 1875--1897
2020
-
[46]
Scott, D. W. (2015): Multivariate density estimation: theory, practice, and visualization, New York (NY): John Wiley & Sons
2015
-
[48]
--- -.1pt --- -.1pt --- (2024): When Should We (Not) Interpret Linear IV Estimands as LATE ? arXiv:2011.06695v7
2024 arXiv
-
[49]
(1964): Smooth Regression Analysis, Sankhy \=a : The Indian Journal of Statistics, Series A , 26, 359--372
Watson, G. (1964): Smooth Regression Analysis, Sankhy \=a : The Indian Journal of Statistics, Series A , 26, 359--372
1964
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.