Pith. sign in

REVIEW 3 major objections 4 minor 135 references

Debiased Machine Learning for Unobserved Heterogeneity: High-Dimensional Panels and Measurement Error Models

T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper proves that the debiased moments for smooth functionals of unobserved heterogeneity are exactly the solutions of two functional equations, with relevance pinned down by a nonzero constant, and turns this into valid cross-fitted…

desk verdict The constructive machinery is real, but the 'full characterization' in Theorem 1 is not necessary as stated: orthogonality only pins down the conditional moment up to an arbitrary function of X. read the letter →

arxiv 2507.13788 v1 pith:SLBSLOUH submitted 2025-07-18 econ.EM

classification econ.EM MSC 62G0562G2062P20
keywords Neymanorthogonalitydebiasedmachinelearningunobservedheterogeneityhigh-dimensionalpaneldatameasurementerrorvalue-addedmodelsfunctionaldifferencingpartialidentification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks when a researcher can draw valid conclusions about unobserved heterogeneity (UH)—the permanent, person-specific traits that drive behavior in panel data—without estimating the distribution of those traits. The paper's central claim is a complete characterization: a debiased (Neyman-orthogonal) moment condition for a smooth functional of UH exists and carries information exactly when a candidate moment function solves a conditional-expectation equation and an adjoint-score equation. If the claim is right, applied researchers get a constructive checklist for building machine-learning-based tests and confidence intervals for common parameters, average marginal effects, variances, and measurement-error factor moments. The same characterization also proves that some popular policy parameters, such as threshold-based teacher value-added measures, admit no debiased moment, so the usual plug-in inference for them cannot be locally robust.

What carries the argument

The load-bearing object is Theorem 1, the characterization of relevant Neyman-orthogonal moments. In the model above, let $S_\theta\delta=E[\ell_\theta\delta|Z]$ and $S_\eta b=E[b(\alpha,X)|Z]$ be the score operators for the common parameters and for the UH distribution, and let $S^*_\theta$ and $S^*_\eta g=E[g|\alpha,X]$ be their adjoints. The theorem says an orthogonal moment exists exactly when the pair $(r_\theta, r-\psi_0)$ lies in the range of the joint adjoint operator, up to scale; the scale $c\neq 0$ is what makes it relevant. The paper converts this into a three-step algorithm: solve for $g_0$ with $E[g_0|\alpha,X]=r(\alpha,X,\theta_0)$, solve for $g_1$ with $E[g_1|\alpha,X]=0$, then choose $\Gamma_0$ so that $S^*_\theta g_0-r_\theta=\Gamma_0 S^*_\theta g_1$; the final moment is $g_0-\psi_0-\Gamma_0 g_1$. This machinery produces every application in the paper and the cross-fitted DML tests.

What would settle it

Take the teacher value-added model $Y=\alpha+\varepsilon$ with $\varepsilon\sim N(0,\theta_0^2)$ and the policy functional $\psi_0=-E[\alpha\,1\{\alpha\le F_\alpha^{-1}(\phi)\}]$, and try to solve the integral equation $\int g_0(y,\theta_0)\phi_{\theta_0}(y-\alpha)dy=(F_\alpha^{-1}(\phi)-\alpha)1\{\alpha\le F_\alpha^{-1}(\phi)\}$ for a square-integrable $g_0$. The paper predicts no solution exists because the right-hand side is not an analytic function of $\alpha$; exhibiting such a solution would disprove the characterization.

Watch

Extended reading notes

Core claim

The paper proves Theorem 1: in a model with density $f_{\lambda_0}(z)=\int f_{Y|\alpha,X}(y|\alpha,x;\theta_0)\eta_0(\alpha|x)d\alpha$ and target $\psi(\lambda_0)=E[r(\alpha,X,\theta_0)]$, a non-zero zero-mean moment function $g$ is Neyman-orthogonal if and only if $E[g(Z,\lambda_0)|\alpha,X]=c(r(\alpha,X,\theta_0)-\psi_0)$ almost surely and $S^*_\theta g=c\,r_\theta$, where $c$ is a constant; the moment is relevant (informative about the target) if and only if $c\neq 0$. The necessity part is the hard step, and the conditions are constructive: they reduce debiasing to solving two functional equations rather than to a particular estimation strategy. Under support conditions, solutions to the first equation are globally robust to the distribution of UH, so the debiased moment does not require estimating that distribution at all.

Load-bearing premise

The practical claims hold only if the target functional is pathwise differentiable and every nuisance estimator used in the debiased moment converges at the $n^{-1/4}$ mean-square rate; if a lasso or Moore-Penrose step falls short of that rate, the proposed tests can be mis-sized.

Editorial extensions

If this is right

  • Functional differencing for common parameters in conditionally parametric panel models becomes a special case of the first equation, and the second equation supplies the extra orthogonality that functional differencing moments lack for average marginal effects and variances.
  • Cross-fitted debiased tests based on the constructed moments have correct asymptotic size and non-trivial local power when nuisance estimators converge at the $n^{-1/4}$ mean-square rate, covering high-dimensional random-coefficient panels estimated by lasso and truncated Moore-Penrose inverses.
  • In the Kotlarski model with a factor loading, debiased inference for moments $E[\alpha^k]$ is possible through a recursive closed-form moment, avoiding the slow logarithmic rates of nonparametric deconvolution estimators.
  • For teacher value-added models, smooth analytic functionals admit debiased moments built from Hermite polynomials, while CDFs, quantiles, and replacement-policy parameters do not admit any relevant orthogonal moment.
  • The empirical application finds that existing estimates of the average and variance effects of maternal smoking on birth weight are robust to flexible high-dimensional controls, with slightly more estimated variability across mothers.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: because the characterization is necessary and sufficient for the whole mixture class, other latent-variable models of the same form (auctions, duration models, production functions) can be screened for debiased inference by checking the same two equations, so the paper's three applications are a sample rather than the boundary.
  • Editorial inference: the nonexistence result for CDFs, quantiles, and replacement-policy functionals suggests that first-order orthogonality is too demanding for threshold-type targets; second-order orthogonality or distribution-robust bounds are the natural next step for those parameters.
  • Editorial inference: the first equation is a linear inverse problem, so the standard toolkit for ill-posed integral equations could be imported to supply rate-optimal numerical solvers where no closed-form $g_0$ is available.
  • Editorial inference: a direct testable prediction is that the efficiency gains reported in the Kotlarski Monte Carlo should also appear for analytic teacher value-added functionals when Hermite-series moments are truncated and compared with plug-in estimates.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper develops a debiased machine learning framework for inference on functionals of nonparametric unobserved heterogeneity in mixture models of the form f_λ(z)=∫ f_{Y|α,X}(y|α,x;θ_0) η_0(α|x)dα. The main theoretical claim is a complete characterization of all relevant Neyman-orthogonal moments: a nonzero moment function g is a relevant orthogonal moment for a smooth functional ψ(λ_0)=E[r(α,X,θ_0)] if and only if it satisfies the two functional equations E[g(Z,λ_0)|α,X]=c(r(α,X,θ_0)-ψ_0) and S*_θ g=c r_θ (Theorem 1). The paper proposes a three-step construction algorithm, develops DML inference theory with partial identification, and applies the method to high-dimensional random coefficient panel models, the Kotlarski model with a factor loading, and teacher value-added models, with Monte Carlo simulations and an empirical application to the effect of maternal smoking on birth weight.

Significance. If Theorem 1 were correct, the paper would provide a powerful unifying characterization of debiased moments in models with nonparametric unobserved heterogeneity, connecting functional differencing with modern DML theory and enabling constructive two-step estimation. The sufficiency direction, the specific moments derived for the three applications, the careful adaptation of the asymptotic framework of Chernozhukov et al. (2022), and the empirical robustness analysis are genuine contributions. However, the necessity direction of Theorem 1 is incorrect, so the paper's central claim—a full characterization of all relevant orthogonal moments—and the nonexistence results for CDFs, quantiles, and policy parameters are not established. The DML procedures themselves are not affected, as they verify orthogonality directly, but the paper's main intellectual contribution is invalid as stated.

major comments (3)
  1. [Section 4.4, Eq. (4.9), and Appendix A (proof of Theorem 1)] The necessity direction of Theorem 1 is false. The proof applies a Lagrange-multiplier argument on the space H_θ × H_η, but the η-score directions are restricted to B(η_0) with ∫ b(α,x) η_0(α|x) dα = 0 a.s., whose closure is the subspace of L2(η_0 × f_X) of functions with zero conditional mean given X. The orthogonal complement of this subspace is L2(f_X), the space of functions of X alone. Consequently, the necessity argument only establishes E[g|α,X] = c r(α,X,θ_0) + d(X) for some d ∈ L2(f_X), not the pointwise equality (4.9). A concrete counterexample is the model Y_t = α + ε_t for t=1,2, with X ~ Bernoulli(1/2) independent of (α,ε), and ψ_0 = E[α]; the moment g = 0.5(Y_1+Y_2) - ψ_0 + (X - 0.5) satisfies E[g]=0, is orthogonal on every path with ψ(λ_τ)=ψ_0, and is relevant for ψ, yet E[g|α,X] = α - ψ_0 + X - 0.5, which is not proportional to α - ψ_0. Theorem 1 as stated is therefore incorrect.
  2. [Section 7.3.2, Proposition 11] The proof that an orthogonal moment exists only if the Riesz representer r is analytic relies on the necessity of (4.9) through equation (7.12), E[g_0(Y,θ_0)|α] = r(α,θ_0). With the correction identified above, the relevant equation becomes E[g_0|α] = c r(α,θ_0) + d(X). When the model includes covariates X, the additive function d(X) can change the existence conclusions, so the paper's policy-relevant nonexistence claims—for CDFs, quantiles, and teacher-replacement policy parameters—are unproven in general models with covariates. Even if the conclusion might survive in the no-covariate teacher value-added example, the given proof does not establish the stated result.
  3. [Section 5 and Corollary 2] The three-step algorithm and Corollary 2 rely on solving the conditional moment equation (4.13), E[m_0(Z,λ_0)|α,X] = r(α,X,θ_0). Because the true orthogonality condition only pins down the conditional expectation up to an arbitrary function of X, the construction in Step 1 may fail to have a solution even when a relevant orthogonal moment exists; conversely, any solution of (4.13) can be modified by adding a function of X while preserving orthogonality. Thus the algorithm and the associated characterization in Remark 3 do not cover 'all' relevant orthogonal moments, and the claimed necessity of (4.13) is not correct.
minor comments (4)
  1. [Section 6, hypothesis statement] The statement 'H0 : ψ(λ0) = ψ0 vs H0 : ψ(λ0) ≠ ψ0' should use H1 for the alternative hypothesis.
  2. [Section 3.6 and Remark 3] The heuristic summary and Remark 3 state that the set of orthogonal moments is characterized 'up to scale'; given the issue identified in Theorem 1, this should read 'up to scale and addition of a function of X'.
  3. [References] The reference 'Hanusheck' should be 'Hanushek' for the entry 'Teacher Deselection'.
  4. [Section 7.1.3, Assumption 6] Assumption 6 restates 'HV = Iq' after redefining H as the Kronecker product operator H ⊗ H in the same section; the notation should be disambiguated to avoid confusion.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central characterization is a theorem with a self-contained proof, and the constructed moments are verified rather than assumed.

full rationale

The paper's main claim, Theorem 1, is a necessary-and-sufficient characterization of relevant Neyman-orthogonal moments, and its proof in Appendix A derives the two equations (4.9)-(4.10) from the definition of orthogonality and the score-operator representation rather than importing them as assumptions. The moments constructed in Sections 7.1-7.3 are checked directly against those equations, which is a verification of the theorem's conditions, not a circular reduction. The Monte Carlo study compares the proposed debiased moment against plug-in and functional-differencing benchmarks, and the empirical application validates robustness of existing estimates; neither involves fitting a target outcome and then relabeling the fit as a prediction. The paper does cite the authors' earlier arXiv preprint (Argañaraz and Escanciano, 2023), but the cited statements are contextual or point to extensions, and the proof of Theorem 1 does not depend on that prior work. Any concern that the necessity direction of Theorem 1 may be false because orthogonality only pins down the conditional moment up to a function of X would be a mathematical-correctness objection, not a circularity of the paper's own derivation chain. The asymptotic inference results require nuisance estimators with n^{-1/4} rates, but that is a stated regularity condition, not an input smuggled in as a prediction.

Assumptions & free parameters 3 free parameters · 7 assumptions · 0 invented entities

The central theorem rests on standard semiparametric regularity plus domain-specific restrictions (smooth functionals, support conditions, identification assumptions in each application). No new physical entities are introduced; the 'relevant orthogonal moment' is a theoretical construct defined by the paper.

free parameters (3)
  • lasso penalty parameter cn and loadings = c=1.1, γ=0.1/log(p∨(n−nℓ)T)
    Used in Algorithm 1 for estimating β in high-dimensional panel; affects the n^{-1/4} rate in Lemma 20.
  • eigenvalue truncation thresholds νn and πn = (log(T^2)/n)^{1/2} in empirical application
    Used to estimate Moore-Penrose inverses of W and M; choice affects asymptotic variance estimator.
  • Monte Carlo tuning (ks=10, nz=1000, nα=100, L=4) = ks=10, nz=1000, nα=100, L=4
    Specific values used in the spectral cut-off and discretization in Section G.1; reproduce the simulation results but are not derived from theory.
assumptions (7)
  • domain assumption Regularity of the model (Assumption 1 and Assumption 11): DQM, linear tangent space, bounded score operator
    Needed for the score operators Sθ and Sη and their adjoints, which define the orthogonality conditions in Theorem 1.
  • domain assumption Smoothness of the functional (Assumption 3): pathwise derivative of θ↦E[r(α,X,θ)] exists with representer rθ
    Required for the Riesz representer rψ and the necessity part of Theorem 1; excludes CDFs and quantiles.
  • standard math Regularity of moments (Assumption 2): interchange of derivative and expectation for g under paths
    Standard differentiability condition for moment functions used throughout the asymptotic analysis.
  • domain assumption Support condition (Remark 14 and Eq. 5.5): a known set A contains the support of the UH distribution
    Needed for global robustness of the orthogonal moments to the distribution of UH; without it moments may depend on η0 through its support.
  • domain assumption High-dimensional panel model: exact sparsity of β0, restricted eigenvalue conditions on M, and exogeneity E[ε|α,X]=0
    Used in Lemma 20 to establish n^{-1/4} rates for the lasso estimator in Section E.1.
  • domain assumption Kotlarski model: α independent of ε, components of ε independent with zero mean, finite moments
    The recursive construction of LR moments for moments of α requires these identifying conditions.
  • domain assumption Teacher value-added: errors Gaussian with common variance, functional analytic (entire) in α, exponential moment condition E[exp(cY^2)]<∞
    The Hermite polynomial expansion of g0 and the finiteness of variance in Proposition 10 rely on these.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Debiased Machine Learning for Unobserved Heterogeneity: High-Dimensional Panels and Measurement Error Models." pith.science (2026). https://pith.science/paper/SLBSLOUH

@misc{pith2026250713788,
  author       = {Pith},
  title        = {Pith review of: Debiased Machine Learning for Unobserved Heterogeneity: High-Dimensional Panels and Measurement Error Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SLBSLOUH}},
  note         = {Machine review of arXiv:2507.13788}
}
read the original abstract

Developing robust inference for models with nonparametric Unobserved Heterogeneity (UH) is both important and challenging. We propose novel Debiased Machine Learning (DML) procedures for valid inference on functionals of UH, allowing for partial identification of multivariate target and high-dimensional nuisance parameters. Our main contribution is a full characterization of all relevant Neyman-orthogonal moments in models with nonparametric UH, where relevance means informativeness about the parameter of interest. Under additional support conditions, orthogonal moments are globally robust to the distribution of the UH. They may still involve other high-dimensional nuisance parameters, but their local robustness reduces regularization bias and enables valid DML inference. We apply these results to: (i) common parameters, average marginal effects, and variances of UH in panel data models with high-dimensional controls; (ii) moments of the common factor in the Kotlarski model with a factor loading; and (iii) smooth functionals of teacher value-added. Monte Carlo simulations show substantial efficiency gains from using efficient orthogonal moments relative to ad-hoc choices. We illustrate the practical value of our approach by showing that existing estimates of the average and variance effects of maternal smoking on child birth weight are robust.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

135 extracted references · 76 canonical work pages

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in " " * FUNCTION format....

  2. [2]

    Aaronson, Daniel, Lisa Barrow, and William Sander (2007): Teachers and student achievement in the Chicago public high schools , Journal of labor Economics, 25 (1), 95--135

  3. [3]

    Abowd, John M, Francis Kramarz, and David N Margolis (1999): High wage workers and high wage firms , Econometrica, 67 (2), 251--333

  4. [4]

    Abrevaya, Jason (2006): Estimating the effect of smoking on birth outcomes using a matched panel data approach , Journal of Applied Econometrics, 21 (4), 489--519

  5. [5]

    Aguirregabiria, Victor and Jes \'u s M Carro (2024): Identification of average marginal effects in fixed effects dynamic discrete choice models , Review of Economics and Statistics, 1--46

  6. [6]

    rep., National Bureau of Economic Research

    Alvarez, Fernando, Katar \' na Borovi c kov \'a , and Robert Shimer (2016): Decomposing duration dependence in a stopping time model , Tech. rep., National Bureau of Economic Research

  7. [7]

    Andersen, Erling Bernhard (1970): Asymptotic properties of conditional maximum-likelihood estimators , Journal of the Royal Statistical Society Series B: Statistical Methodology, 32 (2), 283--301

  8. [8]

    Anderson, Theodore W and Herman Rubin (1949): Estimation of the parameters of a single equation in a complete system of stochastic equations , The Annals of mathematical statistics, 20 (1), 46--63

Show all 135 references
  1. [9]

    Andrews, Donald WK (1987): Asymptotic results for generalized Wald tests , Econometric Theory, 3 (3), 348--358

  2. [10]

    Angrist, Joshua, Peter Hull, Parag Pathak, and Christopher Walters (2016): Interpreting tests of school VAM validity , American Economic Review, 106 (5), 388--392

  3. [11]

    Angrist, Joshua D, Peter D Hull, Parag A Pathak, and Christopher R Walters (2017): Leveraging lotteries for school value-added: Testing and estimation , The Quarterly Journal of Economics, 132 (2), 871--919

  4. [12]

    Arellano, Manuel, Richard Blundell, and St \'e phane Bonhomme (2017): Earnings and consumption dynamics: a nonlinear panel data framework , Econometrica, 85 (3), 693--734

  5. [13]

    Arellano, Manuel and St \'e phane Bonhomme (2012): Identifying distributional characteristics in random coefficients panel data models , The Review of Economic Studies, 79 (3), 987--1020

  6. [14]

    Arellano, Manuel and Jinyong Hahn (2007): Understanding bias in nonlinear panel models: Some recent developments , Econometric Society Monographs, 43, 381

  7. [15]

    Arga \ n araz, Facundo and Juan Carlos Escanciano (2023): On the Existence and Information of Orthogonal Moments For Inference , arXiv preprint arXiv:2303.11418

  8. [16]

    Argañaraz, Facundo (2024): Automatic Debiased Machine Learning of Structural Parameters with General Conditional Moments, Working paper

  9. [17]

    Begun, Janet M, W Jackson Hall, Wei-Min Huang, and Jon A Wellner (1983): Information and asymptotic efficiency in parametric-nonparametric models , The Annals of Statistics, 11 (2), 432--452

  10. [18]

    Belloni, Alexandre, Daniel Chen, Victor Chernozhukov, and Christian Hansen (2012): Sparse models and methods for optimal instruments with an application to eminent domain , Econometrica, 80 (6), 2369--2429

  11. [19]

    Belloni, Alexandre, Victor Chernozhukov, Christian Hansen, and Damian Kozbur (2016): Inference in high-dimensional panel models with an application to gun control , Journal of Business & Economic Statistics, 34 (4), 590--605

  12. [20]

    Bennett, Andrew, Nathan Kallus, Xiaojie Mao, Whitney Newey, Vasilis Syrgkanis, and Masatoshi Uehara (2023): Source condition double robust inference on functionals of inverse problems , arXiv preprint arXiv:2307.13793

  13. [21]

    Bickel, Peter J (1982): On adaptive estimation , The Annals of Statistics, 647--671

  14. [22]

    (2011): Panel Data, Inverse Problems, and the Estimation of Policy Parameters , Unpublished manuscript

    Bonhomme, S. (2011): Panel Data, Inverse Problems, and the Estimation of Policy Parameters , Unpublished manuscript

  15. [23]

    Bonhomme, St \'e phane (2012): Functional Differencing, Econometrica, 80 (4), 1337--1385

  16. [24]

    Bonhomme, St \'e phane and Kevin Dano (2024): Functional Differencing in Networks, Revue \'e conomique , 75 (1), 147--175

  17. [25]

    Bonhomme, St \'e phane and Angela Denis (2024): Estimating heterogeneous effects: applications to labor economics , arXiv preprint arXiv:2404.01495

  18. [26]

    Bonhomme, St \'e phane, Koen Jochmans, and Martin Weidner (2024): A neyman-orthogonalization approach to the incidental parameter problem , arXiv preprint arXiv:2412.10304

  19. [27]

    Bonhomme, St \'e phane and Jean-Marc Robin (2010): Generalized non-parametric deconvolution with an application to earnings dynamics , The Review of Economic Studies, 77 (2), 491--533

  20. [28]

    Borusyak, Kirill, Xavier Jaravel, and Jann Spiess (2024): Revisiting event-study designs: robust and efficient estimation , Review of Economic Studies, 91 (6), 3253--3285

  21. [29]

    Browning, Martin and Jesus Carro (2007): Heterogeneity and Microeconometrics Modeling, Cambridge University Press, 47–74, Econometric Society Monographs

  22. [30]

    rep., National Bureau of Economic Research

    Bunting, Jackson, Paul Diegert, and Arnaud Maurel (2024): Heterogeneity, Uncertainty and Learning: Semiparametric Identification and Estimation, Tech. rep., National Bureau of Economic Research

  23. [31]

    6 of Handbook of Econometrics, Elsevier

    Carrasco, Marine, Jean-Pierre Florens, and Eric Renault (2007): Chapter 77 Linear Inverse Problems in Structural Econometrics Estimation Based on Spectral Decomposition and Regularization , vol. 6 of Handbook of Econometrics, Elsevier

  24. [32]

    Carroll, Raymond J and Peter Hall (1988): Optimal rates of convergence for deconvolving a density , Journal of the American Statistical Association, 83 (404), 1184--1186

  25. [33]

    Chamberlain, Gary (1984): Panel Data, Handbook of econometrics, 2, 1247--1318

  26. [34]

    --- -.1pt --- -.1pt --- (1992): Efficiency bounds for semiparametric regression , Econometrica: Journal of the Econometric Society, 567--596

  27. [35]

    Chen, Xiaohong and Andres Santos (2018): Overidentification in regular models , Econometrica, 86 (5), 1771--1817

  28. [36]

    Chernozhukov, Victor, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins (2018): Double/debiased machine learning for treatment and structural parameters , The Econometrics Journal, 21, C1--C68

  29. [37]

    Chernozhukov, Victor, Juan Carlos Escanciano, Hidehiko Ichimura, Whitney K Newey, and James M Robins (2022): Locally robust semiparametric estimation , Econometrica, 90 (4), 1501--1535

  30. [38]

    Chernozhukov, Victor, Iv \'a n Fern \'a ndez-Val, Jinyong Hahn, and Whitney Newey (2013): Average and quantile effects in nonseparable panel models , Econometrica, 81 (2), 535--580

  31. [39]

    Chetty, Raj, John N Friedman, and Jonah E Rockoff (2014): Measuring the impacts of teachers II: Teacher value-added and student outcomes in adulthood , American economic review, 104 (9), 2633--2679

  32. [40]

    Chetty, Raj and Nathaniel Hendren (2018): The impacts of neighborhoods on intergenerational mobility II: County-level estimates , The Quarterly Journal of Economics, 133 (3), 1163--1228

  33. [41]

    Clarke, Paul S and Annalivia Polselli (2025): Double machine learning for static panel models with fixed effects , The Econometrics Journal, utaf011

  34. [42]

    Cooper, Russell W and John C Haltiwanger (2006): On the nature of capital adjustment costs , The Review of Economic Studies, 73 (3), 611--633

  35. [43]

    Corman, Hope (1995): The effects of low birthweight and other medical risk factors on resource utilization in the pre-school years ,

  36. [44]

    Corman, Hope and Stephen Chaikind (1998): The effect of low birthweight on the school performance and behavior of school-aged children , Economics of Education Review, 17 (3), 307--316

  37. [45]

    Cunha, Flavio, James J Heckman, and Susanne M Schennach (2010): Estimating the technology of cognitive and noncognitive skill formation , Econometrica, 78 (3), 883--931

  38. [46]

    Currie, Janet and Rosemary Hyson (1999): Is the impact of health shocks cushioned by socioeconomic status? The case of low birthweight , American economic review, 89 (2), 245--250

  39. [47]

    Daubechies, Ingrid, Michel Defrise, and Christine De Mol (2004): An iterative thresholding algorithm for linear inverse problems with a sparsity constraint , Communications on Pure and Applied Mathematics: A Journal Issued by the Courant Institute of Mathematical Sciences, 57 ...

  40. [48]

    Davezies, Laurent, Xavier D'Haultfoeuille, and Louise Laage (2021): Identification and estimation of average marginal effects in fixed effects logit models , arXiv preprint arXiv:2105.00879

  41. [49]

    Davis, Tom P (2024): A general expression for Hermite expansions with applications , The Mathematics Enthusiast, 21 (1), 71--87

  42. [50]

    Dobbelaere, Sabien and Jacques Mairesse (2013): Panel data estimates of the production function and product and labor market imperfections , Journal of Applied Econometrics, 28 (1), 1--46

  43. [51]

    Dobronyi, Christopher, Jiaying Gu, and Kyoo il Kim (2021): Identification of Dynamic Panel Logit Models with Fixed Effects ,

  44. [52]

    Russell (2024): Identification of Dynamic Panel Logit Models with Fixed Effects,

    Dobronyi, Christopher, Jiaying Gu, Kyoo il Kim, and Thomas M. Russell (2024): Identification of Dynamic Panel Logit Models with Fixed Effects,

  45. [53]

    Doyle, Joseph, John Graves, and Jonathan Gruber (2019): Evaluating measures of hospital quality: Evidence from ambulance referral patterns , Review of Economics and Statistics, 101 (5), 841--852

  46. [54]

    Duffy, John, Chris Papageorgiou, and Fidel Perez-Sebastian (2004): Capital-skill complementarity? Evidence from a panel of countries , Review of Economics and Statistics, 86 (1), 327--344

  47. [55]

    D’Haultfoeuille, Xavier (2011): On the completeness condition in nonparametric instrumental problems , Econometric Theory, 27 (3), 460--471

  48. [56]

    Efron, Bradley (2016): Empirical Bayes deconvolution estimates , Biometrika, 103 (1), 1--20

  49. [57]

    375 of Mathematics and Its Applications, Dordrecht: Kluwer Academic Publishers

    Engl, Heinz Werner, Martin Hanke, and Andreas Neubauer (1996): Regularization of Inverse Problems, vol. 375 of Mathematics and Its Applications, Dordrecht: Kluwer Academic Publishers

  50. [58]

    Escanciano, Juan Carlos (2022): Semiparametric identification and fisher information , Econometric Theory, 38 (2), 301--338

  51. [59]

    Escanciano, Juan Carlos and Wei Li (2021): Optimal linear instrumental variables approximations , Journal of Econometrics, 221 (1), 223--246

  52. [60]

    Evdokimov, Kirill (2010): Identification and estimation of a nonparametric panel data model with unobserved heterogeneity , Department of Economics, Princeton University, 1 (5), 23

  53. [61]

    Evdokimov, Kirill and Halbert White (2012): Some extensions of a lemma of Kotlarski , Econometric Theory, 28 (4), 925--932

  54. [62]

    Fan, Jianqing (1991): On the optimal rates of convergence for nonparametric deconvolution problems , The Annals of Statistics, 1257--1272

  55. [63]

    Ferreira, JC and Valdir Ant \^o nio Menegatto (2013): Positive Definiteness, Reproducing Kernel Hilbert Spaces and Beyond, Annals of Functional Analysis, 4 (1)

  56. [64]

    Finkelstein, Amy, Matthew Gentzkow, Peter Hull, and Heidi Williams (2017): Adjusting risk adjustment--accounting for variation in diagnostic intensity , The New England journal of medicine, 376 (7), 608

  57. [65]

    rep., National Bureau of Economic Research

    Fletcher, Jason M, Leora I Horwitz, and Elizabeth Bradley (2014): Estimating the value added of attending physicians on patient outcomes , Tech. rep., National Bureau of Economic Research

  58. [66]

    Fox, Jeremy T, Kyoo il Kim, and Chenyu Yang (2016): A simple nonparametric approach to estimating the distribution of random coefficients in structural models , Journal of Econometrics, 195 (2), 236--254

  59. [67]

    Freyberger, Joachim (2018): Non-parametric panel data models with interactive fixed effects , The Review of Economic Studies, 85 (3), 1824--1851

  60. [68]

    Friedman, Jerome, Trevor Hastie, Holger H \"o fling, and Robert Tibshirani (2007): Pathwise coordinate optimization , The annals of applied statistics, 1 (2), 302--332

  61. [69]

    Friedman, Jerome, Trevor Hastie, and Rob Tibshirani (2010): Regularization paths for generalized linear models via coordinate descent , Journal of statistical software, 33 (1), 1

  62. [70]

    Fruehwirth, Jane Cooley, Salvador Navarro, and Yuya Takahashi (2016): How the timing of grade retention affects outcomes: Identification and estimation of time-varying treatment effects , Journal of Labor Economics, 34 (4), 979--1021

  63. [71]

    Fu, Wenjiang J (1998): Penalized regressions: the bridge versus the lasso , Journal of computational and graphical statistics, 7 (3), 397--416

  64. [72]

    rep., National Bureau of Economic Research

    Gilraine, Michael, Jiaying Gu, and Robert McMillan (2020): A new method for estimating teacher value-added , Tech. rep., National Bureau of Economic Research

  65. [73]

    Goodman-Bacon, Andrew (2021): Difference-in-differences with variation in treatment timing , Journal of Econometrics, 225 (2), 254--277

  66. [74]

    Irregular

    Graham, Bryan S and James L Powell (2012): Identification and Estimation of Average Partial Effects in “Irregular” Correlated Random Coefficient Panel Data Models, Econometrica, 80 (5), 2105--2152

  67. [75]

    Guvenen, Fatih (2009): An empirical investigation of labor income processes , Review of Economic dynamics, 12 (1), 58--79

  68. [76]

    Hack, Maureen, Nancy K Klein, and H Gerry Taylor (1995): Long-term developmental outcomes of low birth weight infants , The future of children, 176--196

  69. [77]

    Haider, Steven and Gary Solon (2006): Life-cycle variation in the association between current and lifetime earnings , American Economic Review, 96 (4), 1308--1320

  70. [78]

    (2009): Teacher Deselection, Urban Institute Press, 47–74

    Hanusheck, Eric A. (2009): Teacher Deselection, Urban Institute Press, 47–74

  71. [79]

    Hanushek, Eric A (2011): The economic value of higher teacher quality , Economics of Education review, 30 (3), 466--479

  72. [80]

    Heckman, James and Burton Singer (1984 a ): A method for minimizing the impact of distributional assumptions in econometric models for duration data , Econometrica: Journal of the Econometric Society, 271--320

  73. [81]

    --- -.1pt --- -.1pt --- (1984 b ): The identifiability of the proportional hazard model , The Review of Economic Studies, 51 (2), 231--241

  74. [82]

    Heckman, James J (2001): Micro data, heterogeneity, and the evaluation of public policy: Nobel lecture , Journal of political Economy, 109 (4), 673--748

  75. [83]

    Honor \'e , Bo E (1992): Trimmed LAD and least squares estimation of truncated and censored regression models with fixed effects , Econometrica: Journal of the Econometric Society, 533--565

  76. [84]

    Honor \'e , Bo E and Elie Tamer (2006): Bounds on parameters in panel dynamic discrete choice models , Econometrica, 74 (3), 611--629

  77. [85]

    Honor \'e , Bo E and Martin Weidner (2024): Moment conditions for dynamic panel logit models with fixed effects , Review of Economic Studies

  78. [86]

    and Martin Weidner (2023): Moment Conditions for Dynamic Panel Logit Models with Fixed Effects ,

    Honoré, Bo E. and Martin Weidner (2023): Moment Conditions for Dynamic Panel Logit Models with Fixed Effects ,

  79. [87]

    Horowitz, Joel L (1999): Semiparametric estimation of a proportional hazard model with unobserved heterogeneity , Econometrica, 67 (5), 1001--1028

  80. [88]

    Horowitz, Joel L and Marianthi Markatou (1996): Semiparametric estimation of regression models for panel data , The Review of Economic Studies, 63 (1), 145--168

  81. [89]

    Hu, Yingyao and Susanne M Schennach (2008): Instrumental variable treatment of nonclassical measurement error models , Econometrica, 76 (1), 195--216

  82. [90]

    Ibragimov, I.A. and R.Z. Khasminskii (1981): Statistical Estimation: Asymptotic Theory , Springer-Verlag, New York

  83. [91]

    Ishwaran, Hemant (1999): Information in semiparametric mixtures of exponential families , The Annals of Statistics, 27 (1), 159--177

  84. [92]

    rep., National Bureau of Economic Research

    Kane, Thomas J and Douglas O Staiger (2008): Estimating teacher impacts on student achievement: An experimental evaluation , Tech. rep., National Bureau of Economic Research

  85. [93]

    Kato, Kengo, Yuya Sasaki, and Takuya Ura (2021): Robust inference in deconvolution , Quantitative Economics, 12 (1), 109--142

  86. [94]

    Kiefer, Jack and Jacob Wolfowitz (1956): Consistency of the maximum likelihood estimator in the presence of infinitely many incidental parameters , The Annals of Mathematical Statistics, 887--906

  87. [95]

    Klaassen, Chris AJ (1987): Consistent estimation of the influence function of locally asymptotically linear estimators , The Annals of Statistics, 15 (4), 1548--1562

  88. [96]

    Kline, Patrick, Evan K Rose, and Christopher R Walters (2022): Systemic discrimination among large US employers , The Quarterly Journal of Economics, 137 (4), 1963--2036

  89. [97]

    Klosin, Sylvia and Max Vilgalys (2022): Estimating continuous treatment effects in panel data using machine learning with an agricultural application , arXiv preprint arXiv:2207.08789

  90. [98]

    Koedel, Cory and Jonah E Rockoff (2015): Value-added modeling: A review , Economics of Education Review, 47, 180--195

  91. [99]

    Koenker, Roger and Ivan Mizera (2014): Convex optimization, shape constraints, compound decisions, and empirical Bayes rules , Journal of the American Statistical Association, 109 (506), 674--685

  92. [100]

    and Harold R

    Krantz, Steven G. and Harold R. Parks (2002): A Primer of Real Analytic Functions , Springer Science & Business Media, New York

  93. [101]

    Krasnokutskaya, Elena (2011): Identification and estimation of auction models with unobserved heterogeneity , The Review of Economic Studies, 78 (1), 293--327

  94. [102]

    Kress, Rainer (1989): Linear Integral Equations , Springer Berlin, Heidelberg

  95. [103]

    Lee, Adam (2024): Locally Regular and Efficient Tests in Non-Regular Semiparametric Models , arXiv preprint arXiv:2403.05999

  96. [104]

    Lee, Adam and Geert Mesters (2024): Locally robust inference for non-Gaussian linear simultaneous equations models , Journal of Econometrics, 240 (1), 105647

  97. [105]

    and Joseph P

    Lehmann, Erich L. and Joseph P. Romano (2005): Testing Statistical Hypotheses, Springer Texts in Statistics, Springer, 3rd ed

  98. [106]

    Lewbel, Arthur (2022): Kotlarski with a factor loading , Journal of Econometrics, 229 (1), 176--179

  99. [107]

    Lewbel, Arthur and Krishna Pendakur (2017): Unobserved preference heterogeneity in demand using generalized random coefficients , Journal of Political Economy, 125 (4), 1100--1148

  100. [108]

    Lewbel, Arthur, Susanne M Schennach, and Linqi Zhang (2024): Identification of a triangular two equation system without instruments , Journal of Business & Economic Statistics, 42 (1), 14--25

  101. [109]

    Li, Tong and Quang Vuong (1998): Nonparametric estimation of the measurement error model using multiple indicators , Journal of Multivariate Analysis, 65 (2), 139--165

  102. [110]

    Lillard, Lee A and Yoram Weiss (1979): Components of variation in panel earnings data: American scientists 1960-70 , Econometrica: Journal of the Econometric Society, 437--454

  103. [111]

    Luenberger, David G (1997): Optimization by vector space methods , John Wiley & Sons

  104. [112]

    Magnus, J. R. and H. Neudecker (2019): Matrix Differential Calculus with Applications in Statistics and Econometrics , John Wiley & Sons

  105. [113]

    Mairesse, Jacques and Zvi Griliches (1990): Heterogeneity in Panel Data: Are There Stable Production Functions? in Essays in Honor of Edmond Malinvaud, Vol. 3, ed. by Paul Champsaur, Michel Deleau, Jean-Michel Grandmont, Guy Laroque, Roger Guesnerie, Claude Henry, Jean-Jacques...

  106. [114]

    Masten, Matthew A (2018): Random coefficients on endogenous variables in simultaneous equations models , The Review of Economic Studies, 85 (2), 1193--1250

  107. [115]

    Mundlak, Yair (1978): On the Pooling of Time Series and Cross Section Data, Econometrica: journal of the Econometric Society, 69--85

  108. [116]

    Narasimhan, Balasubramanian and Bradley Efron (2020): deconvolveR: A G-Modeling Program for Deconvolution and Empirical Bayes Estimation, Journal of Statistical Software, 94 (11), 1–20

  109. [117]

    Navarro, Salvador and Jin Zhou (2017): Identifying agent's information sets: An application to a lifecycle model of schooling, consumption and labor supply , Review of Economic Dynamics, 25, 58--92

  110. [118]

    Navjeevan, Manu, Rodrigo Pinto, and Andres Santos (2023): Identification and estimation in a class of potential outcomes models , arXiv preprint arXiv:2310.05311

  111. [119]

    Newey, Whitney K (1990): Semiparametric efficiency bounds , Journal of applied econometrics, 5 (2), 99--135

  112. [120]

    (1959): Optimal Asymptotic Tests of Composite Statistical Hypothesis , Probability and Statistics: The Harald Cramer Volume, 213--234

    Neyman, J. (1959): Optimal Asymptotic Tests of Composite Statistical Hypothesis , Probability and Statistics: The Harald Cramer Volume, 213--234

  113. [121]

    Nybom, Martin and Jan Stuhler (2016): Heterogeneous income profiles and lifecycle bias in intergenerational mobility estimation , Journal of Human Resources, 51 (1), 239--268

  114. [122]

    Rafi, Ahnaf (2024): Nonparametric inference for a class of functionals in the random coefficients logit model , Working paper

  115. [123]

    Rao, BLS Prakasa (1983): Nonparametric functional estimation , Academic Press

  116. [124]

    1, 601--620

    Rao, C Radhakrishna and Sujit Kumar Mitra (1972): Generalized inverse of a matrix and its applications , in Proceedings of the Berkeley Symposium on Mathematical Statistics and Probability, University of California Press, Berkeley, vol. 1, 601--620

  117. [125]

    Santos, Andres (2011): Instrumental variable methods for recovering continuous linear functionals , Journal of Econometrics, 161 (2), 129--146

  118. [126]

    Schick, Anton (1986): On asymptotically efficient estimation in semiparametric models , The Annals of Statistics, 1139--1151

  119. [127]

    Semenova, Vira, Matt Goldman, Victor Chernozhukov, and Matt Taddy (2023): Inference on heterogeneous treatment effects in high-dimensional dynamic panels under weak dependence , Quantitative Economics, 14 (2), 471--510

  120. [128]

    Sz \'e kely, GJ and CR Rao (2000): Identifiability of distributions of independent random variables by linear combinations and moments , Sankhy \=a : The Indian Journal of Statistics, Series A , 193--202

  121. [129]

    Tseng, Paul (2001): Convergence of a block coordinate descent method for nondifferentiable minimization , Journal of optimization theory and applications, 109, 475--494

  122. [130]

    5, 3381--3460

    Van den Berg, Gerard J (2001): Duration models: specification, identification and multiple durations , in Handbook of Econometrics, Elsevier, vol. 5, 3381--3460

  123. [131]

    CWI tract 44, Centrum voor Wiskunde en Informatics, Amsterdam

    Van der Vaart, AW (1988): Statistical estimation in large parameter spaces , Thesis. CWI tract 44, Centrum voor Wiskunde en Informatics, Amsterdam

  124. [132]

    Van der Vaart, Aad (1991): On Differentiable Functionals, The Annals of Statistics, 178--204

  125. [133]

    Van der Vaart, A. W. (1998): Asymptotic Statistics , Cambridge University Press, New York

  126. [134]

    (2019): High-Dimensional Statistics: A Non-Asymptotic Viewpoint , Cambridge University Press

    Wainwright, Martin J. (2019): High-Dimensional Statistics: A Non-Asymptotic Viewpoint , Cambridge University Press

  127. [135]

    Yakusheva, Olga, Richard Lindrooth, and Marianne Weiss (2014): Nurse value-added and patient outcomes in acute care , Health services research, 49 (6), 1767--1786

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.