Pith. sign in

REVIEW 34 references

Bias Correction and Robust Inference in Semiparametric Models

T0 review · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Semiparametric estimators can suffer from a nonlinear bias and an average nonparametric bias when the first-step estimator is imprecise, and this paper provides two correction methods that work under weaker rate conditions than the standard n^{1/4} requirement.

arxiv 1908.00414 v3 pith:QXB453F5 submitted 2019-08-01 math.ST stat.TH

classification math.STstat.TH
keywords semiparametricbiasingredientnonparametricbiasescorrectionestimatorinference
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Many econometric models estimate a parameter of interest in two steps: first estimate an unknown function from the data, then plug that estimate into a formula for the main parameter. Standard theory requires the first-step function to be estimated accurately, at least at a rate faster than n^{-1/4}, so that the two-step estimator behaves as if the function were known. This paper asks what happens when the first step is worse, for example because the bandwidth in a kernel density estimator is chosen in a way that leaves a noticeable bias.

The authors show that two distinct biases can then contaminate the main estimator. One comes from the variance part of the first-step estimate: because the plug-in formula is nonlinear, fluctuations in the function estimate do not average out. The other comes from the smoothed bias of the first-step estimate itself. A third, kernel-specific bias, the singularity bias, arises because the kernel density estimate at an observation point includes that same observation with an unusually large weight. The paper develops two fixes. The multi-scale jackknife re-estimates the whole procedure with several bandwidths and combines the results with weights that cancel the known bias orders. The analytical correction instead computes explicit estimates of the two biases and subtracts them. The asymptotic results require only a faster-than-n^{1/6} first-step rate, which is weaker than the usual n^{1/4} condition. Simulations for the average density, integrated squared density, and density-weighted average derivative estimators show that coverage rates of confidence intervals become flatter across bandwidths.

Extended reading notes

Core claim

Theorem 2 states that if Assumptions 1 to 3 hold, J_n - J_0 = O_P(pG_n(θ0, γhat_n)), s > 1/4, and r > 1/6, then sqrt(n)(θhat_n - θ0 - J_n BNL - J_n BANB) converges in distribution to N(0, Σ_θ). The load-bearing content is that a semiparametric estimator can be asymptotically normal with root-n inference after removing two non-negligible biases even when the nonparametric ingredient converges only faster than n^{1/6}, weaker than the standard faster-than-n^{1/4} requirement.

Load-bearing premise

Assumption 2 (Quadraticity), Section 2.2, requires that g(z_i, θ0, .) admits a stochastic second-order Taylor expansion in the nonparametric ingredient with a remainder controlled by E||g_R|| <= C E||γ - γ0||^3. This smoothness condition is stronger than the linearization used by Newey (1994), and if the criterion function is not twice differentiable in γ, or if the third-order remainder is not small at the required rate, the entire two-bias decomposition, and both correction procedures that build on it, fails.

Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Assumptions & free parameters 1 free parameters · 7 assumptions · 0 invented entities

The central claim rests on a combination of standard semiparametric assumptions and paper-specific high-level conditions. No parameters are fitted to data to produce the main theorems, but the user must choose the jackknife scales in practice, and the theoretical results depend on smoothness and bias-order assumptions that are only partially verified for general estimators.

free parameters (1)
  • Jackknife scale count Q and scale ratios eta_q = Q=2 with eta=(1,5/4) for 2SJ; Q=5 with eta=(3/5,4/5,1,6/5,7/5) for 5SJ
    The multi-scale jackknife requires the user to choose the number of scales and the relative bandwidths. The asymptotic theory works for any choice that keeps the weight equations nonsingular, but the simulation values are chosen by hand and the 5SJ choice for ISD is made after observing that 3SJ performed poorly.
assumptions (7)
  • domain assumption Assumption 1: asymptotic linearity of theta_hat in g with non-degenerate J0 and G(theta0, gamma0)=0.
    Standard in semiparametric two-step estimation; excludes weak identification and assumes an influence-function representation.
  • ad hoc to paper Assumption 2: stochastic quadratic expansion of g in the nonparametric ingredient with third-order remainder bound.
    This is the load-bearing smoothness condition that creates the nonlinear bias; it is not derived from primitive conditions for general g.
  • domain assumption Assumption 3: asymptotic normality of pG_n(theta0, gamma0) + pG1_n(theta0, gamma0, gamma_hat - gamma_bar).
    Justified in the text by U-statistic central limit theory; it holds under moment and degeneracy conditions but is stated as a high-level assumption.
  • ad hoc to paper Assumption 4: bias terms separate into deterministic BNL = O(n^{-2r}) and BANB = O(n^{-s}) with remainders o_P(n^{-1/2}).
    This high-level bias-order condition is verified for kernels and sieves in examples and in Lemma 7, but it is not verified for all semiparametric models.
  • ad hoc to paper Theorem 5 conditions (3.4) and (3.5): the centered U-statistic is asymptotically normal and the twicing-kernel residual bias is o_P(n^{-1/2}).
    These conditions are close to the desired conclusion and are not derived from primitive conditions; the twicing-kernel small bias property makes (3.5) plausible but not automatic.
  • domain assumption MSJ rate conditions: known bias orders h^m and 1/(n h^{d_z}), with dimension satisfying 3d_z < 8m.
    The method requires knowing how bias rates depend on the bandwidth, and it only works when a bandwidth sequence exists that meets both the smoothing-bias and variance-bias conditions.
  • standard math Standard U-statistic central limit theory and projection lemmas.
    Used to prove Assumption 3 and Lemma 7; treated as background results from Hoeffding and related references.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bias Correction and Robust Inference in Semiparametric Models." pith.science (2026). https://pith.science/paper/QXB453F5

@misc{pith2026190800414,
  author       = {Pith},
  title        = {Pith review of: Bias Correction and Robust Inference in Semiparametric Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QXB453F5}},
  note         = {Machine review of arXiv:1908.00414}
}
read the original abstract

This paper analyzes several different biases that emerge from the (possibly) low-precision nonparametric ingredient in a semiparametric model. We show that both the variance part and the bias part of the nonparametric ingredient can lead to some biases in the semiparametric estimator, under conditions weaker than typically required in the literature. We then propose two bias-robust inference procedures, based on multi-scale jackknife and analytical bias correction, respectively. We also extend our framework to the case where the semiparametric estimator is constructed by some discontinuous functionals of the nonparametric ingredient. Simulation study shows that both bias-correction methods have good finite-sample performance.

Figures

Figures reproduced from arXiv: 1908.00414 by the authors.

Figure 1
Figure 1. AD: Decomposition of Mean Squared Error [PITH_FULL_IMAGE:figures/full_fig_p028_1.png] view at source ↗
Figure 2
Figure 2. AD: Empirical Coverage Rates of Confidence Interva [PITH_FULL_IMAGE:figures/full_fig_p029_2.png] view at source ↗
Figure 3
Figure 3. ISD: Decomposition of Mean Squared Error [PITH_FULL_IMAGE:figures/full_fig_p030_3.png] view at source ↗
Figures from the paper (27 more)
Figure 4
Figure 4. Figure 4: ISD: Empirical Coverage Rates of Confidence Interv [PITH_FULL_IMAGE:figures/full_fig_p030_4.png]
Figure 5
Figure 5. Figure 5: DWAD: Decomposition of Mean Squared Error [PITH_FULL_IMAGE:figures/full_fig_p031_5.png]
Figure 6
Figure 6. Figure 6: DWAD: Empirical Coverage Rates of Confidence Inter [PITH_FULL_IMAGE:figures/full_fig_p031_6.png]
Figure 7
Figure 7. Figure 7: AD: Decomposition of Mean Squared Error 0.05 0.1 0.15 0.2 0.25 0.3 0.35 0.4 0.45 0.5 0.8 0.85 0.9 0.95 1 AD: Empirical Coverage Rates of Confidence Intervals (n=50) Nominal Raw ABC 2SJ [PITH_FULL_IMAGE:figures/full_fig_p047_7.png]
Figure 8
Figure 8. Figure 8: AD: Empirical Coverage Rates of Confidence Interva [PITH_FULL_IMAGE:figures/full_fig_p047_8.png]
Figure 9
Figure 9. Figure 9: AD: Densities of t-statistics and standard normal [PITH_FULL_IMAGE:figures/full_fig_p048_9.png]
Figure 10
Figure 10. Figure 10: AD: Decomposition of Mean Squared Error 0.05 0.1 0.15 0.2 0.25 0.3 0.35 0.4 0.45 0.5 0.8 0.85 0.9 0.95 1 AD: Empirical Coverage Rates of Confidence Intervals (n=200) Nominal Raw ABC 2SJ [PITH_FULL_IMAGE:figures/full_fig_p049_10.png]
Figure 11
Figure 11. Figure 11: AD: Empirical Coverage Rates of Confidence Interv [PITH_FULL_IMAGE:figures/full_fig_p049_11.png]
Figure 12
Figure 12. Figure 12: AD: Densities of t-statistics and standard norma [PITH_FULL_IMAGE:figures/full_fig_p050_12.png]
Figure 13
Figure 13. Figure 13: AD: Decomposition of Mean Squared Error 0.05 0.1 0.15 0.2 0.25 0.3 0.8 0.85 0.9 0.95 1 AD: Empirical Coverage Rates of Confidence Intervals (n=1000) Nominal Raw ABC 2SJ [PITH_FULL_IMAGE:figures/full_fig_p051_13.png]
Figure 14
Figure 14. Figure 14: AD: Empirical Coverage Rates of Confidence Interv [PITH_FULL_IMAGE:figures/full_fig_p051_14.png]
Figure 15
Figure 15. Figure 15: AD: Densities of t-statistics and standard norma [PITH_FULL_IMAGE:figures/full_fig_p052_15.png]
Figure 16
Figure 16. Figure 16: ISD: Decomposition of Mean Squared Error [PITH_FULL_IMAGE:figures/full_fig_p054_16.png]
Figure 17
Figure 17. Figure 17: ISD: Empirical Coverage Rates of Confidence Inter [PITH_FULL_IMAGE:figures/full_fig_p054_17.png]
Figure 18
Figure 18. Figure 18: ISD: Densities of t-statistics and standard norm [PITH_FULL_IMAGE:figures/full_fig_p055_18.png]
Figure 19
Figure 19. Figure 19: ISD: Decomposition of Mean Squared Error [PITH_FULL_IMAGE:figures/full_fig_p055_19.png]
Figure 20
Figure 20. Figure 20: ISD: Empirical Coverage Rates of Confidence Inter [PITH_FULL_IMAGE:figures/full_fig_p056_20.png]
Figure 21
Figure 21. Figure 21: ISD: Densities of t-statistics and standard norm [PITH_FULL_IMAGE:figures/full_fig_p057_21.png]
Figure 22
Figure 22. Figure 22: ISD: Decomposition of Mean Squared Error [PITH_FULL_IMAGE:figures/full_fig_p058_22.png]
Figure 23
Figure 23. Figure 23: ISD: Empirical Coverage Rates of Confidence Inter [PITH_FULL_IMAGE:figures/full_fig_p058_23.png]
Figure 24
Figure 24. Figure 24: ISD: Densities of t-statistics and standard norm [PITH_FULL_IMAGE:figures/full_fig_p059_24.png]
Figure 25
Figure 25. Figure 25: DWAD: Decomposition of Mean Squared Error [PITH_FULL_IMAGE:figures/full_fig_p060_25.png]
Figure 26
Figure 26. Figure 26: DWAD: Empirical Coverage Rates of Confidence Inte [PITH_FULL_IMAGE:figures/full_fig_p060_26.png]
Figure 27
Figure 27. Figure 27: DWAD: Decomposition of Mean Squared Error [PITH_FULL_IMAGE:figures/full_fig_p060_27.png]
Figure 28
Figure 28. Figure 28: DWAD: Empirical Coverage Rates of Confidence Inte [PITH_FULL_IMAGE:figures/full_fig_p061_28.png]
Figure 29
Figure 29. Figure 29: DWAD: Decomposition of Mean Squared Error [PITH_FULL_IMAGE:figures/full_fig_p061_29.png]
Figure 30
Figure 30. Figure 30: DWAD: Empirical Coverage Rates of Confidence Inte [PITH_FULL_IMAGE:figures/full_fig_p061_30.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

34 extracted references · 34 canonical work pages

  1. [1]

    Andrews, D. W. K. (1994): Asymptotics for Semiparametric Econometric Models via Stochastic Equicontinuity, Econometrica, 62, 43--72

  2. [2]

    Andrews, I. and A. Mikusheva (2016): Conditional Inference with Functional Nuisance Parameter, Econometrica, 84, 1571--1612

  3. [3]

    Bierens, H. J. (1987): Kernel Estimators of Regression Functions, in Advances in econometrics: Fifth World Congress, vol. 1, 99--144

  4. [4]

    Borovskikh, Y. V. (1996): U-statistics in B anach Spaces , V.S.P. Intl Science

  5. [5]

    Calonico, S., M. D. Cattaneo, and R. Titiunik (2014): Robust nonparametric confidence intervals for regression-discontinuity designs, Econometrica, 82, 2295--2326

  6. [6]

    Cattaneo, M. D., R. K. Crump, and M. Jansson (2010): Robust Data-Driven Inference for Density-Weighted Average Derivatives, Journal of the American Statistical Association, 105, 1070--1083

  7. [7]

    --- -.1pt --- -.1pt --- (2013): Generalized Jackknife Estimators of Weighted Average Derivatives, Journal of the American Statistical Association, 108, 1243--1256

  8. [8]

    --- -.1pt --- -.1pt --- (2014): Small Bandwidth Asymptotics for Density-Weightede Average Derivatives, Econometric Theory, 30, 176--200

Show all 34 references
  1. [9]

    Cattaneo, M. D. and M. Jansson (2018): Kernel-Based Semiparametric Estimators: Small Bandwidth Asymptotics and Bootstrap Consistency, Econometrica, 86, 955--995

  2. [10]

    (2007): Large sample sieve estimation of semi-nonparametric models, in Handbook of Econometrics, ed

    Chen, X. (2007): Large sample sieve estimation of semi-nonparametric models, in Handbook of Econometrics, ed. by J. J. Heckman and E. E. Leamer, Elsevier, vol. 6, 5549--5632

  3. [11]

    Linton, and I

    Chen, X., O. Linton, and I. van Keilegom (2003): Estimation of Semiparametric Models When The Criterion Function Is Not Smooth, Econometrica, 71, 1591--1608

  4. [12]

    Chetverikov, M

    Chernozhukov, V., D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, and W. Newey (2017): Double/Debiased/ N eyman Machine Learning of Treatment Effects, American Economic Review, 107, 261--265

  5. [13]

    Chetverikov, M

    Chernozhukov, V., D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W. Newey, and J. Robins (2018 a ): Double/Debiased Machine Learning for Treatment and Structural Parameters, The Econometrics Journal, 21, C1--C68

  6. [14]

    Chernozhukov, V., J. C. Escanciano, H. Ichimura, W. K. Newey, and J. M. Robins (2018 b ): Locally Robust Semiparametric Estimation, Working paper

  7. [15]

    Chernozhukov, V., W. K. Newey, and J. M. Robins (2018 c ): Double/De-Biased Machine Learning Using Regularized Riesz Representers, Working paper

  8. [16]

    (2006): Limit theorems for dependent U-statistics, in Dependence in Probability and Statistics, Springer, 65--86

    Dehling, H. (2006): Limit theorems for dependent U-statistics, in Dependence in Probability and Statistics, Springer, 65--86

  9. [17]

    Hampel, F. R. (1974): The influence curve and its role in robust estimation, Journal of the American Statistical Association, 69, 383--393

  10. [18]

    Hirano, K., G. W. Imbens, and G. Ridder (2003): Efficient Estimation of Average Treatment Effects Using the Estimated Propensity Score, Econometrica, 71, 1161--1189

  11. [19]

    (1948): A Class of Statistics with Asymptotically Normal Distribution, Annals of Statistics, 19, 293--325

    Hoeffding, W. (1948): A Class of Statistics with Asymptotically Normal Distribution, Annals of Statistics, 19, 293--325

  12. [20]

    Ichimura, H. and W. K. Newey (2017): The Influence Function of Semiparametric Estimators, Working paper

  13. [21]

    Ichimura, H. and P. E. Todd (2007): Implementing Nonparametric and Semiparametric Estimators, in Handbook of Econometrics, ed. by J. J. Heckman and E. E. Leamer, Elsevier, vol. 6, 5549--5632

  14. [22]

    Kollo, T. and D. von Rosen (2006): Advanced multivariate statistics with matrices, vol. 579, Springer Science & Business Media

  15. [23]

    Korolyuk, V. S. and Y. V. Borovskich (1994): Theory of U-statistics, vol. 273 of Mathematics and Its Applications, Springer

  16. [24]

    Liu, and D

    Li, J., Y. Liu, and D. Xiu (2019): Efficient Estimation of Integrated Volatility Functionals via Multiscale Jackknife, The Annals of Statistics, 47, 156--176

  17. [25]

    Nadaraya, E. A. (1964): On estimating regression, Theory of Probability & Its Applications, 9, 141--142

  18. [26]

    Newey, W. K. (1994): The Asymptotic Variance of Semiparametric Estimators, Econometrica, 1349--1382

  19. [27]

    Newey, W. K., F. Hsieh, and J. M. Robins (2004): Twicing Kernels and A Small Bias Property of Semiparametric Estimators, Econometrica, 72, 947--962

  20. [28]

    Newey, W. K. and D. McFadden (1994): Large Sample Estimation and Hypothesis Testing, in Handbook of Econometrics, ed. by R. F. Engle and D. L. McFadden, Elsevier, vol. 4, 2111--2245

  21. [29]

    Powell, J. L., J. H. Stock, and T. M. Stoker (1989): Semiparametric Estimation of Index Coefficients, Econometrica, 57, 1403--1430

  22. [30]

    Quenouille, M. H. (1949): Problems in Plane Sampling, The Annals of Mathematical Statistics, 20, 335--375

  23. [31]

    Schucany, W. and J. P. Sommers (1977): Improvement of Kernel Type Density Estimators, Journal of the American Statistical Association, 72, 420--423

  24. [32]

    Stuetzle, W. and Y. Mittal (1979): Some Comments on The Asymptotic Behavior of Robust Smoothers, in Smoothing Techniques for Curve Estimation, Springer, vol. 757 of Lecture Notes in Mathematics, 191--195

  25. [33]

    Watson, G. S. (1964): Smooth regression analysis, Sankhy \=a : The Indian Journal of Statistics, Series A , 359--372

  26. [34]

    Yang, X. (2020): Semiparametric Estimation in Continuous-Time: Asymptotics for Integrated Volatility Functionals with Small and Large Bandwidths, Journal of Business & Economic Statistics, forthcoming

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.