Pith. sign in

REVIEW 2 major objections 5 minor 3 references

A Response to Recent Critiques of Hainmueller, Mummolo and Xu (2019) on Estimating Conditional Relationships

T0 review · 2 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read Recent critiques of the kernel estimator for interaction effects compared it against the wrong target; once the conditional marginal effect is defined as the estimand, the estimator recovers the true effect in the very simulation said to…

desk verdict The specific rebuttal of Simonsohn's key example is correct, but the appendix overclaims general consistency for the kernel estimator. read the letter →

arxiv 2502.05717 v1 pith:T53RCZVF submitted 2025-02-08 stat.ME

classification stat.ME MSC 62G0862G20
keywords conditionalmarginaleffectinteractionmodelskernelestimatorlocallinearregressionestimandGAMdouble/debiasedmachinelearningcausalmoderation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Recent critiques of the kernel estimator for interaction effects argued that it produces spurious marginal effects when the treatment and moderator are correlated and nonlinearities exist. This response paper argues that the critiques never properly defined their causal target: the quantity the kernel estimator is built to recover is the conditional marginal effect (CME), the expected derivative of the potential outcome with respect to treatment among units with a fixed moderator value. In the critiques' central simulation, the true CME is not zero but $x - 0.5$, and the kernel estimator recovers it, while the linear interaction model is biased. The paper also contends that the alternative the critiques recommend, GAM, estimates a different object (the conditional average partial effect given a treatment level), so its outputs should not be read as marginal effects. The upshot is a call to define estimands explicitly and to adopt doubly robust or machine-learning-based estimators when covariates are high-dimensional.

What carries the argument

The load-bearing object is the conditional marginal effect (CME), defined as $\theta(x) = E[\partial Y_i(d)/\partial d \mid X_i = x]$, which is the causal effect of the treatment on the outcome among units with moderator value $x$, marginalized over the treatment distribution and other covariates. The estimator that carries the argument is the kernel estimator, a local linear regression of $Y$ on $(1, D, X - x_0, D(X - x_0))$ with kernel weights shrinking the bandwidth. Near $x_0$, the interaction term vanishes and the coefficient on $D$ directly approximates the local average derivative of the outcome in $D$, which is exactly the CME. The distinction between CME and the conditional average partial effect (CAPE) that conditions on $D=d$ is what the paper says the critiques conflate.

What would settle it

Simulate a smooth response function $Y = s(D,X)+\varepsilon$ with large $n$, fix a point $x_0$, and draw $D\mid X$ from a skewed or otherwise non-normal distribution whose density changes with $X$ so that the local projection of $Y$ onto $D$ does not equal $E[\partial s/\partial D \mid X=x_0]$. If the kernel estimator's estimated coefficient on $D$ fails to converge to the CME as the bandwidth shrinks, the paper's broad consistency claim is falsified.

Watch

Extended reading notes

Core claim

Stated in the paper's own terms: the conditional marginal effect is $\theta(x) = E[\partial Y_i(d)/\partial d \mid X_i = x]$, and under the unconfoundedness and overlap assumptions it is identified nonparametrically. In the simulation disputed by the critiques, $Y = D^2 - 0.5D + \varepsilon$ with $(D,X)$ bivariate normal and correlation 0.5, the true CME is $\theta(x) = x - 0.5$. The kernel estimator, implemented as a local linear regression of $Y$ on $(1, D, X - x_0, D(X - x_0))$, consistently recovers this curve where data are dense; the linear interaction model does not. The paper generalizes this: for any smooth response function $s(D,X)$, the local-linear coefficient on $D$ converges to the CME in regions with abundant data. The paper concludes that the critiques' asserted spuriousness is an artifact of evaluating the estimator against the wrong estimand.

Load-bearing premise

The general claim that the kernel estimator consistently estimates the CME for any smooth response function rests on the unproven assumption that the local linear regression slope converges to the average derivative $E[\partial s(D,x_0)/\partial D]$, which is not implied by local polynomial regression alone and holds in the paper's example because the conditional distribution of $D$ given $X$ is normal.

Editorial extensions

If this is right

  • Applied researchers studying effect heterogeneity can treat the kernel estimator as a valid tool for the CME when the sample is large enough to provide dense support on the moderator.
  • Linear interaction models should not be used as the benchmark for judging whether a marginal effect is real, because they are themselves biased under nonlinearity.
  • GAM-based 'marginal effects' computed by finite differences should be reported as conditional on a chosen treatment level, not as average causal effects.
  • The recommended workflow—state the estimand, check overlap, use diagnostics, and switch to doubly robust or DML estimators with many covariates—provides a concrete template for interaction research.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the CME/CAPE distinction becomes standard, many past studies that reported marginal effects from finite differences of predicted outcomes may need to be re-read as answering a conditional question, not an average one.
  • The paper's general consistency proof relies on the local projection slope equaling the average derivative; this holds in the worked example because $D\mid X$ is normal. A natural test is to repeat the simulation with skewed or heteroskedastic $D\mid X$ and see whether the kernel estimate drifts.
  • The debate suggests a broader lesson for method comparisons: a simulation can only convict an estimator if the target quantity is specified in advance; otherwise the same numbers can support opposite conclusions.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper is a response to Simonsohn (2024a,b), which criticized the binning and kernel estimators of Hainmueller, Mummolo and Xu (2019). The authors define the conditional marginal effect (CME) as E[∂Y_i(d)/∂d | X_i=x], distinguish it from the conditional average partial effect (CAPE) and from causal moderation, and argue that Simonsohn's simulations use the wrong benchmark. They reanalyze Simonsohn's main DGP (Y=D^2−0.5D, Corr(D,X)=0.5), derive the true CME θ(x)=x−0.5 and the population linear interaction slope −0.5+0.8x, and argue that the HMX kernel estimator recovers θ(x), while Simonsohn's GAM-based procedure estimates a CAPE. They also recommend doubly robust and DML estimators. An appendix claims a general consistency result for smooth response functions s(D,X).

Significance. If the central claims are correct, the paper would be a useful clarification of estimands in interaction-effects research and would rebut the claim that the HMX kernel estimator is inherently biased in the criticized scenarios. The paper's strengths include exact population derivations for the quadratic example, a clear table of estimands, explicit simulation DGPs, and a correct demonstration that Simonsohn's GAM procedure as implemented targets a CAPE rather than a CME. However, the paper's broad generalization beyond the quadratic example rests on an unproven property of local polynomial regression, and this is load-bearing for the claim that the kernel estimator works in 'most, if not all' of Simonsohn's examples. The conceptual distinction between CME and CAPE is valuable, but the manuscript's defense of the kernel estimator is currently too sweeping and needs a precise theorem or a clearly stated set of sufficient conditions.

major comments (2)
  1. [Appendix A.1, Extension to more complex settings] The proof's final step, 'By the properties of local polynomial regression, β̂1 → ∂s(D,X)/∂D|X=x0', is not a valid property of local polynomial regression. In a shrinking X-window, the coefficient on D in the local linear fit converges to the least-squares projection slope Cov(D, s(D,x0) | X=x0)/Var(D | X=x0), not generally to E[∂s(D,x0)/∂D | X=x0]. The equality holds under special conditions, such as conditional normality of D given X (by Stein's lemma) or s affine in D, but the manuscript neither states nor proves such a condition. Since the main text uses this appendix to claim that the kernel estimator 'can still consistently estimate the CME in regions of X where data are sufficiently abundant' for any smooth s(D,X), and to conjecture that it works in 'most, if not all' of Simonsohn's examples, the broad claim is unsupported. The quadratic example is saved by conditional normality, but that fact is not acknowledged as a condition.
  2. [Appendix A.1, The kernel estimator can recover the CME] The derivation of the kernel estimator's limit, β̂1(x0) ≈ Cov(D,Y|X≈x0)/Var(D|X≈x0) ≈ x0−0.5, relies on the bivariate normal assumption for (D,X). For non-Gaussian D|X, this ratio can differ from the CME, and the subsequent paragraph generalizes the claim to arbitrary smooth s(D,X) without adding any distributional condition. The Taylor expansion around X=x0 only linearizes in X; s(D,x0) remains an arbitrary function of D, while the local regression is linear in D. A least-squares slope of s(D,x0) on D is a projection, not a derivative. The paper should either state the conditional-normality (or equivalent) condition as part of the theorem, or restrict the consistency claim to settings where the local polynomial includes sufficient flexibility in D.
minor comments (5)
  1. [GAM is Not an Appealing Alternative] The statement that 'mgcv can handle at most two variables, including D' is inaccurate; mgcv supports smooths of more than two covariates via tensor-product terms. Please correct or qualify this claim.
  2. [Figure 2 and Figure 3 captions] The text and captions mention both pointwise and uniform confidence intervals but do not consistently specify which is the shaded region and which is the dashed line; please clarify in each caption.
  3. [Assumption 1 and references] Assumption 1 is labeled 'SUTV A' (a typo for SUTVA), and the Brambor, Clark and Golder reference has a typo in the title ('Mdels'); please proofread.
  4. [Equation (1)] Equation (1) calls ε an idiosyncratic error without stating assumptions; adding E[ε|D,X,Z]=0 or an equivalent condition would make the population regression calculations in Appendix A.1 more self-contained.
  5. [Binning estimator discussion] The claim that the binning estimator 'will be biased but will have less bias than the linear estimator' is asserted without derivation; since the binning estimator is presented as a diagnostic tool, a one-sentence justification or a reference would be helpful.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation: the kernel estimator is validated against an externally defined CME; self-citations are contextual, not load-bearing.

full rationale

The paper's central claims are not circular. The true CME in Simonsohn's example, θ(x)=x−0.5, is derived from the stated DGP (Y=D^2−0.5D+ε with (D,X) bivariate normal), and the kernel estimates in Figure 2 are produced by an independent local-linear algorithm; the estimator is not fit to the target curve and no fitted parameter is relabeled as a prediction. The derivation for the quadratic example is a substantive calculation: the local projection slope equals the average derivative because D|X is normal, not because the estimand was inserted into the estimator. The paper does contain self-citations (HMX 2019; Liu, Liu and Xu, under contract; Hainmueller and Hazlett 2014), but these identify the estimator and support practical recommendations; they are not the load-bearing justification for the numerical or analytic results, which stand on the external DGP and independent simulation. One flagged concern is the appendix's 'Extension to more complex settings' (Appendix A.1): the assertion that β1 → ∂s(D,X)/∂D | X=x0 'by the properties of local polynomial regression' is mathematically unsupported — the local-linear slope is generally Cov(g(D),D|X=x0)/Var(D|X=x0), equal to the average derivative only under extra conditions such as conditional normality — but this is an invalid proof step, not a circularity, because β1 is not defined in terms of θ(x0). Overall, the derivation chain is self-contained against external benchmarks, so the appropriate circularity finding is minimal (score 1), with the appendix proof gap belonging to correctness risk rather than circularity.

Assumptions & free parameters 1 free parameters · 6 assumptions · 0 invented entities

The central derivations introduce no new entities or fitted constants. The main unstated input is the conditional distribution of the treatment given the moderator, which makes the broad kernel consistency claim distribution-dependent. The kernel bandwidth is an unreported tuning parameter.

free parameters (1)
  • Kernel bandwidth h(x0) = not reported
    The finite-sample simulation figures depend on the bandwidth of the kernel estimator, which is not specified in the paper. The asymptotic argument requires h→0 with n·h→∞, but the actual choices behind Figures 2 through 4 are not given.
assumptions (6)
  • standard math SUTVA (Assumption 1): potential outcome of unit i depends only on its own treatment
    Standard stable unit treatment value assumption, stated in Section 'Properly Defining the Estimand Is Critical'.
  • domain assumption Unconfoundedness (Assumption 2)
    Used to nonparametrically identify the CME under the Neyman-Rubin framework.
  • domain assumption Strict overlap (Assumption 3)
    Requires positive conditional density of treatment given covariates; used for identification of causal quantities.
  • domain assumption Partially linear model (Assumption 4)
    Y = g(V)D + f(V) + epsilon with E[D|V]=e(V), used to motivate DML and doubly robust estimators.
  • domain assumption Bivariate normality of (D,X) in the critique's DGP
    Used in Appendix A.1 to derive E[D|X=x]=0.5x and the true CME x−0.5; it is also the hidden condition that makes the local linear projection coefficient equal the average derivative.
  • ad hoc to paper Local linear projection slope converges to the average derivative E[∂s(D,X0)/∂D] for smooth s
    Appendix A.1, 'Extension to more complex settings', asserts this as a property of local polynomial regression without proof. It is not true for arbitrary conditional distributions of D.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Response to Recent Critiques of Hainmueller, Mummolo and Xu (2019) on Estimating Conditional Relationships." pith.science (2026). https://pith.science/paper/T53RCZVF

@misc{pith2026250205717,
  author       = {Pith},
  title        = {Pith review of: A Response to Recent Critiques of Hainmueller, Mummolo and Xu (2019) on Estimating Conditional Relationships},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/T53RCZVF}},
  note         = {Machine review of arXiv:2502.05717}
}
read the original abstract

Simonsohn (2024a) and Simonsohn (2024b) critique Hainmueller, Mummolo and Xu (2019, HMX), arguing that failing to model nonlinear relationships between the treatment and moderator leads to biased marginal effect estimates and uncontrolled Type-I error rates. While these critiques highlight the issue of under-modeling nonlinearity in applied research, they are fundamentally flawed in several key ways. First, the causal estimand for interaction effects and the necessary identifying assumptions are not clearly defined in these critiques. Once properly stated, the critiques no longer hold. Second, the kernel estimator HMX proposes recovers the true causal effects in the scenarios presented in these recent critiques, which compared effects to the wrong benchmark, producing misleading conclusions. Third, while Generalized Additive Models (GAM) can be a useful exploratory tool (as acknowledged in HMX), they are not designed to estimate marginal effects, and better alternatives exist, particularly in the presence of additional covariates. Our response aims to clarify these misconceptions and provide updated recommendations for researchers studying interaction effects through the estimation of conditional marginal effects.

Figures

Figures reproduced from arXiv: 2502.05717 by the authors.

Figure 1
Figure 1. Figure adapted from Simonsohn (2024a) (with Modified Labels) Note: In the original plot, the treatment and moderator were labeled as X and Z, respec￾tively; we have modified them to D and X to maintain consistency with the notation used in this response. Merits of the Critique. There is value in these critiques. First, they bring renewed attention to the issue of potential nonlinearity in estimating and testing cond… view at source ↗
Figure 3
Figure 3. Estimating the CME with Various Estimators [PITH_FULL_IMAGE:figures/full_fig_p015_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

3 extracted references · 2 canonical work pages

  1. [2009]

    On the Distinction Between Interaction and Effect Modifica- tion

    “On the Distinction Between Interaction and Effect Modifica- tion.” Epidemiology 20:863–871. 20 A. Appendix A.1. Dissecting the Key Simulated Example We begin by deriving the true CME. We then show that the linear interaction model is biased, while the kernel estimator proposed in HMX (2019) can recover the CME when data are abundant. The true CME. The pa...

  2. [2021]

    Debiased Machine Learning of Condi- tional Average Treatment Effects and Other Causal Functions

    “Debiased Machine Learning of Condi- tional Average Treatment Effects and Other Causal Functions.”The Econometrics Journal 24(2):264–289. Simonsohn, Uri. 2024 a. “Dear Political Scientists: Don’t Bin, GAM Instead.” Data Colada blog. URL: https://datacolada.org/121 Simonsohn, Uri. 2024 b. “Interacting with Curves: How to Validly Test and Probe interac- tio...

  3. [2024]

    LaLonde (1986) after Nearly Four Decades: Lessons Learned

    “LaLonde (1986) after Nearly Four Decades: Lessons Learned.” arXiv preprint arXiv:2406.00827 . Liu, Jiehan, Ziyi Liu and Yiqing Xu. N.d. “Estimating and Interpreting Conditional Marginal Effects.” under contract, Cambrdige University Press. Lundberg, Ian, Rebecca Johnson and Brandon M Stewart

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.