REVIEW 2 major objections 5 minor 3 references
A Response to Recent Critiques of Hainmueller, Mummolo and Xu (2019) on Estimating Conditional Relationships
T0 review · 2 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read Recent critiques of the kernel estimator for interaction effects compared it against the wrong target; once the conditional marginal effect is defined as the estimand, the estimator recovers the true effect in the very simulation said to…
desk verdict The specific rebuttal of Simonsohn's key example is correct, but the appendix overclaims general consistency for the kernel estimator. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the conditional marginal effect (CME), defined as $\theta(x) = E[\partial Y_i(d)/\partial d \mid X_i = x]$, which is the causal effect of the treatment on the outcome among units with moderator value $x$, marginalized over the treatment distribution and other covariates. The estimator that carries the argument is the kernel estimator, a local linear regression of $Y$ on $(1, D, X - x_0, D(X - x_0))$ with kernel weights shrinking the bandwidth. Near $x_0$, the interaction term vanishes and the coefficient on $D$ directly approximates the local average derivative of the outcome in $D$, which is exactly the CME. The distinction between CME and the conditional average partial effect (CAPE) that conditions on $D=d$ is what the paper says the critiques conflate.
What would settle it
Simulate a smooth response function $Y = s(D,X)+\varepsilon$ with large $n$, fix a point $x_0$, and draw $D\mid X$ from a skewed or otherwise non-normal distribution whose density changes with $X$ so that the local projection of $Y$ onto $D$ does not equal $E[\partial s/\partial D \mid X=x_0]$. If the kernel estimator's estimated coefficient on $D$ fails to converge to the CME as the bandwidth shrinks, the paper's broad consistency claim is falsified.
Extended reading notes
Core claim
Stated in the paper's own terms: the conditional marginal effect is $\theta(x) = E[\partial Y_i(d)/\partial d \mid X_i = x]$, and under the unconfoundedness and overlap assumptions it is identified nonparametrically. In the simulation disputed by the critiques, $Y = D^2 - 0.5D + \varepsilon$ with $(D,X)$ bivariate normal and correlation 0.5, the true CME is $\theta(x) = x - 0.5$. The kernel estimator, implemented as a local linear regression of $Y$ on $(1, D, X - x_0, D(X - x_0))$, consistently recovers this curve where data are dense; the linear interaction model does not. The paper generalizes this: for any smooth response function $s(D,X)$, the local-linear coefficient on $D$ converges to the CME in regions with abundant data. The paper concludes that the critiques' asserted spuriousness is an artifact of evaluating the estimator against the wrong estimand.
Load-bearing premise
The general claim that the kernel estimator consistently estimates the CME for any smooth response function rests on the unproven assumption that the local linear regression slope converges to the average derivative $E[\partial s(D,x_0)/\partial D]$, which is not implied by local polynomial regression alone and holds in the paper's example because the conditional distribution of $D$ given $X$ is normal.
Editorial extensions
If this is right
- Applied researchers studying effect heterogeneity can treat the kernel estimator as a valid tool for the CME when the sample is large enough to provide dense support on the moderator.
- Linear interaction models should not be used as the benchmark for judging whether a marginal effect is real, because they are themselves biased under nonlinearity.
- GAM-based 'marginal effects' computed by finite differences should be reported as conditional on a chosen treatment level, not as average causal effects.
- The recommended workflow—state the estimand, check overlap, use diagnostics, and switch to doubly robust or DML estimators with many covariates—provides a concrete template for interaction research.
Reading between the lines
- If the CME/CAPE distinction becomes standard, many past studies that reported marginal effects from finite differences of predicted outcomes may need to be re-read as answering a conditional question, not an average one.
- The paper's general consistency proof relies on the local projection slope equaling the average derivative; this holds in the worked example because $D\mid X$ is normal. A natural test is to repeat the simulation with skewed or heteroskedastic $D\mid X$ and see whether the kernel estimate drifts.
- The debate suggests a broader lesson for method comparisons: a simulation can only convict an estimator if the target quantity is specified in advance; otherwise the same numbers can support opposite conclusions.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a response to Simonsohn (2024a,b), which criticized the binning and kernel estimators of Hainmueller, Mummolo and Xu (2019). The authors define the conditional marginal effect (CME) as E[∂Y_i(d)/∂d | X_i=x], distinguish it from the conditional average partial effect (CAPE) and from causal moderation, and argue that Simonsohn's simulations use the wrong benchmark. They reanalyze Simonsohn's main DGP (Y=D^2−0.5D, Corr(D,X)=0.5), derive the true CME θ(x)=x−0.5 and the population linear interaction slope −0.5+0.8x, and argue that the HMX kernel estimator recovers θ(x), while Simonsohn's GAM-based procedure estimates a CAPE. They also recommend doubly robust and DML estimators. An appendix claims a general consistency result for smooth response functions s(D,X).
Significance. If the central claims are correct, the paper would be a useful clarification of estimands in interaction-effects research and would rebut the claim that the HMX kernel estimator is inherently biased in the criticized scenarios. The paper's strengths include exact population derivations for the quadratic example, a clear table of estimands, explicit simulation DGPs, and a correct demonstration that Simonsohn's GAM procedure as implemented targets a CAPE rather than a CME. However, the paper's broad generalization beyond the quadratic example rests on an unproven property of local polynomial regression, and this is load-bearing for the claim that the kernel estimator works in 'most, if not all' of Simonsohn's examples. The conceptual distinction between CME and CAPE is valuable, but the manuscript's defense of the kernel estimator is currently too sweeping and needs a precise theorem or a clearly stated set of sufficient conditions.
major comments (2)
- [Appendix A.1, Extension to more complex settings] The proof's final step, 'By the properties of local polynomial regression, β̂1 → ∂s(D,X)/∂D|X=x0', is not a valid property of local polynomial regression. In a shrinking X-window, the coefficient on D in the local linear fit converges to the least-squares projection slope Cov(D, s(D,x0) | X=x0)/Var(D | X=x0), not generally to E[∂s(D,x0)/∂D | X=x0]. The equality holds under special conditions, such as conditional normality of D given X (by Stein's lemma) or s affine in D, but the manuscript neither states nor proves such a condition. Since the main text uses this appendix to claim that the kernel estimator 'can still consistently estimate the CME in regions of X where data are sufficiently abundant' for any smooth s(D,X), and to conjecture that it works in 'most, if not all' of Simonsohn's examples, the broad claim is unsupported. The quadratic example is saved by conditional normality, but that fact is not acknowledged as a condition.
- [Appendix A.1, The kernel estimator can recover the CME] The derivation of the kernel estimator's limit, β̂1(x0) ≈ Cov(D,Y|X≈x0)/Var(D|X≈x0) ≈ x0−0.5, relies on the bivariate normal assumption for (D,X). For non-Gaussian D|X, this ratio can differ from the CME, and the subsequent paragraph generalizes the claim to arbitrary smooth s(D,X) without adding any distributional condition. The Taylor expansion around X=x0 only linearizes in X; s(D,x0) remains an arbitrary function of D, while the local regression is linear in D. A least-squares slope of s(D,x0) on D is a projection, not a derivative. The paper should either state the conditional-normality (or equivalent) condition as part of the theorem, or restrict the consistency claim to settings where the local polynomial includes sufficient flexibility in D.
minor comments (5)
- [GAM is Not an Appealing Alternative] The statement that 'mgcv can handle at most two variables, including D' is inaccurate; mgcv supports smooths of more than two covariates via tensor-product terms. Please correct or qualify this claim.
- [Figure 2 and Figure 3 captions] The text and captions mention both pointwise and uniform confidence intervals but do not consistently specify which is the shaded region and which is the dashed line; please clarify in each caption.
- [Assumption 1 and references] Assumption 1 is labeled 'SUTV A' (a typo for SUTVA), and the Brambor, Clark and Golder reference has a typo in the title ('Mdels'); please proofread.
- [Equation (1)] Equation (1) calls ε an idiosyncratic error without stating assumptions; adding E[ε|D,X,Z]=0 or an equivalent condition would make the population regression calculations in Appendix A.1 more self-contained.
- [Binning estimator discussion] The claim that the binning estimator 'will be biased but will have less bias than the linear estimator' is asserted without derivation; since the binning estimator is presented as a diagnostic tool, a one-sentence justification or a reference would be helpful.
Circularity Check
No circular derivation: the kernel estimator is validated against an externally defined CME; self-citations are contextual, not load-bearing.
full rationale
The paper's central claims are not circular. The true CME in Simonsohn's example, θ(x)=x−0.5, is derived from the stated DGP (Y=D^2−0.5D+ε with (D,X) bivariate normal), and the kernel estimates in Figure 2 are produced by an independent local-linear algorithm; the estimator is not fit to the target curve and no fitted parameter is relabeled as a prediction. The derivation for the quadratic example is a substantive calculation: the local projection slope equals the average derivative because D|X is normal, not because the estimand was inserted into the estimator. The paper does contain self-citations (HMX 2019; Liu, Liu and Xu, under contract; Hainmueller and Hazlett 2014), but these identify the estimator and support practical recommendations; they are not the load-bearing justification for the numerical or analytic results, which stand on the external DGP and independent simulation. One flagged concern is the appendix's 'Extension to more complex settings' (Appendix A.1): the assertion that β1 → ∂s(D,X)/∂D | X=x0 'by the properties of local polynomial regression' is mathematically unsupported — the local-linear slope is generally Cov(g(D),D|X=x0)/Var(D|X=x0), equal to the average derivative only under extra conditions such as conditional normality — but this is an invalid proof step, not a circularity, because β1 is not defined in terms of θ(x0). Overall, the derivation chain is self-contained against external benchmarks, so the appropriate circularity finding is minimal (score 1), with the appendix proof gap belonging to correctness risk rather than circularity.
Assumptions & free parameters
free parameters (1)
- Kernel bandwidth h(x0) =
not reported
assumptions (6)
- standard math SUTVA (Assumption 1): potential outcome of unit i depends only on its own treatment
- domain assumption Unconfoundedness (Assumption 2)
- domain assumption Strict overlap (Assumption 3)
- domain assumption Partially linear model (Assumption 4)
- domain assumption Bivariate normality of (D,X) in the critique's DGP
- ad hoc to paper Local linear projection slope converges to the average derivative E[∂s(D,X0)/∂D] for smooth s
Cite this review
Pith. "Pith review of A Response to Recent Critiques of Hainmueller, Mummolo and Xu (2019) on Estimating Conditional Relationships." pith.science (2026). https://pith.science/paper/T53RCZVF
@misc{pith2026250205717,
author = {Pith},
title = {Pith review of: A Response to Recent Critiques of Hainmueller, Mummolo and Xu (2019) on Estimating Conditional Relationships},
year = {2026},
howpublished = {\url{https://pith.science/paper/T53RCZVF}},
note = {Machine review of arXiv:2502.05717}
}
read the original abstract
Simonsohn (2024a) and Simonsohn (2024b) critique Hainmueller, Mummolo and Xu (2019, HMX), arguing that failing to model nonlinear relationships between the treatment and moderator leads to biased marginal effect estimates and uncontrolled Type-I error rates. While these critiques highlight the issue of under-modeling nonlinearity in applied research, they are fundamentally flawed in several key ways. First, the causal estimand for interaction effects and the necessary identifying assumptions are not clearly defined in these critiques. Once properly stated, the critiques no longer hold. Second, the kernel estimator HMX proposes recovers the true causal effects in the scenarios presented in these recent critiques, which compared effects to the wrong benchmark, producing misleading conclusions. Third, while Generalized Additive Models (GAM) can be a useful exploratory tool (as acknowledged in HMX), they are not designed to estimate marginal effects, and better alternatives exist, particularly in the presence of additional covariates. Our response aims to clarify these misconceptions and provide updated recommendations for researchers studying interaction effects through the estimation of conditional marginal effects.
Figures
Reference graph
Works this paper leans on
-
[2009]
On the Distinction Between Interaction and Effect Modifica- tion
“On the Distinction Between Interaction and Effect Modifica- tion.” Epidemiology 20:863–871. 20 A. Appendix A.1. Dissecting the Key Simulated Example We begin by deriving the true CME. We then show that the linear interaction model is biased, while the kernel estimator proposed in HMX (2019) can recover the CME when data are abundant. The true CME. The pa...
work page 2019
-
[2021]
Debiased Machine Learning of Condi- tional Average Treatment Effects and Other Causal Functions
“Debiased Machine Learning of Condi- tional Average Treatment Effects and Other Causal Functions.”The Econometrics Journal 24(2):264–289. Simonsohn, Uri. 2024 a. “Dear Political Scientists: Don’t Bin, GAM Instead.” Data Colada blog. URL: https://datacolada.org/121 Simonsohn, Uri. 2024 b. “Interacting with Curves: How to Validly Test and Probe interac- tio...
work page 2024
-
[2024]
LaLonde (1986) after Nearly Four Decades: Lessons Learned
“LaLonde (1986) after Nearly Four Decades: Lessons Learned.” arXiv preprint arXiv:2406.00827 . Liu, Jiehan, Ziyi Liu and Yiqing Xu. N.d. “Estimating and Interpreting Conditional Marginal Effects.” under contract, Cambrdige University Press. Lundberg, Ian, Rebecca Johnson and Brandon M Stewart
arXiv 1986
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.