{"id":"6c9ed140-55ff-431d-b4e0-f42a8ae28d1d","arxiv_id":"2502.05717","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"The HMX kernel estimator recovers the conditional marginal effect in Simonsohn's simulated example, while the GAM procedure he recommends estimates a different, treatment-level-specific quantity.","lead":"This paper rebuts recent critiques of the HMX 2019 toolkit for estimating conditional treatment effects, arguing the critiques used the wrong benchmark and that the kernel estimator recovers the intended conditional marginal effect. It also warns that GAM-based simple slopes estimate a different quantity, the conditional average partial effect, not the causal marginal effect.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"General consistency claim in Appendix A.1 is invalid: the local-linear slope on D converges to Cov(D, s(D,x0)|X=x0)/Var(D|X=x0), not to E[∂s/∂D|X=x0], absent conditions like conditional normality.","rationale":"The reader and I converge on the same weakest point. The response's central example is handled correctly: under the stated bivariate normal DGP, D|X=x is Gaussian, so the local OLS slope on D equals the average derivative by Stein's lemma; the linear-model bias calculation and the GAM-as-CAPE argument are also persuasive. The problem is the unqualified generalization in Appendix A.1, which is not a 'simple proof' but a misapplication of local polynomial regression. The local polynomial is linear in D and local only in X; the resulting coefficient is a projection functional of s(D,x0), not the average derivative. Without an assumption such as conditional normality of D|X (or some other condition making the projection equal the derivative), the consistency claim is false, and a concrete uniform-D counterexample exhibits the bias. This does not overturn the paper's specific rebuttal of Simonsohn, but it does undermine the response's broad claim that the kernel estimator recovers the CME in arbitrary smooth nonlinear DGPs, and it is the basis for the conclusion that the critiques are fundamentally flawed. The right remedy is to restrict the claim to settings where the needed projection condition holds (or prove it under explicit conditions), or to recommend estimators that directly target the average derivative. I therefore keep the reader's CONDITIONAL verdict; the appendix needs correction.","tokens_in":11706,"tokens_out":9341,"duration_ms":97517,"concrete_test":"Simulate n = 100,000 with X ~ Unif[-1,1], D|X ~ Unif[0,1] independent of X, and Y = D^3. The true CME is θ(x) = E[3D^2] = 1 at every x. Apply the HMX kernel/interflex estimator at x0 = 0 with bandwidth h ≈ n^{-1/5} (e.g., h = 0.2) and increase n or shrink h to approximate the h→0 limit. If the estimated CME converges to approximately 0.90 rather than 1, the Appendix A.1 consistency claim is falsified, because for uniform D, Cov(D,D^3)/Var(D) = (1/5 − 1/8)/(1/12) = 0.9, while E[3D^2] = 1.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Appendix A.1's 'Extension to more complex settings' is the load-bearing step for the paper's broad assertion that the HMX kernel estimator consistently estimates the CME for arbitrary smooth response functions. The local linear regression Y = β0 + β1D + β2(X−x0) + β3D(X−x0) + ε is only local in X; within a shrinking X-window, the coefficient β1 is a least-squares projection of s(D,x0) on D. As h→0, β1 → Cov(D, s(D,x0) | X=x0) / Var(D | X=x0). This equals E[∂s(D,x0)/∂D | X=x0] only under special conditions, e.g., D|X=x0 Gaussian (Stein's lemma) or s linear in D; it is not a property of local polynomial regression. The paper's proof simply asserts convergence 'by properties of local polynomial regression,' but local polynomial regression recovers derivatives with respect to regressors included in the local polynomial, not the projection slope of an arbitrary function of D. The Simonsohn example is saved only because the DGP makes D|X normal. The broad generalization, and the claim that the kernel estimator recovers the CME in 'most, if not all' of Simonsohn's examples, is therefore unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a response to Simonsohn (2024a,b), which criticized the binning and kernel estimators of Hainmueller, Mummolo and Xu (2019). The authors define the conditional marginal effect (CME) as E[∂Y_i(d)/∂d | X_i=x], distinguish it from the conditional average partial effect (CAPE) and from causal moderation, and argue that Simonsohn's simulations use the wrong benchmark. They reanalyze Simonsohn's main DGP (Y=D^2−0.5D, Corr(D,X)=0.5), derive the true CME θ(x)=x−0.5 and the population linear interaction slope −0.5+0.8x, and argue that the HMX kernel estimator recovers θ(x), while Simonsohn's GAM-based procedure estimates a CAPE. They also recommend doubly robust and DML estimators. An appendix claims a general consistency result for smooth response functions s(D,X).","tokens_in":12008,"tokens_out":10309,"duration_ms":108748,"significance":"If the central claims are correct, the paper would be a useful clarification of estimands in interaction-effects research and would rebut the claim that the HMX kernel estimator is inherently biased in the criticized scenarios. The paper's strengths include exact population derivations for the quadratic example, a clear table of estimands, explicit simulation DGPs, and a correct demonstration that Simonsohn's GAM procedure as implemented targets a CAPE rather than a CME. However, the paper's broad generalization beyond the quadratic example rests on an unproven property of local polynomial regression, and this is load-bearing for the claim that the kernel estimator works in 'most, if not all' of Simonsohn's examples. The conceptual distinction between CME and CAPE is valuable, but the manuscript's defense of the kernel estimator is currently too sweeping and needs a precise theorem or a clearly stated set of sufficient conditions.","major_comments":[{"comment":"The proof's final step, 'By the properties of local polynomial regression, β̂1 → ∂s(D,X)/∂D|X=x0', is not a valid property of local polynomial regression. In a shrinking X-window, the coefficient on D in the local linear fit converges to the least-squares projection slope Cov(D, s(D,x0) | X=x0)/Var(D | X=x0), not generally to E[∂s(D,x0)/∂D | X=x0]. The equality holds under special conditions, such as conditional normality of D given X (by Stein's lemma) or s affine in D, but the manuscript neither states nor proves such a condition. Since the main text uses this appendix to claim that the kernel estimator 'can still consistently estimate the CME in regions of X where data are sufficiently abundant' for any smooth s(D,X), and to conjecture that it works in 'most, if not all' of Simonsohn's examples, the broad claim is unsupported. The quadratic example is saved by conditional normality, but that fact is not acknowledged as a condition.","section":"Appendix A.1, Extension to more complex settings"},{"comment":"The derivation of the kernel estimator's limit, β̂1(x0) ≈ Cov(D,Y|X≈x0)/Var(D|X≈x0) ≈ x0−0.5, relies on the bivariate normal assumption for (D,X). For non-Gaussian D|X, this ratio can differ from the CME, and the subsequent paragraph generalizes the claim to arbitrary smooth s(D,X) without adding any distributional condition. The Taylor expansion around X=x0 only linearizes in X; s(D,x0) remains an arbitrary function of D, while the local regression is linear in D. A least-squares slope of s(D,x0) on D is a projection, not a derivative. The paper should either state the conditional-normality (or equivalent) condition as part of the theorem, or restrict the consistency claim to settings where the local polynomial includes sufficient flexibility in D.","section":"Appendix A.1, The kernel estimator can recover the CME"}],"minor_comments":[{"comment":"The statement that 'mgcv can handle at most two variables, including D' is inaccurate; mgcv supports smooths of more than two covariates via tensor-product terms. Please correct or qualify this claim.","section":"GAM is Not an Appealing Alternative"},{"comment":"The text and captions mention both pointwise and uniform confidence intervals but do not consistently specify which is the shaded region and which is the dashed line; please clarify in each caption.","section":"Figure 2 and Figure 3 captions"},{"comment":"Assumption 1 is labeled 'SUTV A' (a typo for SUTVA), and the Brambor, Clark and Golder reference has a typo in the title ('Mdels'); please proofread.","section":"Assumption 1 and references"},{"comment":"Equation (1) calls ε an idiosyncratic error without stating assumptions; adding E[ε|D,X,Z]=0 or an equivalent condition would make the population regression calculations in Appendix A.1 more self-contained.","section":"Equation (1)"},{"comment":"The claim that the binning estimator 'will be biased but will have less bias than the linear estimator' is asserted without derivation; since the binning estimator is presented as a diagnostic tool, a one-sentence justification or a reference would be helpful.","section":"Binning estimator discussion"}],"recommendation":"major_revision","confidential_remarks":"The paper is a response to a public critique and is partly based on an under-contract book by two of the authors. The core conceptual point (CME vs. CAPE) and the analytic re-derivation of Simonsohn's example are sound, but the appendix's general consistency claim is not. The authors should either prove a correct theorem with explicit conditions or substantially weaken the text's 'most, if not all' claim. I would also note the heavy reliance on self-citations and the under-contract book; this is not a correctness issue, but the manuscript would benefit from a more standard proof or external reference for the local projection claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nI read the HMX response to Simonsohn's critique. The bottom line: the specific rebuttal works. In the quadratic DGP from Simonsohn (2024a), Simonsohn's GAM finite-difference procedure estimates a conditional average partial effect (CAPE) at D=0 and D=0.1, not the conditional marginal effect (CME) θ(x)=E[∂Y/∂D|X=x]. The authors derive the true CME as x−0.5, show the linear interaction model gives a biased slope −0.5+0.8x, and simulate that the HMX kernel estimator tracks x−0.5 with uniform confidence bands. Those derivations are correct as far as I can tell. The estimand taxonomy (CME vs CAPE vs causal moderation) is a genuinely useful clarification, and the point that Simonsohn's benchmark is the wrong estimand is well taken.\n\nThe paper also adds a nice simulated example showing that the GAM procedure yields different CAPEs depending on the chosen D, and it recommends DML/AIPW/PDS-LASSO as more robust alternatives. That's sensible and consistent with recent literature.\n\nNow the soft spots. The appendix (A.1 'Extension to more complex settings') makes a load-bearing overclaim. It says that for any smooth s(D,X), the local linear regression coefficient on D converges to ∂s/∂D at X=x0 'by properties of local polynomial regression.' That's not a property of local polynomial regression. The local-linear slope on D converges to Cov(D, s(D,x0)|X=x0)/Var(D|X=x0), which equals the average derivative only under extra conditions, e.g., D|X normal or s linear in D. The quadratic example works because D|X is normal. The broad claim that the kernel estimator recovers the CME in 'most, if not all' Simonsohn examples is therefore unsupported. The paper should either restrict the claim to DGPs where the local projection condition holds or provide a real proof.\n\nAlso minor: the paper says mgcv can handle at most two variables—that's false. mgcv routinely fits smooths in three or more variables. It's a small error but it weakens the GAM critique.\n\nOverall, this is a serious response worth sending to a referee. The central rebuttal stands, but the appendix generalization needs repair. I'd accept with major revision and ask the authors to fix the consistency claim and tone down the speculation about Simonsohn's other examples.","headline":"The specific rebuttal of Simonsohn's key example is correct, but the appendix overclaims general consistency for the kernel estimator.","tokens_in":12514,"tokens_out":2519,"would_cite":true,"duration_ms":23463,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G08","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Recent critiques of the kernel estimator for interaction effects compared it against the wrong target; once the conditional marginal effect is defined as the estimand, the estimator recovers the true effect in the very simulation said to…","keywords":["conditional marginal effect","interaction models","kernel estimator","local linear regression","estimand","GAM","double/debiased machine learning","causal moderation"],"falsifier":"Simulate a smooth response function $Y = s(D,X)+\\varepsilon$ with large $n$, fix a point $x_0$, and draw $D\\mid X$ from a skewed or otherwise non-normal distribution whose density changes with $X$ so that the local projection of $Y$ onto $D$ does not equal $E[\\partial s/\\partial D \\mid X=x_0]$. If the kernel estimator's estimated coefficient on $D$ fails to converge to the CME as the bandwidth shrinks, the paper's broad consistency claim is falsified.","tokens_in":11512,"feed_emoji":"📈","tokens_out":8915,"duration_ms":80433,"temperature":0.7,"pith_summary":"Recent critiques of the kernel estimator for interaction effects argued that it produces spurious marginal effects when the treatment and moderator are correlated and nonlinearities exist. This response paper argues that the critiques never properly defined their causal target: the quantity the kernel estimator is built to recover is the conditional marginal effect (CME), the expected derivative of the potential outcome with respect to treatment among units with a fixed moderator value. In the critiques' central simulation, the true CME is not zero but $x - 0.5$, and the kernel estimator recovers it, while the linear interaction model is biased. The paper also contends that the alternative the critiques recommend, GAM, estimates a different object (the conditional average partial effect given a treatment level), so its outputs should not be read as marginal effects. The upshot is a call to define estimands explicitly and to adopt doubly robust or machine-learning-based estimators when covariates are high-dimensional.","feed_headline":"Kernel estimator recovers true interaction effect","feed_subtitle":"A response argues the disputed results vanish once the causal estimand is specified correctly.","key_machinery":"The load-bearing object is the conditional marginal effect (CME), defined as $\\theta(x) = E[\\partial Y_i(d)/\\partial d \\mid X_i = x]$, which is the causal effect of the treatment on the outcome among units with moderator value $x$, marginalized over the treatment distribution and other covariates. The estimator that carries the argument is the kernel estimator, a local linear regression of $Y$ on $(1, D, X - x_0, D(X - x_0))$ with kernel weights shrinking the bandwidth. Near $x_0$, the interaction term vanishes and the coefficient on $D$ directly approximates the local average derivative of the outcome in $D$, which is exactly the CME. The distinction between CME and the conditional average partial effect (CAPE) that conditions on $D=d$ is what the paper says the critiques conflate.","core_discovery":"Stated in the paper's own terms: the conditional marginal effect is $\\theta(x) = E[\\partial Y_i(d)/\\partial d \\mid X_i = x]$, and under the unconfoundedness and overlap assumptions it is identified nonparametrically. In the simulation disputed by the critiques, $Y = D^2 - 0.5D + \\varepsilon$ with $(D,X)$ bivariate normal and correlation 0.5, the true CME is $\\theta(x) = x - 0.5$. The kernel estimator, implemented as a local linear regression of $Y$ on $(1, D, X - x_0, D(X - x_0))$, consistently recovers this curve where data are dense; the linear interaction model does not. The paper generalizes this: for any smooth response function $s(D,X)$, the local-linear coefficient on $D$ converges to the CME in regions with abundant data. The paper concludes that the critiques' asserted spuriousness is an artifact of evaluating the estimator against the wrong estimand.","pith_inferences":["If the CME/CAPE distinction becomes standard, many past studies that reported marginal effects from finite differences of predicted outcomes may need to be re-read as answering a conditional question, not an average one.","The paper's general consistency proof relies on the local projection slope equaling the average derivative; this holds in the worked example because $D\\mid X$ is normal. A natural test is to repeat the simulation with skewed or heteroskedastic $D\\mid X$ and see whether the kernel estimate drifts.","The debate suggests a broader lesson for method comparisons: a simulation can only convict an estimator if the target quantity is specified in advance; otherwise the same numbers can support opposite conclusions."],"forward_implications":["Applied researchers studying effect heterogeneity can treat the kernel estimator as a valid tool for the CME when the sample is large enough to provide dense support on the moderator.","Linear interaction models should not be used as the benchmark for judging whether a marginal effect is real, because they are themselves biased under nonlinearity.","GAM-based 'marginal effects' computed by finite differences should be reported as conditional on a chosen treatment level, not as average causal effects.","The recommended workflow—state the estimand, check overlap, use diagnostics, and switch to doubly robust or DML estimators with many covariates—provides a concrete template for interaction research."],"supporting_citations":[],"fun_headline_variants":["Kernel estimator recovers true effect when benchmark is right","HMX critiques fail because they compare to wrong estimand","Kernel method recovers conditional effect under correct estimand","Wrong benchmark, not kernel estimator, explains disputed results","Simonsohn critiques refuted by proper causal estimand"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The general claim that the kernel estimator consistently estimates the CME for any smooth response function rests on the unproven assumption that the local linear regression slope converges to the average derivative $E[\\partial s(D,x_0)/\\partial D]$, which is not implied by local polynomial regression alone and holds in the paper's example because the conditional distribution of $D$ given $X$ is normal.","fun_headline_variants_meta":{"raw":{"variants":["Kernel estimator recovers true effect when benchmark is right","HMX critiques fail because they compare to wrong estimand","Kernel method recovers conditional effect under correct estimand","Wrong benchmark, not kernel estimator, explains disputed results","Simonsohn critiques refuted by proper causal estimand"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000743,"raw_usage":{"total_tokens":3320,"prompt_tokens":955,"completion_tokens":2365,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":571,"completion_tokens_details":{"reasoning_tokens":2285}},"tokens_in":571,"tokens_out":2365,"duration_ms":14996,"temperature":1.0,"reasoning_tokens":2285,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T18:14:58.518832+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a smooth response function $Y = s(D,X)+\\varepsilon$ with large $n$, fix a point $x_0$, and draw $D\\mid X$ from a skewed or otherwise non-normal distribution whose density changes with $X$ so that the local projection of $Y$ onto $D$ does not equal $E[\\partial s/\\partial D \\mid X=x_0]$. If the kernel estimator's estimated coefficient on $D$ fails to converge to the CME as the bandwidth shrinks, the paper's broad consistency claim is falsified.","supporting_citations":[],"review_version":1}