REVIEW 2 major objections 4 minor 12 references
A fixed-effects causal forest estimates covariate-conditional group-time treatment effects in staggered adoption without the forbidden-comparison bias that plagues two-way fixed effects and pooled forests.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 12:08 UTC pith:FKOCZ6GW
load-bearing objection A genuinely new estimator and an honest method paper, but the Medicaid poverty gradient is softer than the framing suggests — the pre-trend defense does not rule out post-2013 differential trends correlated with poverty. the 2 major comments →
A Fixed-Effects Causal Forest for Staggered Adoption, with an Application to Medicaid Expansion
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that treating staggered adoption as a collection of clean two-group, two-period blocks — each cohort compared only to units not yet treated at the relevant time, anchored at its pre-adoption period — and fitting honest causal trees with node-level two-way fixed-effects residualization within each block yields consistent estimates of τ_{g,t}(x), the cohort-, period-, and covariate-conditional effect. The residualization absorbs unit and period effects inside the node; splitting on covariate-driven treatment-effect heterogeneity recovers the conditional surface; aggregation with group-time weights recovers the scalar. Because already-treated units never enter a com
What carries the argument
The central mechanism is the block decomposition combined with node-level fixed-effects residualization. For each cohort g and post-period t, the estimator forms a two-period panel of cohort-g units and their not-yet-treated controls at periods g−1 and t, then grows honest causal trees on that block; inside each node it residualizes outcome and treatment on unit and period fixed effects by iterative two-way demeaning before splitting on treatment-effect heterogeneity. This transforms the identification problem so that, within a block, the base-period differencing makes treatment unconfounded given covariates under the assumed parallel trends. Aggregating block surfaces with group-time weight
Load-bearing premise
For the Medicaid headline, the load-bearing premise is that, conditional on the 2013 poverty rate, income, and region, counties in expansion states would have followed the same uninsured-rate path as counties in non-expansion states after 2013 had they not expanded.
What would settle it
Re-estimate the Medicaid surface on the 2011–2023 trimmed panel and on a state-level-trends specification using only never-treated counties as controls; if the poverty-rate correlation (−0.55) weakens or changes sign when the recession-era divergent states are excluded or allowed separate trends, the conditional gradient is an artifact of trend divergence rather than treatment heterogeneity.
If this is right
- Applied researchers using staggered difference-in-differences can estimate for whom a policy works without modeling propensity scores or outcome regressions.
- Under timing-heterogeneous effects, the block construction should replace pooled causal forest implementations: the latter inherit the same negative-weight contamination as two-way fixed effects.
- For Medicaid expansion, the conditional surface says coverage gains were largest in poorer and lower-income counties, so aggregate effects understate the program's redistributive content.
- The scalar estimate can serve as a check: because it equals the standard group-time aggregate by construction, a researcher can validate the forest implementation against the published aggregate.
- For designs with late cohorts and thin controls, overlap-failing blocks are reported as dropped rather than extrapolated, which exposes where the data cannot identify the conditional effect.
Where Pith is reading between the lines
- If the conditional surface is accurate, cost-benefit targeting could shift: expansion resources aimed at high-poverty counties are where the coverage return is largest — a targeting rule the paper does not itself derive.
- The block-wise fixed-effects residualization could be paired with doubly-robust nuisance adjustment to protect against within-block residual confounding; the paper treats these as alternatives rather than complements.
- A natural stress test: re-estimate the Medicaid surface using only counties in states that expanded in 2014 and separately in later cohorts; if the poverty gradient is driven by the 2014 wave alone, the heterogeneity claim is narrower than it appears.
- The estimator's dependence on large blocks suggests it will perform best in county- or firm-level panels; for country-level panels with few cohorts, the conditional surface should be interpreted cautiously.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a fixed-effects causal forest for staggered difference-in-differences, targeting the covariate-conditional group-time treatment effect τ_{g,t}(x) of Callaway and Sant'Anna. The estimator forms clean (g,t) comparison blocks, excludes already-treated units, residualizes outcomes and treatment on unit and period fixed effects within each tree node, and aggregates block-level surfaces with CS weights. The authors are explicit that the estimand is not new (MLDID; Imai et al.) and that the scalar ATT equals the CS aggregate by construction; the claimed contribution is the conditional surface. Monte Carlo evidence shows that under timing-heterogeneous effects the staggered forest avoids the large forbidden-comparison bias of TWFE and a pooled causal forest, and the Medicaid application reports a -2.25 pp average effect plus a poverty gradient in the conditional surface.
Significance. If the method works as claimed, it is a useful complement to existing staggered-DiD tools: it provides a nonparametric conditional object under parallel trends rather than unconfoundedness, and the clean-block construction is a principled fix for forbidden-comparison bias in forest-based DiD. The paper ships a reproducible archive and is transparent about the estimand's provenance. The Monte Carlo design is informative, especially the DGP with cohort-varying timing effects, where the staggered forest's advantage is large and the honesty sweep in Table 2 supports the claim that the block construction, not the honesty knob, drives the gain. However, the central consistency proof is only sketched and the implementation switches honesty off in small blocks, and the empirical headline — the poverty gradient in Medicaid effects — is not subjected to the same sensitivity analysis as the average effect, despite visible pre-trend divergence in the paper's own placebo plot.
major comments (2)
- [§3.3, Proposition 1] The consistency claim is not established for the estimator as implemented. The proof asserts that node-level two-way demeaning is algebraically base-period differencing and then imports Wager and Athey (2018) Theorems 1 and 3, but those theorems require honest trees, whereas §3.3 and Table 2 state that the implementation splits non-honestly in small blocks. If consistency is meant asymptotically as block size grows so all blocks eventually meet the honesty threshold, that argument should be made explicit. As written, the object that is proved consistent differs from the object run in the Monte Carlo and applications. The sketch also needs to verify the Wager–Athey conditions on the residualized block (e.g., Lipschitzness of the transformed regression) rather than asserting them. Please provide a formal reduction or clearly state the result as conditional on honesty and quantify finite-bl
- [§5.3, Table 7, Figure 1] The empirical headline — poorer counties gained more (Table 6, corr(τ̂, poverty)=−0.55) — is identified only under conditional parallel trends. The deep pre-period placebos in Figure 1 (up to −1.98 pp at exposures ≤ −9) show that expansion and non-expansion states were on divergent paths through the Great Recession. The paper's response — that these years never enter an estimation block because every cohort's reference period is g−1, and that trimming to 2011–2023 leaves the ATT unchanged — does not address the possibility that the same trend divergence continues after 2013 and is correlated with poverty. The Rambachan–Roth analysis in §5.3 is applied to the overall ATT only, not to the conditional gradient. A robustness check that bounds the gradient under plausible post-treatment differential trends, or a placebo analysis interacted with poverty, is needed before the poverty gradient c
minor comments (4)
- [§4, Table 1] In DGP (a) the staggered forest reports coverage 0.85 for the overall ATT, below nominal, while CS reports 0.96. The text attributes this to bootstrap calibration on small two-period blocks, but the paper's broader claim of 'valid coverage' should be qualified, and the result should be discussed in the main text rather than only in a footnote.
- [§4.1, footnote 2] The MLDID benchmark is a Python reimplementation that omits the minimax reweighting of the original MLDID, and the original package is not run because of installation errors. The reported MLDID bias (+0.381) and coverage (0.33) in DGP (c) may therefore not reflect the actual method. Since the head-to-head is not the paper's core claim, please soften the comparative statements or make the reimplementation's limitations more prominent.
- [§3.4] The pointwise bootstrap bands for the conditional surface are not uniform. The paper acknowledges this, but given that the conditional surface is the main empirical object, a sentence explaining the practical implications for interpreting the reported gradient would be helpful.
- [General] Typos and formatting: 'causalfePython' in the Data and code section is missing a space; the JEL code formatting and the figure axis labels could be cleaned up. These do not affect substance.
Circularity Check
No significant circularity: the only by-construction identity (scalar = Callaway–Sant'Anna aggregate) is explicitly acknowledged and non-load-bearing; the conditional-surface claim is independently benchmarked.
full rationale
The derivation chain is self-contained. The identifying equation (5) is the standard conditional parallel-trends DiD result, and the estimator is an honest causal forest applied to base-period-differenced outcomes within each clean (g,t) block; consistency is imported from Wager and Athey under explicit block regularity, not from a circular definition. The key evaluation of the conditional surface uses Monte Carlo CATE-RMSE comparisons that are independent of the estimator's own construction, and the method is benchmarked against external estimators (TWFE, CS, MLDID). The one by-construction relationship is the overall scalar's numerical equality to the Callaway–Sant'Anna aggregate; the paper states this explicitly ('the overall (covariate-averaged) effect our estimator reports is numerically identical, by construction, to the CS aggregate... we therefore lead with the conditional surface and treat the scalar as a check against CS rather than as a contribution'), so the scalar is not repackaged as an independent prediction. There are no self-citations; the load-bearing references (Wager–Athey, Callaway–Sant'Anna, Kattenberg et al., Gavrilova et al., Rambachan–Roth) are external, and the replication archive makes the Monte Carlo and applications reproducible. The skeptical reader's pre-trend concern is an external-validity threat to the empirical gradient, not a circularity of the estimator's derivation. No claimed derivation reduces to its own input.
Axiom & Free-Parameter Ledger
free parameters (2)
- Honesty block-size threshold =
adaptive; 'a few hundred observations' (Table 2 sweeps 50/100/200/400)
- Forest hyperparameters (tree count, mtry, min node size, bootstrap draws) =
not reported in text; package defaults in causalfe
axioms (6)
- domain assumption Assumption 1: conditional parallel trends for not-yet-treated units, E[Y_t(∞)-Y_{g-1}(∞)|G=g,X=x] equals the same for C_{g,t}.
- domain assumption Assumption 2: no anticipation.
- domain assumption Assumption 3: within-block overlap for every (g,t) and covariate region used by the forest.
- domain assumption Assumption 4: block regularity — continuous covariates on bounded support, Lipschitz conditional means in X, bounded second moments, independent units — plus the Wager-Athey regularity conditions.
- standard math External theory: Wager-Athey causal-forest consistency/normality and the Goodman-Bacon decomposition.
- domain assumption State-level treatment assignment and state-clustered block bootstrap for the Medicaid application.
read the original abstract
Difference-in-differences with staggered adoption identifies group-time average treatment effects ATT(g,t) by comparing each cohort to units not yet treated, which avoids the "forbidden comparisons" that bias two-way fixed-effects estimators when effects are heterogeneous. This paper studies the covariate-conditional version of that object, tau_{g,t}(x), and estimates it with a fixed-effects causal forest. Within each (g,t) comparison block, the outcome and treatment are residualized on unit and period fixed effects inside each tree node, and honest causal trees split on treatment-effect heterogeneity in the covariates. The estimand is not new: Hatamyar, Kreif, Rocha and Huber (2023) introduced it using a doubly-robust R-learner, and Imai, Qin and Yanagi (2023) study it for a single continuous covariate. What we add is a different way to estimate it. Where those methods remove confounding by modeling nuisance functions, we remove it by differencing out unit and period effects within each tree node, following the fixed-effects residualization of Kattenberg, Scheer and Thiel (2023) and Gavrilova, Langorgen and Zoutman (2025) and carrying it into the Callaway-Sant'Anna group-time structure. In Monte Carlo experiments the estimator is the only forest-based method that stays unbiased and correctly covered for the overall effect under staggered timing with cohort-varying effects; two-way fixed effects and a pooled causal forest inherit large forbidden-comparison bias. We apply the method to the Callaway-Sant'Anna minimum-wage panel as a validation and to the staggered county-level rollout of the ACA Medicaid expansion, where it recovers an average 2.25 percentage-point fall in the uninsured rate and a conditional surface on which poorer and lower-income counties gained substantially more coverage -- heterogeneity measured along socioeconomic covariates that are not lags of the outcome.
Figures
Reference graph
Works this paper leans on
-
[1]
Generalized random forests
Susan Athey, Julie Tibshirani, and Stefan Wager. Generalized random forests. The Annals of Statistics, 47 0 (2): 0 1148--1178, 2019
2019
-
[2]
Brantly Callaway and Pedro H. C. Sant'Anna. Difference-in-differences with multiple time periods. Journal of Econometrics, 225 0 (2): 0 200--230, 2021
2021
-
[3]
Two-way fixed effects estimators with heterogeneous treatment effects
Cl \'e ment de Chaisemartin and Xavier D'Haultf uille. Two-way fixed effects estimators with heterogeneous treatment effects. American Economic Review, 110 0 (9): 0 2964--2996, 2020
2020
-
[4]
Evelina Gavrilova, Audun Lang rgen, and Floris T. Zoutman. Difference-in-difference causal forests, with an application to payroll tax incidence in norway. Journal of Applied Econometrics, 40 0 (7): 0 727--740, 2025
2025
-
[5]
Difference-in-differences with variation in treatment timing
Andrew Goodman-Bacon. Difference-in-differences with variation in treatment timing. Journal of Econometrics, 225 0 (2): 0 254--277, 2021
2021
-
[6]
Machine learning for staggered difference-in-differences and dynamic treatment effect heterogeneity
Julia Hatamyar, Noemi Kreif, Rudi Rocha, and Martin Huber. Machine learning for staggered difference-in-differences and dynamic treatment effect heterogeneity. arXiv preprint arXiv:2310.11962, 2023
Pith/arXiv arXiv 2023
-
[7]
Shunsuke Imai, Lei Qin, and Takahide Yanagi. Doubly robust uniform confidence bands for group-time conditional average treatment effects in difference-in-differences. arXiv preprint arXiv:2305.02185, 2023
Pith/arXiv arXiv 2023
-
[8]
Causal forests with fixed effects for treatment effect heterogeneity in difference-in-differences
Mark Kattenberg, Bas Scheer, and Jurre Thiel. Causal forests with fixed effects for treatment effect heterogeneity in difference-in-differences. CPB Discussion Paper 452, CPB Netherlands Bureau for Economic Policy Analysis, 2023
2023
-
[9]
Estimating individual treatment effects in difference-in-differences frameworks
Xiaomeng Lu, Xinkun Nie, and Stefan Wager. Estimating individual treatment effects in difference-in-differences frameworks. arXiv preprint arXiv:1902.09166, 2019
Pith/arXiv arXiv 1902
-
[10]
A more credible approach to parallel trends
Ashesh Rambachan and Jonathan Roth. A more credible approach to parallel trends. Review of Economic Studies, 90 0 (5): 0 2555--2591, 2023
2023
-
[11]
Estimating dynamic treatment effects in event studies with heterogeneous treatment effects
Liyang Sun and Sarah Abraham. Estimating dynamic treatment effects in event studies with heterogeneous treatment effects. Journal of Econometrics, 225 0 (2): 0 175--199, 2021
2021
-
[12]
Estimation and inference of heterogeneous treatment effects using random forests
Stefan Wager and Susan Athey. Estimation and inference of heterogeneous treatment effects using random forests. Journal of the American Statistical Association, 113 0 (523): 0 1228--1242, 2018
2018
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.