REVIEW 2 major objections 4 minor 1 cited by
Better Understanding Triple Differences Estimators
T0 review · 2 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Once covariates matter, triple-differences estimation requires three difference-in-differences contrasts, not two, and the paper supplies doubly robust estimators that remain valid in staggered designs.
desk verdict Solid DDD theory with a real new result, but the pre-trend diagnostic in Remark 4.4 does not test the identifying assumption and should be fixed before publication. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine is Assumption DDD-CPT (triple-differences conditional parallel trends): for a treated cohort $g$, every available comparison cohort $g_c$, and each post period $t$, the difference in outcome evolution between eligible ($Q=1$) and ineligible ($Q=0$) units must coincide across cohorts after conditioning on covariates $X$. The assumption deliberately permits group-specific and eligibility-specific trend violations as long as they are stable across cohorts, which is what makes DDD more credible than plain DiD. On top of it, Theorem 4.1 builds the doubly robust DDD estimand: a weighted sum of three doubly robust DiD components that pit treated units ($S=g$, $Q=1$) against each of the three untreated cells ($S=g$, $Q=0$), ($S=g_c$, $Q=1$), and ($S=g_c$, $Q=0$), with covariates averaged over the treated group's own distribution. A GMM step then converts the resulting over-identification into efficiency, weighting cohort-specific estimates by the inverse variance-covariance matrix of their re-centered influence functions.
What would settle it
Estimate $\text{ATT}(g,t)$ from the same data using two different not-yet-treated comparison cohorts, for instance never-treated units and units in a cohort treated later, and compare the answers: DDD-CPT implies both identify the same quantity, so a material divergence is direct evidence that the assumption fails. A controlled version is to simulate outcomes in which the eligible–ineligible trend gap differs across cohorts even after conditioning on $X$; the paper's estimators should then drift away from the true effect, and the paper's own DGP 4 marks the boundary where all working models are misspecified and even the doubly robust DDD estimator is biased.
Extended reading notes
Core claim
The paper's central claim is that group-time average treatment effects $\text{ATT}(g,t)$ in triple-differences designs are identified under DDD-CPT — the condition that the eligible-versus-ineligible difference in outcome trends is the same across treatment cohorts conditional on covariates — and that three distinct estimands, regression adjustment, inverse probability weighting, and a doubly robust combination, all recover $\text{ATT}(g,t)$ for any not-yet-treated comparison cohort (Theorem 4.1). Because many comparison cohorts are valid, the design is over-identified, and the paper constructs an optimally weighted GMM estimator from re-centered influence functions that no weighted average of cohort-specific estimators can beat asymptotically (Theorem 4.2). The paper also claims that the familiar equivalence 'DDD equals the difference of two DiDs' survives only in the no-covariate, two-period case; with covariates one needs three DiD terms, and standard three-way fixed-effects regressions, as well as pooled not-yet-treated comparisons in staggered designs, are generally biased.
Load-bearing premise
The load-bearing premise is that the gap between eligible and ineligible units' outcome trends is exactly the same for the treated cohort and every comparison cohort once covariates are accounted for; if this double-differenced parallel-trends condition fails, all the proposed estimators are inconsistent, and no amount of pre-treatment data can prove that it holds.
Editorial extensions
If this is right
- DDD studies that currently report three-way fixed-effects coefficients, or differences of two DiD estimates, should switch to the paper's regression adjustment, inverse probability weighting, or doubly robust estimators whenever covariates help justify the design; the simulations show the standard estimators can be badly biased while the DR DDD estimator stays centered.
- In staggered DDD designs, pooling all not-yet-treated units into one comparison group is not generally valid; using each not-yet-treated cohort separately and combining the estimates with optimal GMM weights recovers the ATT and shortens confidence intervals, with the never-treated-only estimator showing roughly 50% wider intervals in the simulations.
- Event-study analyses remain available: the paper provides event-study estimators and their asymptotic distributions, so dynamic policy effects and pre-treatment placebo checks can still be constructed.
- Consistency is multiply robust: as long as each of the three DiD components has either a correctly specified outcome model or a correctly specified generalized propensity-score model, the DR DDD estimator is consistent — eight working-model combinations in all.
- The three empirical applications show the corrections bite: the three-way fixed-effects effect on low-carbon patents becomes statistically insignificant when estimated with the proposed DDD approach, and the genetically-modified-crop yield result survives dropping never-treated countries only when the DDD estimators are used.
Reading between the lines
- The over-identification the paper exposes suggests a cheap specification test: since every comparison cohort $g_c$ must deliver the same $\text{ATT}(g,t)$ under DDD-CPT, disagreement across cohort-specific estimates is a direct post-treatment probe of the identifying assumption, complementing pre-trend event-study plots.
- The pooling bias the paper traces to composition differences implies a practical diagnostic: when the share of eligible ($Q=1$) units is similar across enabling groups, the pooled not-yet-treated shortcut may be nearly unbiased, whereas when shares differ sharply, the per-cohort estimators are the safer choice.
- The documented precision gains — three-way fixed-effects estimation would need up to roughly 54% more observations to match the DR DDD precision in one application — suggest that DDD studies with limited samples should prefer designs with many not-yet-treated cohorts and use the GMM combination.
- The efficient influence function derived for the two-period case is a stepping stone toward semiparametric efficiency bounds for staggered DDD designs, a direction the paper itself flags for future work.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops identification, estimation, and inference tools for triple-differences (DDD) designs in panel data with covariates and staggered treatment adoption. The target is the group-time average treatment effect ATT(g,t) among eligible units in an enabling group g. Under random sampling, strong overlap, no anticipation, and a conditional parallel trends assumption DDD-CPT that equalizes double-differenced outcome changes across eligible and ineligible units in treated and comparison groups, Theorem 4.1 shows ATT(g,t) equals regression adjustment, inverse probability weighting, and doubly robust estimands for any comparison group that enables treatment after t. Theorem 4.2 establishes consistency and asymptotic normality for the doubly robust estimator and shows that an optimally weighted GMM combination across comparison groups is weakly more precise than any fixed weighted average. The paper also argues, mainly through simulations and informal reasoning, that three-way fixed-effects regressions and pooled not-yet-treated DDD estimators are generally biased. Three empirical applications illustrate the proposed estimators.
Significance. If the results hold, the paper fills a real gap in the DDD literature by providing covariate-adjusted DDD estimators with double robustness, multiple comparison groups, event-study aggregations, and inference, analogous to Callaway and Sant'Anna (2021) for DiD. The identification theorems are derived under explicit assumptions, and the paper provides detailed proofs in the supplemental appendix, extensive Monte Carlo evidence, three empirical applications, and a companion R package. The main substantive weakness is the proposed pre-trend diagnostic: Remark 4.4 recommends event-study checks that, as stated, cannot assess the key identifying assumption DDD-CPT. The negative claims about standard 3WFE and pooled not-yet-treated estimators are also stated more strongly than what is formally proved.
major comments (2)
- [Section 2.2 / Remark 4.4, Eq. (4.7)] The pre-trend diagnostic recommended in Remark 4.4 is invalid as a test of Assumption DDD-CPT. DDD-CPT is stated only for periods t >= g, but for event time e < 0 the estimand ATT^dr,gc(g,g+e) uses periods t < g and is a sum of one-period changes over periods t+1, ..., g-1, none of which satisfy t >= g. Consequently, DDD-CPT imposes no restriction on these pre-treatment coefficients: nonzero pre-trends are fully compatible with DDD-CPT, and zero pre-trends are not implied by it. The event-study statements in Sections 6.2 and 6.3 that interpret 'no very serious violation of pre-treatment trends' as supporting DDD-CPT are therefore not justified and should be revised or removed.
- [Section 3.1, Section 3.2, Abstract] The abstract and Section 3 claim that common DDD implementations are 'generally invalid' when covariates enter the parallel trends assumption or when treatment adoption is staggered. These claims are supported by the Monte Carlo designs in Figures 1-4 and by informal reasoning in Remark 4.2, but not by a formal theorem characterizing the bias of three-way fixed-effects estimators or pooled not-yet-treated estimators. As written, the claims are stronger than what is proved. The authors should either provide formal counterexamples or explicit conditions under which these estimators fail, or qualify the claims as demonstrations in specific DGPs rather than general invalidity results.
minor comments (4)
- [Equation (4.13)] There is a typographical error in the displayed formula for P_n(G=g | G+e in [1,T]): a stray 'L' appears in the denominator and should be removed.
- [Figures 5-7 notes] The notation 'yES dr,gc(e)' and 'yES dr,gmm(e)' is inconsistent with the main text; using \widehat{ES}_{dr,gc}(e) and \widehat{ES}_{dr,gmm}(e) would improve readability.
- [Section 5.2 and Table OA-2] It is worth noting explicitly in the main text that the pooled not-yet-treated estimator is unbiased for ATT(2,3) and ATT(3,3) in the simulation because by period 3 there is only one available comparison group; this helps readers understand why only ATT(2,2) demonstrates the bias.
- [Remark 4.6] The term 're-centered influence function' is used in the main text before being formally defined only in a later remark; a one-sentence definition at first use would help the reader.
Circularity Check
No significant circularity: identification and estimation results are derived from explicitly stated causal assumptions with self-contained proofs.
full rationale
The paper's central theoretical contribution is Theorem 4.1, which states that under Assumptions S, SO, NA, and DDD-CPT, ATT(g,t) equals the RA, IPW, and doubly robust DDD estimands. Assumption DDD-CPT is formulated entirely in terms of untreated potential outcomes and does not reference ATT(g,t), the estimated weights, or the outcome-regression/propensity-score working models. The proof in Lemma A.1 uses DDD-CPT and NA to express the conditional ATT as a double difference of conditional means, and the proof of Theorem 4.1 then shows by iterated expectations and simple algebra that the RA, IPW, and DR estimands all equal that same expression. No fitted parameter is renamed as a prediction; the Monte Carlo DGPs are constructed to satisfy the identifying assumptions, which is standard simulation practice and does not feed back into the theoretical derivation. The paper's citations to Callaway and Sant'Anna (2021) and Sant'Anna and Zhao (2020) are to published, peer-reviewed results and are used as building blocks for asymptotic representations and as methodological analogs, not as an unverified self-citation chain that forces the main conclusion. The supplemental appendix contains self-contained proofs of the main theorems. One passage deserves comment: Remark 4.4 recommends using pre-treatment event-study coefficients to assess the plausibility of DDD-CPT, even though DDD-CPT is only stated for post-treatment periods t >= g; this is a validity concern about the diagnostic rather than a circularity, because it does not make any derived estimand equivalent to its inputs by construction. The paper also honestly reports in DGP 4 that the proposed estimator is biased when all working models are misspecified, further indicating that the results are not circularly guaranteed. Overall, no load-bearing step reduces, by the paper's own equations or by self-citation, to its own inputs.
Assumptions & free parameters
assumptions (6)
- domain assumption Assumption DDD-CPT (conditional parallel trends for DDD)
- domain assumption Assumption SO (strong overlap)
- domain assumption Assumption NA (no anticipation)
- standard math Assumption S (random sampling)
- domain assumption Existence of never-enabled units
- standard math Fixed-T, large-n asymptotics
Cite this review
Pith. "Pith review of Better Understanding Triple Differences Estimators." pith.science (2026). https://pith.science/paper/ASFRO626
@misc{pith2026250509942,
author = {Pith},
title = {Pith review of: Better Understanding Triple Differences Estimators},
year = {2026},
howpublished = {\url{https://pith.science/paper/ASFRO626}},
note = {Machine review of arXiv:2505.09942}
}
read the original abstract
Triple Differences (DDD) designs are widely used in empirical work to relax parallel trends assumptions in Difference-in-Differences (DiD) settings. This paper highlights that common DDD implementations -- such as taking the difference between two DiDs or applying three-way fixed effects regressions -- are generally invalid when identification requires conditioning on covariates. In staggered adoption settings, the common DiD practice of pooling all not-yet-treated units as a comparison group can introduce additional bias, even when covariates are not required for identification. These insights challenge conventional empirical strategies and underscore the need for estimators tailored specifically to DDD structures. We develop regression adjustment, inverse probability weighting, and doubly robust estimators that remain valid under covariate-adjusted DDD parallel trends. For staggered designs, we demonstrate how to effectively utilize multiple comparison groups to obtain more informative inferences. Simulations and three empirical applications highlight bias reductions and precision gains relative to standard approaches. A companion R package is available.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
Beyond Parallel Trends in Staggered Difference-in-Differences: Identification under Higher-Order Parallelism
The paper establishes point identification of cohort-specific and aggregate treatment effects in staggered DiD under a hierarchy of higher-order parallel trends conditions, with an aggregation theorem for mixed-order ...
Reference graph
Works this paper leans on
-
[1]
Difference-in-differences with multiple time periods,
Callaway, Brantly and Pedro HC Sant’Anna , “Difference-in-differences with multiple time periods,” Journal of econometrics , 2021, 225 (2), 200–230. Chen, Xiaohong, Han Hong, and Alessandro Tarozzi , “Semiparametric efficiency in GMM models with auxiliary data,” Annals of Statistics , 2008, 36 (2), 808 –
work page 2021
-
[539]
Semiparametric efficiency bounds,
Newey, Whitney K., “Semiparametric efficiency bounds,” Journal of Applied Econometrics, 1990, 5 (2), 99–135. and Daniel McFadden, “Large sample estimation and hypothesis testing,” in “Handbook of Econometrics,” Vol. 4, Amsterdam: North-Holland: Elsevier, 1994, chapter 36, pp. 2111–2245. Sant’Anna, Pedro H.C. and Jun Zhao, “Doubly Robust Difference-in-Diff...
work page 1990
-
[843]
Cited by: 119; All Open Access, Bronze Open Access, Green Open Access. Hahn, Jinyong, “On the Role of the Propensity Score in Efficient Semiparametric Estimation of Average Treatment Effects,” Econometrica, 1998, 66 (2), 315–331. Kang, Joseph D. Y. and Joseph L. Schafer, “Demystifying Double Robustness: A Comparison of Alternative Strategies for Estimatin...
work page 1998
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.