REVIEW 1 major objections 2 minor 11 references
Stacked Triple Differences
T0 review · 1 major / 2 minor · reviewed 2026-05-21 · grok-4.3
Pith's one-line read A regression with fully saturated fixed effects on stacked triple-difference data identifies a cell-size-weighted average of conditional average treatment effects.
desk verdict Stacked DDD adapts the stacked DiD framework to triple differences with a claimed identification result for a cell-size weighted average under saturated fixed effects, but the cross-stack fixed effects specification needs explicit verification to confirm the result holds. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Self-contained stacks consisting of four cells over an event window—treated and clean comparison cohorts, each with treatment-eligible and treatment-ineligible units—appended into a unified dataset for estimation.
What would settle it
If the coefficient from the saturated regression on the stacked data does not match the cell-size-weighted average of the stack-level effects computed separately, the identification claim would be falsified.
Extended reading notes
Core claim
At each post-treatment event-time, a linear regression with fully saturated fixed-effects applied to the stacked dataset identifies a strictly positive, cell-size-weighted average of stack-level conditional average treatment effects, with stack weights proportional to stack-level cell sizes. Building on this characterization, alternative weighting schemes recover causal estimands with clear interpretations.
Load-bearing premise
Within each stack the clean comparison cohort satisfies the parallel changes-in-trends condition relative to the treated cohort, and appending stacks does not introduce cross-stack contamination in the fixed effects.
Editorial extensions
If this is right
- Alternative weighting schemes can recover causal estimands with clear interpretations.
- The approach relies on pairwise rather than global parallel changes-in-trends conditions.
- Researchers obtain direct control over the comparison group for each treated unit and over the aggregation weights.
- Stacked DDD trades some efficiency for regression-based transparency relative to GMM and imputation frameworks.
Reading between the lines
- The stacking logic may extend to other multi-way difference designs with eligibility rules.
- Monte Carlo studies could test how the method performs when stack-level parallel trends hold only approximately.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces stacked triple differences (stacked DDD) for staggered adoption settings. Self-contained stacks are formed, each consisting of four cells (treated and clean comparison cohorts, each split into treatment-eligible and ineligible units) over an event window. Appending stacks into one dataset, the central claim is that OLS with fully saturated fixed effects identifies a strictly positive, cell-size-weighted average of stack-level conditional average treatment effects (CATEs), with weights proportional to stack cell sizes. Alternative weighting schemes are outlined, and two empirical illustrations are presented where results differ from standard procedures.
Significance. If the identification result is correct, the approach supplies a regression-transparent estimator that grants explicit control over comparison cohorts and aggregation weights while relying only on pairwise parallel changes-in-trends within each stack. It complements GMM and imputation frameworks by emphasizing interpretability and direct weight selection over efficiency. The empirical examples illustrate that the choice of estimator can materially affect quantitative conclusions.
major comments (1)
- [Identification proof] Identification section (around the proof of the weighted-average result): the claim that fully saturated fixed effects recover the cell-size-weighted average of stack-level CATEs assumes stacks remain self-contained after appending. If any unit or time period appears in more than one stack, the global unit, time, or eligibility fixed effects are estimated from observations belonging to different stacks. This creates potential cross-stack leakage that mixes the within-stack contrasts, so the coefficient on the post-treatment interaction is no longer guaranteed to equal the claimed weighted average. The manuscript should clarify whether fixed effects are interacted with stack indicators or provide an explicit argument that the result survives without such interactions.
minor comments (2)
- [Abstract] The abstract states that two empirical illustrations are provided but gives no indication of the policy settings or the direction of the differences relative to existing estimators; a brief clause summarizing the applications would improve reader orientation.
- [Notation and setup] Notation for the event-time-specific estimands and the cell-size weights could be introduced with an explicit equation early in the identification section to make the subsequent weighting discussion easier to follow.
Simulated Author's Rebuttal
We thank the referee for the careful reading and constructive feedback on our manuscript introducing stacked triple differences. We address the major comment below.
read point-by-point responses
-
Referee: [Identification proof] Identification section (around the proof of the weighted-average result): the claim that fully saturated fixed effects recover the cell-size-weighted average of stack-level CATEs assumes stacks remain self-contained after appending. If any unit or time period appears in more than one stack, the global unit, time, or eligibility fixed effects are estimated from observations belonging to different stacks. This creates potential cross-stack leakage that mixes the within-stack contrasts, so the coefficient on the post-treatment interaction is no longer guaranteed to equal the claimed weighted average. The manuscript should clarify whether fixed effects are interacted with stack indicators or provide an explicit argument that the result survives without such interactions.
Authors: We appreciate the referee highlighting this important implementation detail. The manuscript defines stacks as self-contained to ensure that identification relies on pairwise within-stack changes in trends. To eliminate any possibility of cross-stack leakage when units or periods are shared across stacks, we will revise the identification section to specify that the saturated fixed effects (for units, time periods, and eligibility) are interacted with stack indicators. This guarantees that all contrasts remain strictly within each stack, so the OLS coefficient on the post-treatment interaction continues to recover the cell-size-weighted average of the stack-level CATEs as claimed. We will also add a brief discussion of this choice and its implications for the estimator. revision: yes
Circularity Check
No circularity in the identification derivation
full rationale
The paper derives the identification result from the explicit construction of self-contained stacks (each with treated/clean cohorts and eligible/ineligible units) followed by appending and applying a linear regression with fully saturated fixed effects. This produces the cell-size-weighted average of stack-level CATEs as a consequence of the within-stack contrasts and the fixed-effects absorption, rather than by defining the target quantity in terms of the fitted coefficients themselves or by renaming a fitted input. No self-citation chain, ansatz smuggling, or uniqueness theorem imported from prior author work is invoked to force the central claim. The derivation is therefore self-contained against the stated assumptions about pairwise parallel trends within stacks.
Assumptions & free parameters
assumptions (2)
- domain assumption Within each stack, the clean comparison cohort satisfies parallel changes-in-trends with the treated cohort for the eligible and ineligible subgroups.
- domain assumption Appending independent stacks does not induce bias through the shared fixed effects structure.
Cite this review
Pith. "Pith review of Stacked Triple Differences." pith.science (2026). https://pith.science/paper/26RXUZKT
@misc{pith2026260422982,
author = {Pith},
title = {Pith review of: Stacked Triple Differences},
year = {2026},
howpublished = {\url{https://pith.science/paper/26RXUZKT}},
note = {Machine review of arXiv:2604.22982}
}
read the original abstract
Triple differences (DDD) is a workhorse quasi-experimental design in applied economics. But, under staggered adoption, its conventional three-way fixed-effects (3WFE) implementation inherits the interpretation issues now well understood in the difference-in-differences literature. I introduce stacked DDD. I extend the stacked difference-in-differences approach to the DDD setting by creating self-contained stacks, each consisting of four cells over an event window: treated and clean comparison cohorts, each with treatment-eligible and treatment-ineligible units. Appending these stacks yields a unified dataset for estimating treatment effects. I prove that, at each post-treatment event-time, a linear regression with fully saturated fixed-effects applied to the stacked dataset identifies a strictly positive, cell-size-weighted average of stack-level conditional average treatment effects, with stack weights proportional to stack-level cell sizes. Building on this characterization, I outline alternative weighting schemes that recover causal estimands with clear interpretations. Stacked DDD complements recent GMM and imputation-based frameworks by trading efficiency for regression-based transparency, pairwise (rather than global) parallel changes-in-trends, and direct control over both the comparison group for each treated unit and the aggregation weights. I provide two empirical illustrations where stacked DDD yields substantially different quantitative conclusions compared to existing procedures.
Lean theorems connected to this paper
-
IndisputableMonolith/Cost/FunctionalEquation.leanwashburn_uniqueness_aczel unclear?
unclearRelation between the paper passage and the cited Recognition theorem.
a linear regression with fully saturated fixed-effects applied to the stacked dataset identifies a strictly positive, cell-size-weighted average of stack-level conditional average treatment effects
-
IndisputableMonolith/Foundation/RealityFromDistinction.leanreality_from_one_distinction unclear?
unclearRelation between the paper passage and the cited Recognition theorem.
pairwise (rather than global) parallel changes-in-trends
What do these tags mean?
- matches
- The paper's claim is directly supported by a theorem in the formal canon.
- supports
- The theorem supports part of the paper's argument, but the paper may add assumptions or extra steps.
- extends
- The paper goes beyond the formal theorem; the theorem is a base layer rather than the whole result.
- uses
- The paper appears to rely on the theorem as machinery.
- contradicts
- The paper's claim conflicts with a theorem or certificate in the canon.
- unclear
- Pith found a possible connection, but the passage is too broad, indirect, or ambiguous to say the theorem truly supports the claim.
Reference graph
Works this paper leans on
-
[1]
Abadie, A., Athey, S., Imbens, G. W., and Wooldridge, J. M. (2023). When Should You Adjust Standard Errors for Clustering?Quarterly Journal of Economics, 138(1):1–35. (Cited on page 3.) Borusyak, K., Jaravel, X., and Spiess, J. (2024). Revisiting Event-Study Designs: Robust and Efficient Estima- tion.Review of Economic Studies, 91(6):3253–3285. (Cited on ...
work page 2023
-
[2]
(Cited on page 5.) Sant’Anna, P. H. C. and Zhao, J. (2020). Doubly robust difference-in-differences estimators.Journal of Econo- metrics, 219(1):101–122. (Cited on page 3.) Shastry, G. K. and Tortorice, D. L. (2025). Effective health aid: Evidence from gavi’s vaccine program.Amer- ican Economic Journal: Economic Policy, 17(1):540–74. (Cited on pages 4, 31...
-
[3]
Computing the FWL residual within each stack. By Step 1, the OLS estimator forτ e reduces to a bivariate FWL regression (B.12)bτ e = P g∈Gtrg(e) P i∈Sg eRe(i, g+e, g) g∆Y i,g+e,g P g∈Gtrg(e) P i∈Sg eRe(i, g+e, g) 2 , where the time summation collapses entirely tot=g+ebecause eRe(i, t, g) = 0fort̸=g+e. I now derive the FWL residual eRe and the resulting we...
work page 2021
-
[4]
APPENDIXE. THREE-WAYFIXED-EFFECTS INEVENT-STUDYDESIGNS The conventional approach to DDD estimation is the three-way fixed effects (3WFE) regression (E.1)Y i,t =α i +γ t +δ Si,t +θD i,t +ϵ i,t , whereα i are unit fixed effects,γ t are time fixed effects,δ Si,t are group-by-time fixed effects, andD i,t = 1{t≥S i}Qi is the treatment indicator. Despite its si...
work page 2021
-
[5]
The unit-level time mean is Di,· = (T−g+ 1)/T. The group-time mean Dg,t equals the fraction of group-g units that are eligible, Dg,t =n g,1/ng,· fort≥g(since all eligible group-gunits are treated) and Dg,t = 0for t < g. Thus Dg,· = ((T−g+ 1)/T)(n g,1/ng,·). STACKED TRIPLE DIFFERENCES 23 Now consider a second cohortg ′ > gand a not-yet-treated eligible uni...
work page 2020
-
[6]
The stacked estimator avoids both pathologies. The stacked estimator avoids negative weights because each stack produces a singleATT(g, t)estimate, and the aggregation weightsω g are chosen by the researcher to be non-negative. It avoids forbidden comparisons because each stack restricts the comparison group to units withS i =g c > g+K, ensuring no compar...
work page 2023
-
[7]
Property (i), own-period weights sum to one
I now prove each property by summing (E.12) overg. Property (i), own-period weights sum to one. Fixℓ=eand sum (E.12) overg∈ G trg X g∈Gtrg ωe g,e =e ⊤ e X t E[ ¨Ri,t ¨R⊤ i,t] −1 E ¨Ri,g+e X g∈Gtrg Rg,e(i, g+e) . SinceP g Rg,e(i, t) =R e(i, t)for all(i, t)(summing over cohorts recovers the aggregate event-time indicator), this becomes the coefficien...
work page 2021
-
[8]
Thereforeµ e = ATTe.■ Theorem E.5 clarifies the conditions under which the 3WFE event-study regression is valid. The specifi- cation (E.6) recovers interpretable causal parameters only under the joint restrictions of DDD-PCT, treatment effect homogeneity across cohorts, and no anticipation. In practice, these conditions are rarely satisfied simul- taneous...
work page 2021
Show all 11 references
-
[9]
The event- time indicatorR 0(i,2015)lights up for cohort 2015 but not for cohort 2013; conversely,R 2(i,2015)lights up for cohort 2013 but not
2015
-
[10]
forbidden comparisons
After removing unit and time fixed effects, these indicators retain residual correlation because the relative-time composition of the sample changes across time. The group-by-time fixed effectsδ Si,t absorb some of this variation but do not eliminate it, since the within-group...
2021
-
[11]
forbidden comparison
The group mean for groupgis DSi=g,· = 1 ngT X j:Sj=g TX t=1 Dj,t =p Q=1|S=g T−g+ 1 T . Substituting into (E.19) for a unit withS i =g,Q i = 1: eDi,t =1{t≥g} − T−g+ 1 T −p Q=1|S=g 1{t≥g}+p Q=1|S=g T−g+ 1 T = (1−p Q=1|S=g) 1{t≥g} − T−g+ 1 T .(E.20) The factor(1−p Q=1|S=g)is the ...
2021
Reviewed May 21, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.