REVIEW 3 major objections 3 minor 1 cited by
Difference-in-differences Design with Outcomes Missing Not at Random
T0 review · 3 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Under principal-strata parallel trends and parallel trends of missingness, the average treatment effect for always-respondents is partially identified by trimming the treated-respondent distribution, giving sharp bounds that do not…
desk verdict A genuinely useful partial-identification result for DID with MNAR outcomes, but the paper needs to fix a sign error and, more importantly, close the gap between Theorem 2 and its empirical application's pre-treatment missingness. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the principal stratification of the response indicator, classifying each unit as always-respondent, if-treated-respondent, if-control-respondent, or never-respondent based on potential response under both treatment arms. The argument chains two assumptions together: principal-strata parallel trends (within each latent stratum, the no-treatment time trend is the same in treated and control groups) and parallel trends of missingness (under control, the change in response rates from the first wave to the second is the same for treated and control units). These assumptions identify the treated group's always-respondent share, $π_{11}(1)$, from observed response rates alone. The trimming step then replaces the unknown distribution of the always-respondent treated units with the upper or lower extreme of the observed treated-respondent distribution, yielding the partial-identification bounds. The paper's secondary machinery is a 'bespoke instrumental variable' argument, where a baseline response indicator acts as an instrument for post-treatment missingness under a bias-homogeneity condition.
What would settle it
Use three or more pre-treatment waves and compare the response-rate trend of eventual treated and control units during the pre-treatment period when neither group is treated; a placebo difference-in-differences of the response indicator between the two groups over adjacent pre-treatment waves should be zero if the parallel-trends-of-missingness assumption holds. A nonzero placebo estimate, or a visible divergence in the pre-treatment response-rate trajectories, would contradict the identifying assumption. Additionally, one can conduct a placebo outcome test: an outcome observed without missingness that is known to be unaffected by treatment should show no DID effect across the treated and control groups.
Extended reading notes
Core claim
The central result is Theorem 2: under principal-strata parallel trends and parallel trends of missingness, the average treatment effect for always-respondents (ATT-AR) is partially identified by trimming the observed treated-respondent distribution, with bounds of the form $LB_{\Delta,1} - E[\Delta \mid D=0,R=1] \le ATT \text{-}AR \le UB_{\Delta,1} - E[\Delta \mid D=0,R=1]$ under monotonicity, where the lower and upper bounds are quantile-trimmed means of the before-after outcome change among treated respondents. The paper establishes this by first showing that, under monotonicity plus the parallel-trends assumption on missingness, the proportion of always-respondents in the treated group is identified from observed response rates in the pre- and post-treatment waves. It then applies the trimming logic of Lee and Zhang-Rubin to the treated-respondent distribution to bound the always-respondent group's mean before-after change, and subtracts the observed control mean to obtain the band for ATT-AR. This result accommodates both dependence between treatment selection and principal strata and heterogeneous effects across strata, which the paper shows are the two core obstacles to identifying the full ATT under outcome MNAR.
Load-bearing premise
The load-bearing premise is that the response-rate trend under the control condition is the same for the treated and control groups: in the absence of treatment, the change in survey response from the first wave to the second would have been parallel between the two arms. If treated and control units would have drifted apart in their response rates even without treatment, the estimated proportion of always-respondents is wrong and the trimming bounds are invalid.
Editorial extensions
If this is right
- Complete-case DID estimators implicitly require that, within each treatment arm, the time trend of the outcome is the same for respondents and nonrespondents; this paper shows that assumption is not justified when treatment selection and missingness patterns are dependent or effects are heterogeneous across strata.
- With only the observed response rates in two waves plus one pre-treatment wave, the proportion of always-respondents in each arm can be identified under monotonicity or a symmetric response-effect assumption, enabling the trimming bounds.
- The identified set for ATT-AR is informative whenever the always-respondent share is large; in the paper's reanalysis of the aid-attitudes survey data, most outcomes have always-respondent shares around 0.8 to 0.9 and correspondingly tight bounds.
- If a baseline binary response indicator (for example, response to an auxiliary pre-treatment question) satisfies relevance, parallel trends, and bias homogeneity, the full ATT can be point-identified without assuming homogeneous effects across principal strata.
- The approach converts the missingness-rate trajectory itself into identifying information, so the method works for MNAR outcomes without requiring imputation models or covariate-based inverse probability weighting.
Reading between the lines
- With three or more pre-treatment waves, the parallel-trends-of-missingness assumption becomes testable: one can compare response-rate trends between eventual treatment and control groups in pre-treatment placebo periods, and a visible divergence would flag a violation.
- The same trimming logic could be extended to multiple post-treatment waves or to covariate-specific trimming, which would likely tighten the bounds when response propensity varies with observed characteristics.
- The bounds are naturally suited to sensitivity analysis: a researcher can report the range of ATT-AR alongside the estimated always-respondent share, showing whether the original complete-case point estimate falls inside the identified set.
- Formal inference for the bounds should account for the fact that the trimming proportions are estimated; a bootstrap or resampling procedure over the estimated principal-strata proportions is a natural next step that the paper does not develop.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies difference-in-differences designs with outcomes missing not at random. It introduces a principal stratification framework with strata defined by potential post-treatment response status, discusses why complete-case analysis requires strong assumptions, and proposes two identification strategies: an instrumental-variable approach using a baseline response indicator, and a partial-identification approach that tailors Lee-type trimming bounds to the ATT for always-respondents. The main theoretical result, Theorem 2, claims to partially identify ATT-AR under principal-strata parallel trends and a parallel-trends assumption on missingness, with or without a monotonicity assumption. The method is illustrated on a reanalysis of Sexton and Zürcher (2024).
Significance. If Theorem 2 is correct, the paper would provide a valuable partial-identification result for a practically important estimand: the average treatment effect among always-respondents under outcome MNAR, without requiring independence between treatment selection and missingness or homogeneous effects across principal strata. The paper also usefully clarifies the assumptions underlying complete-case analysis in DID settings and highlights the role of principal strata in missing-data problems. The empirical illustration, though explicitly illustrative, shows the method can produce informative bounds in realistic settings. However, the central theorem has a load-bearing gap concerning pre-treatment missingness, and the non-monotonicity version of the theorem is not operational as stated because key proportions are only bounded, not point-identified.
major comments (3)
- [3.2, Theorem 2] Theorem 2 defines f_d(y) = Pr(Y_i2 - Y_i1 <= y | D_i = d, R_i2 = 1), but Y_i1 is observed only when R_i1 = 1. The paper initially assumes no missingness in the first wave (Section 2.2), then states in Section 3.2 that the conditioning set omits R_i1 = 1 'for simplicity of notation.' This omission is not harmless: the composition of the observed subsample {D = d, R_i1 = 1, R_i2 = 1} in terms of principal strata differs from that of {D = d, R_i2 = 1} whenever R_i1 depends on the principal stratum, so the trimming proportions in Theorem 2 are not the proportions of always-respondents within the subsample whose difference Y_i2 - Y_i1 is actually observed. This is not a technical edge case: Table 1 reports 69.83% pre-treatment missingness for the treated group in 2016, and the empirical application in Section 4 applies Theorem 2 directly to these data. To make Theorem 2 valid for the motivating data structure, the estimand must be defined conditional on R_i1 = 1 and the principal-strata proportions must be identified within that subpopulation, or an explicit additional assumption must be introduced. Assumption 8 in Appendix B.1 does not fill this gap, because it only asserts a conditional independence for the time trend and does not recover the missing pre-treatment outcomes needed to compute Y_i2 - Y_i1 for units with R_i1 = 0.
- [3.2, Theorem 2, part 2] The non-monotonicity version of Theorem 2 is not operational as stated. Without Assumption 5, Proposition 4 identifies Pr(R_i2(0) = 1 | D_i = 1) and Pr(R_i2(1) = 1 | D_i = 0), but it does not point-identify the principal-strata proportions pi_11(1) and pi_11(0). Appendix B.2.2 only provides bounds on these proportions. Nevertheless, the bounds LBy2-y1,0 and UBy2-y1,0 in Theorem 2 are defined using quantiles that depend on pi_11(0)/Pr(R_i2 = 1 | D_i = 0), so the claimed identified set cannot be computed from the identified quantities. The theorem needs either an additional assumption that point-identifies the proportions, or a derivation of outer bounds that accounts for the unidentified pi_11(d) through a further optimization over their feasible ranges.
- [3.2, ATT-AR decomposition] The displayed identity for ATT-AR immediately before the trimming-bounds discussion contains a sign error: it reads E[Y_i2 - Y_i1 | D_i = 1, R_i2(1) = 1, R_i2(0) = 1] + E[Y_i2 - Y_i1 | D_i = 0, R_i2(1) = 1, R_i2(0) = 1], but the DID contrast requires a minus sign. Theorem 2 itself uses the correct subtraction, so the error is likely typographical, but the displayed equation is central to the derivation and must be corrected.
minor comments (3)
- [3.2, trimming-bounds discussion] The quantile definitions in the introductory paragraph of Section 3.2 are inconsistent with those in Theorem 2. The text says q_low_1(f1) is the pi_10(1)/(pi_11(1)+pi_10(1)) quantile, while the theorem defines q_low_d as the bottom pi_11(d)/Pr(R_i2=1|D_i=d) quantile. For the lower bound on the always-respondent mean, the correct quantity is the bottom pi_11/(pi_11+pi_10) quantile, as in the theorem; the earlier statement should be corrected.
- [Appendix B.2] The proof of Theorem 2 is not provided. Appendix B.2 proves Proposition 4 and gives bounds on the principal-strata proportions, but the trimming-bounds argument that combines these results into Theorem 2 is not formally demonstrated. Given that Theorem 2 is the paper's main identification result, a complete proof or a precise reference to a Lee (2009)-type argument should be included.
- [Section 2.2] The notation R_i is introduced for the post-treatment response indicator, and later R_i1 and R_i2 are used for pre- and post-treatment indicators. The paper should state explicitly at first use that R_i in Sections 2 and 3.1 refers to R_i2, to avoid ambiguity when pre-treatment missingness is introduced.
Circularity Check
No significant circularity: the paper's identification assumptions do not encode the target estimand, and the bounds are derived from distinct observed response-rate and outcome-distribution quantities.
full rationale
The paper's central claims are identification theorems, not fitted predictions. Theorem 2 bounds ATT-AR = E[Yi2(1)-Yi2(0)|Di=1,Ri2(1)=Ri2(0)=1] using two distinct ingredients: (i) Lee-type trimming bounds applied to the observed distribution fd(y)=Pr(Yi2-Yi1≤y|Di=d,Ri2=1), and (ii) trimming proportions π11(d)/Pr(Ri2=1|Di=d) identified from response-rate moments under Assumptions 5 and 6 via Proposition 3. The estimand, the observed outcome distribution, and the proportion parameters are logically separate objects; no equation defines the estimand as equal to a fitted quantity or to an assumption. Assumptions 3 and 6 are substantive parallel-trends conditions on potential outcomes and potential response indicators, and neither contains ATT-AR. The trimming technique is explicitly attributed to Lee (2009) and Zhang and Rubin (2003), and the adaptation is transparent rather than a renaming. The paper contains no self-citations, and no uniqueness result is imported from the author's own prior work. The empirical application is explicitly illustrative and does not claim out-of-sample prediction from parameters fitted to the target quantity. The most serious concern raised in review — that fd(y) may not be fully observed when pre-treatment outcomes are missing — is an identification/correctness issue, not a circularity issue; the paper itself flags the Ri1 conditioning simplification in Section 3.2 and offers Assumption 8 as an alternative route. Because no load-bearing step reduces by construction to its own inputs, the circularity score is 0.
Assumptions & free parameters
assumptions (7)
- domain assumption Consistency and no anticipation: Yit = Di Yit(1) + (1-Di) Yit(0), and Yit(1)=Yit(0) for t=1.
- domain assumption Assumption 1 (Parallel trends): E[Yi2(0) - Yi1 | Di=1] = E[Yi2(0) - Yi1 | Di=0].
- domain assumption Assumption 3 (Principal strata parallel trends): E[Yi2(0) - Yi1 | Di=0, Si=s] = E[Yi2(0) - Yi1 | Di=1, Si=s] for all s.
- domain assumption Assumption 6 (Parallel trends of missingness): E[Ri2(0) - Ri1(0) | Di=1] = E[Ri2(0) - Ri1(0) | Di=0].
- domain assumption Assumption 5 (Monotonicity): Ri2(1) >= Ri2(0) for all i.
- domain assumption Assumption 7 (Equivalence of ATT and ATC on missingness): E[Ri2(1)-Ri2(0)|Di=1] = E[Ri2(1)-Ri2(0)|Di=0].
- domain assumption Assumption 4 (IV approach): eR is relevant to missingness, parallel trends of observed outcome across eR groups, and bias homogeneity δd is constant across eR groups.
Cite this review
Pith. "Pith review of Difference-in-differences Design with Outcomes Missing Not at Random." pith.science (2026). https://pith.science/paper/53NRKXTN
@misc{pith2026241118772,
author = {Pith},
title = {Pith review of: Difference-in-differences Design with Outcomes Missing Not at Random},
year = {2026},
howpublished = {\url{https://pith.science/paper/53NRKXTN}},
note = {Machine review of arXiv:2411.18772}
}
read the original abstract
This paper addresses one of the most prevalent problems encountered by political scientists working with difference-in-differences (DID) design: missingness in panel data. A common practice for handling missing data, known as complete case analysis, is to drop cases with any missing values over time. A more principled approach involves using nonparametric bounds on causal effects or applying inverse probability weighting based on baseline covariates. Yet, these methods are general remedies that often under-utilize the assumptions already imposed on panel structure for causal identification. In this paper, I outline the pitfalls of complete case analysis and propose an alternative identification strategy based on principal strata. To be specific, I impose parallel trends assumption within each latent group that shares the same missingness pattern (e.g., always-respondents, if-treated-respondents) and leverage missingness rates over time to estimate the proportions of these groups. Building on this, I tailor Lee bounds, a well-known nonparametric bounds under selection bias, to partially identify the causal effect within the DID design. Unlike complete case analysis, the proposed method does not require independence between treatment selection and missingness patterns, nor does it assume homogeneous effects across these patterns.
Figures
Figures from the paper (1 more)
Forward citations
Cited by 1 Pith paper
-
Identification of dynamic treatment effects when treatment histories are partially observed
A robust DID estimator recovers path-dependent treatment effects with partially missing treatment histories whenever any two of outcome, propensity, and missingness models are correct.
Reference graph
Works this paper leans on
-
[1]
Plugging in Pr(Ri2(0) = 1 | Di = 1) = Pr(Ri = 1 | Di =
− Pr(Ri(0) = 1 | Di = 1). Plugging in Pr(Ri2(0) = 1 | Di = 1) = Pr(Ri = 1 | Di =
-
[2]
Assumption 9 (Baseline response indicators as IV for time trend)
− Pr(Ri1 = 1| Di = 0) + Pr(Ri1 = 1| Di = 1), we have Pr(Ri(1) = 1| Di = 0) = Pr(Ri = 1| Di = 1)− Pr(Ri1 = 1| Di = 1) + Pr(Ri1 = 1| Di = 0) Accordingly, we have πr1,r0(0) ∈ [0, 1] ∀r0, r1 = 0, 1 π11(0) +π01(0) +π10(0) +π00(0) = 1 π10(0) = Pr(Ri2 = 1| Di = 0)− π11(0) π01(0) = Pr(Ri2(1) = 1| Di = 0)− π11(0) Analogous to the previous c...
-
[3]
Parallel difference in trendsThe difference in time trends of the outcome between respondents and nonrespondents is parallel for two auxiliary variables. E[Yi2 − Yi1 | Di = d, eR(1) i = 1, Ri1 = 1]− E[Yi2 − Yi1 | Di = d, eR(1) i = 0, Ri1 = 1] = E[Yi2 − Yi1 | Di = d, eR(2) i = 1, Ri1 = 1]− E[Yi2 − Yi1 | Di = d, eR(2) i = 0, Ri1 = 1] for d = 0, 1
-
[4]
Bias homogeneity: The bias due to missingness in the time trend of the outcome is 43 homogeneous across the subgroups defined by the baseline response indicators, and is same as the marginalized version. E[Yi2 − Yi1 | Di = d, Ri1 = 1, Ri2 = 1]− E[Yi2 − Yi1 | Di = d, Ri1 = 1, Ri2 = 0] = E[Yi2 − Yi1 | Di = d, eR(1) i = r, Ri1 = 1, Ri2 = 1]− E[Yi2 − Yi1 | Di...
-
[5]
Relevance to missingness: The baseline response indicators are relevant to the miss- ingness of post-treatment outcome. Pr(Ri2 = 0| Di = d, eR(1) i = 0, Ri1 = 1)̸= Pr(Ri2 = 0| Di = d, eR(1) i = 1, Ri1 = 1) Pr(Ri2 = 0| Di = d, eR(2) i = 0, Ri1 = 1)̸= Pr(Ri2 = 0| Di = d, eR(2) i = 1, Ri1 = 1) for d = 0, 1. Lemma 2 Under Assumption 9, we have E[Yi2 − Yi1 | D...
work page 2016
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.