Pith. sign in

REVIEW 3 major objections 3 minor 1 cited by

Difference-in-differences Design with Outcomes Missing Not at Random

T0 review · 3 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Under principal-strata parallel trends and parallel trends of missingness, the average treatment effect for always-respondents is partially identified by trimming the treated-respondent distribution, giving sharp bounds that do not…

desk verdict A genuinely useful partial-identification result for DID with MNAR outcomes, but the paper needs to fix a sign error and, more importantly, close the gap between Theorem 2 and its empirical application's pre-treatment missingness. read the letter →

arxiv 2411.18772 v1 pith:53NRKXTN submitted 2024-11-27 stat.ME econ.EMstat.AP

classification stat.MEecon.EMstat.AP
keywords difference-in-differencesmissingnotatrandomprincipalstratificationLeeboundspaneldataattritionbiaspartialidentificationcausalinference
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that complete-case analysis in difference-in-differences can be biased unless a restrictive parallel-trends assumption holds between respondents and nonrespondents within each treatment arm. The proposed alternative defines latent principal strata by how units would respond under treatment and under control, then imposes parallel trends within each stratum. Using the observed trend in response rates over time, the paper shows how to identify the proportion of always-respondents in the treated group, and then tailors Lee-style trimming bounds to partially identify the treatment effect for that subgroup. A secondary instrumental-variable approach uses a baseline response indicator, such as response to an auxiliary survey question, to point-identify the full ATT under a bias-homogeneity condition. If correct, the method gives DID researchers a way to handle outcomes missing not at random without imputation, covariate weighting, or assuming that treatment selection is independent of the missingness pattern.

What carries the argument

The central object is the principal stratification of the response indicator, classifying each unit as always-respondent, if-treated-respondent, if-control-respondent, or never-respondent based on potential response under both treatment arms. The argument chains two assumptions together: principal-strata parallel trends (within each latent stratum, the no-treatment time trend is the same in treated and control groups) and parallel trends of missingness (under control, the change in response rates from the first wave to the second is the same for treated and control units). These assumptions identify the treated group's always-respondent share, $π_{11}(1)$, from observed response rates alone. The trimming step then replaces the unknown distribution of the always-respondent treated units with the upper or lower extreme of the observed treated-respondent distribution, yielding the partial-identification bounds. The paper's secondary machinery is a 'bespoke instrumental variable' argument, where a baseline response indicator acts as an instrument for post-treatment missingness under a bias-homogeneity condition.

What would settle it

Use three or more pre-treatment waves and compare the response-rate trend of eventual treated and control units during the pre-treatment period when neither group is treated; a placebo difference-in-differences of the response indicator between the two groups over adjacent pre-treatment waves should be zero if the parallel-trends-of-missingness assumption holds. A nonzero placebo estimate, or a visible divergence in the pre-treatment response-rate trajectories, would contradict the identifying assumption. Additionally, one can conduct a placebo outcome test: an outcome observed without missingness that is known to be unaffected by treatment should show no DID effect across the treated and control groups.

Watch

Extended reading notes

Core claim

The central result is Theorem 2: under principal-strata parallel trends and parallel trends of missingness, the average treatment effect for always-respondents (ATT-AR) is partially identified by trimming the observed treated-respondent distribution, with bounds of the form $LB_{\Delta,1} - E[\Delta \mid D=0,R=1] \le ATT \text{-}AR \le UB_{\Delta,1} - E[\Delta \mid D=0,R=1]$ under monotonicity, where the lower and upper bounds are quantile-trimmed means of the before-after outcome change among treated respondents. The paper establishes this by first showing that, under monotonicity plus the parallel-trends assumption on missingness, the proportion of always-respondents in the treated group is identified from observed response rates in the pre- and post-treatment waves. It then applies the trimming logic of Lee and Zhang-Rubin to the treated-respondent distribution to bound the always-respondent group's mean before-after change, and subtracts the observed control mean to obtain the band for ATT-AR. This result accommodates both dependence between treatment selection and principal strata and heterogeneous effects across strata, which the paper shows are the two core obstacles to identifying the full ATT under outcome MNAR.

Load-bearing premise

The load-bearing premise is that the response-rate trend under the control condition is the same for the treated and control groups: in the absence of treatment, the change in survey response from the first wave to the second would have been parallel between the two arms. If treated and control units would have drifted apart in their response rates even without treatment, the estimated proportion of always-respondents is wrong and the trimming bounds are invalid.

Editorial extensions

If this is right

  • Complete-case DID estimators implicitly require that, within each treatment arm, the time trend of the outcome is the same for respondents and nonrespondents; this paper shows that assumption is not justified when treatment selection and missingness patterns are dependent or effects are heterogeneous across strata.
  • With only the observed response rates in two waves plus one pre-treatment wave, the proportion of always-respondents in each arm can be identified under monotonicity or a symmetric response-effect assumption, enabling the trimming bounds.
  • The identified set for ATT-AR is informative whenever the always-respondent share is large; in the paper's reanalysis of the aid-attitudes survey data, most outcomes have always-respondent shares around 0.8 to 0.9 and correspondingly tight bounds.
  • If a baseline binary response indicator (for example, response to an auxiliary pre-treatment question) satisfies relevance, parallel trends, and bias homogeneity, the full ATT can be point-identified without assuming homogeneous effects across principal strata.
  • The approach converts the missingness-rate trajectory itself into identifying information, so the method works for MNAR outcomes without requiring imputation models or covariate-based inverse probability weighting.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • With three or more pre-treatment waves, the parallel-trends-of-missingness assumption becomes testable: one can compare response-rate trends between eventual treatment and control groups in pre-treatment placebo periods, and a visible divergence would flag a violation.
  • The same trimming logic could be extended to multiple post-treatment waves or to covariate-specific trimming, which would likely tighten the bounds when response propensity varies with observed characteristics.
  • The bounds are naturally suited to sensitivity analysis: a researcher can report the range of ATT-AR alongside the estimated always-respondent share, showing whether the original complete-case point estimate falls inside the identified set.
  • Formal inference for the bounds should account for the fact that the trimming proportions are estimated; a bootstrap or resampling procedure over the estimated principal-strata proportions is a natural next step that the paper does not develop.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper studies difference-in-differences designs with outcomes missing not at random. It introduces a principal stratification framework with strata defined by potential post-treatment response status, discusses why complete-case analysis requires strong assumptions, and proposes two identification strategies: an instrumental-variable approach using a baseline response indicator, and a partial-identification approach that tailors Lee-type trimming bounds to the ATT for always-respondents. The main theoretical result, Theorem 2, claims to partially identify ATT-AR under principal-strata parallel trends and a parallel-trends assumption on missingness, with or without a monotonicity assumption. The method is illustrated on a reanalysis of Sexton and Zürcher (2024).

Significance. If Theorem 2 is correct, the paper would provide a valuable partial-identification result for a practically important estimand: the average treatment effect among always-respondents under outcome MNAR, without requiring independence between treatment selection and missingness or homogeneous effects across principal strata. The paper also usefully clarifies the assumptions underlying complete-case analysis in DID settings and highlights the role of principal strata in missing-data problems. The empirical illustration, though explicitly illustrative, shows the method can produce informative bounds in realistic settings. However, the central theorem has a load-bearing gap concerning pre-treatment missingness, and the non-monotonicity version of the theorem is not operational as stated because key proportions are only bounded, not point-identified.

major comments (3)
  1. [3.2, Theorem 2] Theorem 2 defines f_d(y) = Pr(Y_i2 - Y_i1 <= y | D_i = d, R_i2 = 1), but Y_i1 is observed only when R_i1 = 1. The paper initially assumes no missingness in the first wave (Section 2.2), then states in Section 3.2 that the conditioning set omits R_i1 = 1 'for simplicity of notation.' This omission is not harmless: the composition of the observed subsample {D = d, R_i1 = 1, R_i2 = 1} in terms of principal strata differs from that of {D = d, R_i2 = 1} whenever R_i1 depends on the principal stratum, so the trimming proportions in Theorem 2 are not the proportions of always-respondents within the subsample whose difference Y_i2 - Y_i1 is actually observed. This is not a technical edge case: Table 1 reports 69.83% pre-treatment missingness for the treated group in 2016, and the empirical application in Section 4 applies Theorem 2 directly to these data. To make Theorem 2 valid for the motivating data structure, the estimand must be defined conditional on R_i1 = 1 and the principal-strata proportions must be identified within that subpopulation, or an explicit additional assumption must be introduced. Assumption 8 in Appendix B.1 does not fill this gap, because it only asserts a conditional independence for the time trend and does not recover the missing pre-treatment outcomes needed to compute Y_i2 - Y_i1 for units with R_i1 = 0.
  2. [3.2, Theorem 2, part 2] The non-monotonicity version of Theorem 2 is not operational as stated. Without Assumption 5, Proposition 4 identifies Pr(R_i2(0) = 1 | D_i = 1) and Pr(R_i2(1) = 1 | D_i = 0), but it does not point-identify the principal-strata proportions pi_11(1) and pi_11(0). Appendix B.2.2 only provides bounds on these proportions. Nevertheless, the bounds LBy2-y1,0 and UBy2-y1,0 in Theorem 2 are defined using quantiles that depend on pi_11(0)/Pr(R_i2 = 1 | D_i = 0), so the claimed identified set cannot be computed from the identified quantities. The theorem needs either an additional assumption that point-identifies the proportions, or a derivation of outer bounds that accounts for the unidentified pi_11(d) through a further optimization over their feasible ranges.
  3. [3.2, ATT-AR decomposition] The displayed identity for ATT-AR immediately before the trimming-bounds discussion contains a sign error: it reads E[Y_i2 - Y_i1 | D_i = 1, R_i2(1) = 1, R_i2(0) = 1] + E[Y_i2 - Y_i1 | D_i = 0, R_i2(1) = 1, R_i2(0) = 1], but the DID contrast requires a minus sign. Theorem 2 itself uses the correct subtraction, so the error is likely typographical, but the displayed equation is central to the derivation and must be corrected.
minor comments (3)
  1. [3.2, trimming-bounds discussion] The quantile definitions in the introductory paragraph of Section 3.2 are inconsistent with those in Theorem 2. The text says q_low_1(f1) is the pi_10(1)/(pi_11(1)+pi_10(1)) quantile, while the theorem defines q_low_d as the bottom pi_11(d)/Pr(R_i2=1|D_i=d) quantile. For the lower bound on the always-respondent mean, the correct quantity is the bottom pi_11/(pi_11+pi_10) quantile, as in the theorem; the earlier statement should be corrected.
  2. [Appendix B.2] The proof of Theorem 2 is not provided. Appendix B.2 proves Proposition 4 and gives bounds on the principal-strata proportions, but the trimming-bounds argument that combines these results into Theorem 2 is not formally demonstrated. Given that Theorem 2 is the paper's main identification result, a complete proof or a precise reference to a Lee (2009)-type argument should be included.
  3. [Section 2.2] The notation R_i is introduced for the post-treatment response indicator, and later R_i1 and R_i2 are used for pre- and post-treatment indicators. The paper should state explicitly at first use that R_i in Sections 2 and 3.1 refers to R_i2, to avoid ambiguity when pre-treatment missingness is introduced.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's identification assumptions do not encode the target estimand, and the bounds are derived from distinct observed response-rate and outcome-distribution quantities.

full rationale

The paper's central claims are identification theorems, not fitted predictions. Theorem 2 bounds ATT-AR = E[Yi2(1)-Yi2(0)|Di=1,Ri2(1)=Ri2(0)=1] using two distinct ingredients: (i) Lee-type trimming bounds applied to the observed distribution fd(y)=Pr(Yi2-Yi1≤y|Di=d,Ri2=1), and (ii) trimming proportions π11(d)/Pr(Ri2=1|Di=d) identified from response-rate moments under Assumptions 5 and 6 via Proposition 3. The estimand, the observed outcome distribution, and the proportion parameters are logically separate objects; no equation defines the estimand as equal to a fitted quantity or to an assumption. Assumptions 3 and 6 are substantive parallel-trends conditions on potential outcomes and potential response indicators, and neither contains ATT-AR. The trimming technique is explicitly attributed to Lee (2009) and Zhang and Rubin (2003), and the adaptation is transparent rather than a renaming. The paper contains no self-citations, and no uniqueness result is imported from the author's own prior work. The empirical application is explicitly illustrative and does not claim out-of-sample prediction from parameters fitted to the target quantity. The most serious concern raised in review — that fd(y) may not be fully observed when pre-treatment outcomes are missing — is an identification/correctness issue, not a circularity issue; the paper itself flags the Ri1 conditioning simplification in Section 3.2 and offers Assumption 8 as an alternative route. Because no load-bearing step reduces by construction to its own inputs, the circularity score is 0.

Assumptions & free parameters 0 free parameters · 7 assumptions · 0 invented entities

The central claims depend on a set of untestable parallel-trends and monotonicity assumptions on the response process. No free parameters are fitted: the proportions of principal strata are identified nonparametrically from observed response rates under these assumptions. No new entities are introduced beyond the standard principal-strata latent types.

assumptions (7)
  • domain assumption Consistency and no anticipation: Yit = Di Yit(1) + (1-Di) Yit(0), and Yit(1)=Yit(0) for t=1.
    Standard potential-outcomes assumptions underlying the DID framework, invoked in Section 2.2.
  • domain assumption Assumption 1 (Parallel trends): E[Yi2(0) - Yi1 | Di=1] = E[Yi2(0) - Yi1 | Di=0].
    Canonical DID identifying assumption, used in Theorem 1 and in the complete-case analysis result (Proposition 1).
  • domain assumption Assumption 3 (Principal strata parallel trends): E[Yi2(0) - Yi1 | Di=0, Si=s] = E[Yi2(0) - Yi1 | Di=1, Si=s] for all s.
    Core identifying assumption for the partial identification approach in Section 3.2; it replaces canonical parallel trends with within-stratum parallel trends.
  • domain assumption Assumption 6 (Parallel trends of missingness): E[Ri2(0) - Ri1(0) | Di=1] = E[Ri2(0) - Ri1(0) | Di=0].
    Identifies the counterfactual response probability Pr(Ri2(0)=1|Di=1), which is needed to estimate the proportion of always-respondents in the treated group.
  • domain assumption Assumption 5 (Monotonicity): Ri2(1) >= Ri2(0) for all i.
    Eliminates the if-control-respondent stratum, simplifying the trimming bounds; can be replaced by Assumption 7 if not plausible.
  • domain assumption Assumption 7 (Equivalence of ATT and ATC on missingness): E[Ri2(1)-Ri2(0)|Di=1] = E[Ri2(1)-Ri2(0)|Di=0].
    Alternative to monotonicity that identifies counterfactual response probabilities in both treatment arms; needed for the non-monotonicity version of Theorem 2.
  • domain assumption Assumption 4 (IV approach): eR is relevant to missingness, parallel trends of observed outcome across eR groups, and bias homogeneity δd is constant across eR groups.
    The three conditions that make the baseline response indicator a valid instrument for post-treatment missingness in Theorem 1.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Difference-in-differences Design with Outcomes Missing Not at Random." pith.science (2026). https://pith.science/paper/53NRKXTN

@misc{pith2026241118772,
  author       = {Pith},
  title        = {Pith review of: Difference-in-differences Design with Outcomes Missing Not at Random},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/53NRKXTN}},
  note         = {Machine review of arXiv:2411.18772}
}
read the original abstract

This paper addresses one of the most prevalent problems encountered by political scientists working with difference-in-differences (DID) design: missingness in panel data. A common practice for handling missing data, known as complete case analysis, is to drop cases with any missing values over time. A more principled approach involves using nonparametric bounds on causal effects or applying inverse probability weighting based on baseline covariates. Yet, these methods are general remedies that often under-utilize the assumptions already imposed on panel structure for causal identification. In this paper, I outline the pitfalls of complete case analysis and propose an alternative identification strategy based on principal strata. To be specific, I impose parallel trends assumption within each latent group that shares the same missingness pattern (e.g., always-respondents, if-treated-respondents) and leverage missingness rates over time to estimate the proportions of these groups. Building on this, I tailor Lee bounds, a well-known nonparametric bounds under selection bias, to partially identify the causal effect within the DID design. Unlike complete case analysis, the proposed method does not require independence between treatment selection and missingness patterns, nor does it assume homogeneous effects across these patterns.

Figures

Figures reproduced from arXiv: 2411.18772 by the authors.

Figure 1
Figure 1. Illustration of the Parallel Trends Assumption 1 and Assumption 2. (Left): Canonical parallel trends assumption between treated and control groups. Each dot repre￾sents the expected mean outcome for the treated (red) and control (blue) groups. Since these groups are a mix of respondents and nonrespondents, each quantity in post-treatment period is not directly observed. (Middle): Parallel trends assumption between r… view at source ↗
Figure 2
Figure 2. Toy Example of the Principal Strata Parallel Trends Assumption 3. Each dot represents the expected mean outcome for the treated (yellow) and control (green) groups within each principal strata. Filled dots represent the observed quantity and empty dots represent the unobserved quantity. Dashed lines represent the outcome trends under no administration of treatment for the treated group, i.e. line connecting E[Yi1 | … view at source ↗
Figure 3
Figure 3. A DAG example of Randomized Incentive (Re) and Post-treatment Response Indicator (R). Omitted D for simplicity. selection bias due to missingness (also see Tchetgen Tchetgen and Wirth, 2017; Richard￾son and Tchetgen Tchetgen, 2022). Here, I adapt the basic idea of (1) introducing an IV that satisfies a certain exclusion restriction and (2) assuming that the bias is homoge￾neous across the subgroups defined by the IV… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Illustration of the trimming bounds approach for the ATT-AR. The un-shaded area represents the lower and upper bounds of E[Yi2 − Yi1 | Di = 1, Ri2(1) = 1, Ri2(0) = 1]. q high 0 (f0) is the π11(0) π11(0)+π01(0) quantile of f0(·). Consequently, our remaining task involve…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Identification of dynamic treatment effects when treatment histories are partially observed

    econ.EM 2025-01 accept novelty 5.0 of 10

    A robust DID estimator recovers path-dependent treatment effects with partially missing treatment histories whenever any two of outcome, propensity, and missingness models are correct.

Reference graph

Works this paper leans on

5 extracted references · 5 canonical work pages · cited by 1 Pith paper

  1. [1]

    Plugging in Pr(Ri2(0) = 1 | Di = 1) = Pr(Ri = 1 | Di =

    − Pr(Ri(0) = 1 | Di = 1). Plugging in Pr(Ri2(0) = 1 | Di = 1) = Pr(Ri = 1 | Di =

  2. [2]

    Assumption 9 (Baseline response indicators as IV for time trend)

    − Pr(Ri1 = 1| Di = 0) + Pr(Ri1 = 1| Di = 1), we have Pr(Ri(1) = 1| Di = 0) = Pr(Ri = 1| Di = 1)− Pr(Ri1 = 1| Di = 1) + Pr(Ri1 = 1| Di = 0) Accordingly, we have    πr1,r0(0) ∈ [0, 1] ∀r0, r1 = 0, 1 π11(0) +π01(0) +π10(0) +π00(0) = 1 π10(0) = Pr(Ri2 = 1| Di = 0)− π11(0) π01(0) = Pr(Ri2(1) = 1| Di = 0)− π11(0) Analogous to the previous c...

  3. [3]

    Parallel difference in trendsThe difference in time trends of the outcome between respondents and nonrespondents is parallel for two auxiliary variables. E[Yi2 − Yi1 | Di = d, eR(1) i = 1, Ri1 = 1]− E[Yi2 − Yi1 | Di = d, eR(1) i = 0, Ri1 = 1] = E[Yi2 − Yi1 | Di = d, eR(2) i = 1, Ri1 = 1]− E[Yi2 − Yi1 | Di = d, eR(2) i = 0, Ri1 = 1] for d = 0, 1

  4. [4]

    Bias homogeneity: The bias due to missingness in the time trend of the outcome is 43 homogeneous across the subgroups defined by the baseline response indicators, and is same as the marginalized version. E[Yi2 − Yi1 | Di = d, Ri1 = 1, Ri2 = 1]− E[Yi2 − Yi1 | Di = d, Ri1 = 1, Ri2 = 0] = E[Yi2 − Yi1 | Di = d, eR(1) i = r, Ri1 = 1, Ri2 = 1]− E[Yi2 − Yi1 | Di...

  5. [5]

    Relevance to missingness: The baseline response indicators are relevant to the miss- ingness of post-treatment outcome. Pr(Ri2 = 0| Di = d, eR(1) i = 0, Ri1 = 1)̸= Pr(Ri2 = 0| Di = d, eR(1) i = 1, Ri1 = 1) Pr(Ri2 = 0| Di = d, eR(2) i = 0, Ri1 = 1)̸= Pr(Ri2 = 0| Di = d, eR(2) i = 1, Ri1 = 1) for d = 0, 1. Lemma 2 Under Assumption 9, we have E[Yi2 − Yi1 | D...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.