{"id":"9545e21e-a615-41ba-8119-8be136d29090","arxiv_id":"2411.18772","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A DID estimator with bounding and IV strategies that use principal strata for missingness and parallel trends in response rates to identify ATT or ATT among always-respondents without assuming homogeneous effects across missingness patterns.","lead":"This paper develops new ways to estimate difference-in-differences effects when outcome data are missing for some panel units, using latent response groups and trends in missingness rates. Missing survey responses are common in political science panels, and the proposed bounds and instrumental-variable estimators promise to correct bias under weaker assumptions than complete-case analysis.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 2's trimming bounds implicitly require fully observed pre-treatment outcomes, yet the empirical application has extensive pre-treatment missingness; the distributions and trimming proportions in the theorem are not identified under the stated assumptions.","rationale":"The reader's CONDITIONAL verdict is appropriate: the paper's core trimming argument is standard and the algebra of Proposition 3 is correct when pre-treatment outcomes are fully observed, so the theoretical contribution is not fatally flawed. However, the reader's weakest_assumption focuses on Assumption 6 (parallel trends of missingness), which is indeed needed to identify the trimming proportions, whereas the more immediately load-bearing gap for the central claim is the paper's treatment of pre-treatment missingness. The manuscript explicitly claims to relax the no-missingness-in-first-wave assumption in Section 3, but Theorem 2 and Proposition 3 are stated without conditioning on R_i1=1, and all estimands f_d(y) are defined on Y2-Y1, which is only observed when Y1 is observed. The motivating applications in Table 1 show substantial pre-treatment missingness, so this is not an abstract concern: the proposed method is being demonstrated on data where the theorem's inputs are not identified under the stated assumptions. This is a scope/assumption gap rather than a mathematical contradiction, so it does not warrant rejection; it strengthens the case for a CONDITIONAL verdict requiring the author to either restrict Theorem 2 to fully observed Y1 or add and formalize an explicit assumption on pre-treatment response (e.g., an exclusion restriction like Assumption 8 extended to the trimming estimand, or MCAR for R1). The proposed simulation would settle whether the claim as written is valid when pre-treatment missingness is present, and would also give a concrete check for the required additional condition.","tokens_in":28896,"tokens_out":7288,"duration_ms":65297,"concrete_test":"Simulate a two-period DGP with treatment D, latent principal strata S, monotone response R2(1) ≥ R2(0), and pre-treatment missingness R1 that depends on both S and D (e.g., P(R1=1 | S, D) varies across strata). Generate Y1 only for units with R1=1, and Y2 with stratum-specific trends and effects. Compute the claimed bounds from Theorem 2 using only the complete-case differences Y2-Y1 among units with R2=1 (ignoring R1 in the conditioning set) and the proportions from Proposition 3. Repeat many times and record whether the true ATT-AR falls within the nominal bounds. If coverage is materially below 100% for configurations satisfying Assumptions 3, 5, and 6, the theorem requires an explicit condition such as R_i1=1 for all units or an additional ignorability assumption for pre-treatment response.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central identification result, Theorem 2, defines f_d(y) as Pr(Y_i2 - Y_i1 ≤ y | D_i=d, R_i2=1) and uses trimming proportions such as π_10(1)/(π_11(1)+π_10(1)) derived from the full treated population. If Y_i1 is missing for some units with R_i2=1, the difference Y_i2 - Y_i1 is not observed for those units, so f_d(y) is not identified without additional assumptions. The paper's Section 3.2 says 'we are interested in the ATT for the population who responded to the survey at time 1, while omitting R_i1=1 in the conditioning set for simplicity of notation,' but no formal argument shows that this omission is harmless. Under the stated Assumptions 3 and 5-6, the composition of the observed subsample {D=1, R_i1=1, R_i2=1} in terms of principal strata differs from that of {D=1, R_i2=1} whenever R_i1 depends on the principal stratum, so the identified trimming proportion is not the proportion of if-treated-respondents within the subsample whose Δ is observed. This is not a purely technical edge case: Table 1 reports pre-treatment missingness of 69.83% for the treated group in 2016, and the empirical application in Section 4 applies Theorem 2 directly to those data. The same gap affects the control-side bound E[Δ | D=0, R_2=1], which is treated as fully observed but is only identified among units with R_i1=1. Thus, as stated, Theorem 2 does not cover the paper's motivating data structure unless an additional assumption on pre-treatment missingness is made explicit.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies difference-in-differences designs with outcomes missing not at random. It introduces a principal stratification framework with strata defined by potential post-treatment response status, discusses why complete-case analysis requires strong assumptions, and proposes two identification strategies: an instrumental-variable approach using a baseline response indicator, and a partial-identification approach that tailors Lee-type trimming bounds to the ATT for always-respondents. The main theoretical result, Theorem 2, claims to partially identify ATT-AR under principal-strata parallel trends and a parallel-trends assumption on missingness, with or without a monotonicity assumption. The method is illustrated on a reanalysis of Sexton and Zürcher (2024).","tokens_in":29239,"tokens_out":10420,"duration_ms":89590,"significance":"If Theorem 2 is correct, the paper would provide a valuable partial-identification result for a practically important estimand: the average treatment effect among always-respondents under outcome MNAR, without requiring independence between treatment selection and missingness or homogeneous effects across principal strata. The paper also usefully clarifies the assumptions underlying complete-case analysis in DID settings and highlights the role of principal strata in missing-data problems. The empirical illustration, though explicitly illustrative, shows the method can produce informative bounds in realistic settings. However, the central theorem has a load-bearing gap concerning pre-treatment missingness, and the non-monotonicity version of the theorem is not operational as stated because key proportions are only bounded, not point-identified.","major_comments":[{"comment":"Theorem 2 defines f_d(y) = Pr(Y_i2 - Y_i1 <= y | D_i = d, R_i2 = 1), but Y_i1 is observed only when R_i1 = 1. The paper initially assumes no missingness in the first wave (Section 2.2), then states in Section 3.2 that the conditioning set omits R_i1 = 1 'for simplicity of notation.' This omission is not harmless: the composition of the observed subsample {D = d, R_i1 = 1, R_i2 = 1} in terms of principal strata differs from that of {D = d, R_i2 = 1} whenever R_i1 depends on the principal stratum, so the trimming proportions in Theorem 2 are not the proportions of always-respondents within the subsample whose difference Y_i2 - Y_i1 is actually observed. This is not a technical edge case: Table 1 reports 69.83% pre-treatment missingness for the treated group in 2016, and the empirical application in Section 4 applies Theorem 2 directly to these data. To make Theorem 2 valid for the motivating data structure, the estimand must be defined conditional on R_i1 = 1 and the principal-strata proportions must be identified within that subpopulation, or an explicit additional assumption must be introduced. Assumption 8 in Appendix B.1 does not fill this gap, because it only asserts a conditional independence for the time trend and does not recover the missing pre-treatment outcomes needed to compute Y_i2 - Y_i1 for units with R_i1 = 0.","section":"3.2, Theorem 2"},{"comment":"The non-monotonicity version of Theorem 2 is not operational as stated. Without Assumption 5, Proposition 4 identifies Pr(R_i2(0) = 1 | D_i = 1) and Pr(R_i2(1) = 1 | D_i = 0), but it does not point-identify the principal-strata proportions pi_11(1) and pi_11(0). Appendix B.2.2 only provides bounds on these proportions. Nevertheless, the bounds LBy2-y1,0 and UBy2-y1,0 in Theorem 2 are defined using quantiles that depend on pi_11(0)/Pr(R_i2 = 1 | D_i = 0), so the claimed identified set cannot be computed from the identified quantities. The theorem needs either an additional assumption that point-identifies the proportions, or a derivation of outer bounds that accounts for the unidentified pi_11(d) through a further optimization over their feasible ranges.","section":"3.2, Theorem 2, part 2"},{"comment":"The displayed identity for ATT-AR immediately before the trimming-bounds discussion contains a sign error: it reads E[Y_i2 - Y_i1 | D_i = 1, R_i2(1) = 1, R_i2(0) = 1] + E[Y_i2 - Y_i1 | D_i = 0, R_i2(1) = 1, R_i2(0) = 1], but the DID contrast requires a minus sign. Theorem 2 itself uses the correct subtraction, so the error is likely typographical, but the displayed equation is central to the derivation and must be corrected.","section":"3.2, ATT-AR decomposition"}],"minor_comments":[{"comment":"The quantile definitions in the introductory paragraph of Section 3.2 are inconsistent with those in Theorem 2. The text says q_low_1(f1) is the pi_10(1)/(pi_11(1)+pi_10(1)) quantile, while the theorem defines q_low_d as the bottom pi_11(d)/Pr(R_i2=1|D_i=d) quantile. For the lower bound on the always-respondent mean, the correct quantity is the bottom pi_11/(pi_11+pi_10) quantile, as in the theorem; the earlier statement should be corrected.","section":"3.2, trimming-bounds discussion"},{"comment":"The proof of Theorem 2 is not provided. Appendix B.2 proves Proposition 4 and gives bounds on the principal-strata proportions, but the trimming-bounds argument that combines these results into Theorem 2 is not formally demonstrated. Given that Theorem 2 is the paper's main identification result, a complete proof or a precise reference to a Lee (2009)-type argument should be included.","section":"Appendix B.2"},{"comment":"The notation R_i is introduced for the post-treatment response indicator, and later R_i1 and R_i2 are used for pre- and post-treatment indicators. The paper should state explicitly at first use that R_i in Sections 2 and 3.1 refers to R_i2, to avoid ambiguity when pre-treatment missingness is introduced.","section":"Section 2.2"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses an important and timely problem, and the principal-strata parallel-trends idea is promising. The main obstacle to publication is that Theorem 2 is not identified under the stated assumptions when the pre-treatment outcome has missingness, which is precisely the situation in the paper's own application. The fix may be feasible by reformulating the estimand and trimming proportions for the R_i1 = 1 subpopulation or by adding a credible assumption that recovers the distribution of Y_i2 - Y_i1 among R_i1 = 0 units, but this is a substantive change rather than a copy-editing issue. The non-monotonicity version of the theorem also needs to be reconciled with the fact that the principal-strata proportions are only bounded. I would not recommend rejection, because the core ideas are sound and the issues appear addressable within the paper's scope, but the revision will need to be substantial."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague —\n\nThis paper has a real idea in it. Shin takes principal-stratification logic and applies it to outcome missingness in two-period DID. The genuinely new move is using parallel trends in the response indicator over time to identify the proportions of principal strata in the treated group, then feeding those proportions into Lee-type trimming bounds to partially identify the ATT for always-respondents. That is a clean, useful contribution. The paper also gives a clear account of why complete-case DID needs extra assumptions, and it offers an IV adaptation using baseline response indicators. The citations look fair, and the proofs are consistent with the stated assumptions.\n\nThe soft spots are real but fixable. First, there is a sign error in the display in Section 3.2: ATT-AR is written with a plus between the treated and control terms, when it should be a minus. That looks like a typo, but it sits right at the main estimand.\n\nSecond, and more substantively, the stress-test note is on target. Theorem 2 is stated for distributions of Δ_i = Y_i2 − Y_i1 conditioning only on R_i2=1, and the strata proportions are identified for the full treated population. But the estimand ATT-AR is defined for always-respondents, and the paper later says it is interested in those who responded at time 1, \"omitting R_i1=1 in the conditioning set for simplicity.\" That omission is not harmless when pre-treatment outcomes are missing, and the empirical application has exactly that situation — the treated group's 2016 missingness is around 70% for one outcome. As stated, Theorem 2 does not cover the application without an additional assumption like the one in the appendix. The fix is straightforward: either state the theorem conditional on R_i1=1 and derive the corresponding conditional stratum proportions, or maintain the no-pre-treatment-missingness assumption and apply the estimator only to the R_i1=1 subsample with a clear argument. But it is a point the paper needs to address.\n\nLess important but worth mentioning: no sampling uncertainty is reported for the bounds, and Assumption 6 (parallel trends of missingness) is untestable with a single pre-treatment wave. Neither is disqualifying; they are normal limitations in this kind of identification paper.\n\nOverall, this is a paper that a serious referee should engage with. The core strategy is plausible, the novelty is clear, and the gaps are presentation-level and an under-specified generalization, not a load-bearing contradiction. I would send it out, with a request to fix the sign error, sort out the R_i1 conditioning, and add a paragraph on inference.\n\nBring it to the reading group; there is plenty to discuss.","headline":"A genuinely useful partial-identification result for DID with MNAR outcomes, but the paper needs to fix a sign error and, more importantly, close the gap between Theorem 2 and its empirical application's pre-treatment missingness.","tokens_in":29711,"tokens_out":5451,"would_cite":true,"duration_ms":45496,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Under principal-strata parallel trends and parallel trends of missingness, the average treatment effect for always-respondents is partially identified by trimming the treated-respondent distribution, giving sharp bounds that do not…","keywords":["difference-in-differences","missing not at random","principal stratification","Lee bounds","panel data","attrition bias","partial identification","causal inference"],"falsifier":"Use three or more pre-treatment waves and compare the response-rate trend of eventual treated and control units during the pre-treatment period when neither group is treated; a placebo difference-in-differences of the response indicator between the two groups over adjacent pre-treatment waves should be zero if the parallel-trends-of-missingness assumption holds. A nonzero placebo estimate, or a visible divergence in the pre-treatment response-rate trajectories, would contradict the identifying assumption. Additionally, one can conduct a placebo outcome test: an outcome observed without missingness that is known to be unaffected by treatment should show no DID effect across the treated and control groups.","tokens_in":28707,"feed_emoji":"📊","tokens_out":5377,"duration_ms":51611,"temperature":0.7,"pith_summary":"This paper argues that complete-case analysis in difference-in-differences can be biased unless a restrictive parallel-trends assumption holds between respondents and nonrespondents within each treatment arm. The proposed alternative defines latent principal strata by how units would respond under treatment and under control, then imposes parallel trends within each stratum. Using the observed trend in response rates over time, the paper shows how to identify the proportion of always-respondents in the treated group, and then tailors Lee-style trimming bounds to partially identify the treatment effect for that subgroup. A secondary instrumental-variable approach uses a baseline response indicator, such as response to an auxiliary survey question, to point-identify the full ATT under a bias-homogeneity condition. If correct, the method gives DID researchers a way to handle outcomes missing not at random without imputation, covariate weighting, or assuming that treatment selection is independent of the missingness pattern.","feed_headline":"New bounds salvage DID estimates when survey responses vanish","feed_subtitle":"Using response-rate trends and principal strata, the method identifies the effect on always-respondents without imputation.","key_machinery":"The central object is the principal stratification of the response indicator, classifying each unit as always-respondent, if-treated-respondent, if-control-respondent, or never-respondent based on potential response under both treatment arms. The argument chains two assumptions together: principal-strata parallel trends (within each latent stratum, the no-treatment time trend is the same in treated and control groups) and parallel trends of missingness (under control, the change in response rates from the first wave to the second is the same for treated and control units). These assumptions identify the treated group's always-respondent share, $π_{11}(1)$, from observed response rates alone. The trimming step then replaces the unknown distribution of the always-respondent treated units with the upper or lower extreme of the observed treated-respondent distribution, yielding the partial-identification bounds. The paper's secondary machinery is a 'bespoke instrumental variable' argument, where a baseline response indicator acts as an instrument for post-treatment missingness under a bias-homogeneity condition.","core_discovery":"The central result is Theorem 2: under principal-strata parallel trends and parallel trends of missingness, the average treatment effect for always-respondents (ATT-AR) is partially identified by trimming the observed treated-respondent distribution, with bounds of the form $LB_{\\Delta,1} - E[\\Delta \\mid D=0,R=1] \\le ATT \\text{-}AR \\le UB_{\\Delta,1} - E[\\Delta \\mid D=0,R=1]$ under monotonicity, where the lower and upper bounds are quantile-trimmed means of the before-after outcome change among treated respondents. The paper establishes this by first showing that, under monotonicity plus the parallel-trends assumption on missingness, the proportion of always-respondents in the treated group is identified from observed response rates in the pre- and post-treatment waves. It then applies the trimming logic of Lee and Zhang-Rubin to the treated-respondent distribution to bound the always-respondent group's mean before-after change, and subtracts the observed control mean to obtain the band for ATT-AR. This result accommodates both dependence between treatment selection and principal strata and heterogeneous effects across strata, which the paper shows are the two core obstacles to identifying the full ATT under outcome MNAR.","pith_inferences":["With three or more pre-treatment waves, the parallel-trends-of-missingness assumption becomes testable: one can compare response-rate trends between eventual treatment and control groups in pre-treatment placebo periods, and a visible divergence would flag a violation.","The same trimming logic could be extended to multiple post-treatment waves or to covariate-specific trimming, which would likely tighten the bounds when response propensity varies with observed characteristics.","The bounds are naturally suited to sensitivity analysis: a researcher can report the range of ATT-AR alongside the estimated always-respondent share, showing whether the original complete-case point estimate falls inside the identified set.","Formal inference for the bounds should account for the fact that the trimming proportions are estimated; a bootstrap or resampling procedure over the estimated principal-strata proportions is a natural next step that the paper does not develop."],"forward_implications":["Complete-case DID estimators implicitly require that, within each treatment arm, the time trend of the outcome is the same for respondents and nonrespondents; this paper shows that assumption is not justified when treatment selection and missingness patterns are dependent or effects are heterogeneous across strata.","With only the observed response rates in two waves plus one pre-treatment wave, the proportion of always-respondents in each arm can be identified under monotonicity or a symmetric response-effect assumption, enabling the trimming bounds.","The identified set for ATT-AR is informative whenever the always-respondent share is large; in the paper's reanalysis of the aid-attitudes survey data, most outcomes have always-respondent shares around 0.8 to 0.9 and correspondingly tight bounds.","If a baseline binary response indicator (for example, response to an auxiliary pre-treatment question) satisfies relevance, parallel trends, and bias homogeneity, the full ATT can be point-identified without assuming homogeneous effects across principal strata.","The approach converts the missingness-rate trajectory itself into identifying information, so the method works for MNAR outcomes without requiring imputation models or covariate-based inverse probability weighting."],"supporting_citations":[{"why":"Supplies the trimming-bounds method the paper tailors to DID, bounding the always-respondent mean by trimming the observed treated-respondent distribution.","marker":"Lee (2009)"},{"why":"Provides the principal-strata trimming framework for outcomes truncated by death that underlies the bounds on ATT-AR.","marker":"Zhang and Rubin (2003)"},{"why":"Defines principal stratification, the latent response-type classification on which the identification strategy is built.","marker":"Frangakis and Rubin (2002)"},{"why":"Supplies the repeated-measures MNAR strategies and the 'bespoke instrumental variable' idea that the paper adapts to observational DID.","marker":"Dukes, Richardson and Tchetgen Tchetgen (2022)"},{"why":"Provides the general instrumental-variable framework for outcome MNAR whose bias-homogeneity condition is imported as Assumption 4.","marker":"Tchetgen Tchetgen and Wirth (2017)"},{"why":"Documents the prevalence and nonrandomness of missingness in political-science panel DID studies, motivating the paper's problem.","marker":"Chiu et al. (2023)"},{"why":"Offers a changes-in-changes approach to attrition bias that the paper positions against, using treatment-response subpopulations.","marker":"Ghanem et al. (2022)"}],"fun_headline_variants":["DID gets sharp bounds under outcome missingness","Principal strata rescue DID from missing outcomes","Trimmed DID bounds for always-respondents","Response rates unlock DID bounds under MNAR","Partial ID for DID when outcomes are MNAR"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the response-rate trend under the control condition is the same for the treated and control groups: in the absence of treatment, the change in survey response from the first wave to the second would have been parallel between the two arms. If treated and control units would have drifted apart in their response rates even without treatment, the estimated proportion of always-respondents is wrong and the trimming bounds are invalid.","fun_headline_variants_meta":{"raw":{"variants":["DID gets sharp bounds under outcome missingness","Principal strata rescue DID from missing outcomes","Trimmed DID bounds for always-respondents","Response rates unlock DID bounds under MNAR","Partial ID for DID when outcomes are MNAR"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000173,"raw_usage":{"total_tokens":1311,"prompt_tokens":1010,"completion_tokens":301,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":626,"completion_tokens_details":{"reasoning_tokens":233}},"tokens_in":626,"tokens_out":301,"duration_ms":3670,"temperature":1.0,"reasoning_tokens":233,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:53:48.049715+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Use three or more pre-treatment waves and compare the response-rate trend of eventual treated and control units during the pre-treatment period when neither group is treated; a placebo difference-in-differences of the response indicator between the two groups over adjacent pre-treatment waves should be zero if the parallel-trends-of-missingness assumption holds. A nonzero placebo estimate, or a visible divergence in the pre-treatment response-rate trajectories, would contradict the identifying assumption. Additionally, one can conduct a placebo outcome test: an outcome observed without missingness that is known to be unaffected by treatment should show no DID effect across the treated and control groups.","supporting_citations":[],"review_version":1}