{"id":"f5c1a344-c859-4d8a-bc6a-b959a7455b0f","arxiv_id":"2501.04853","paper_version":2,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A robust DID estimator recovers path-dependent treatment effects with partially missing treatment histories whenever any two of outcome, propensity, and missingness models are correct.","lead":"This paper builds a difference-in-differences estimator for the effects of two-period treatment paths when the first-period treatment is missing for some units. The estimator stays valid if any two of three working models, outcome, propensity score, and missingness, are correctly specified.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Efficiency bound in Theorem 2 appears incorrect; R variance in Table SE.3 falls below the claimed bound, indicating the semiparametric efficiency claim is unsupported.","rationale":"The reader's verdict ACCEPT with moderate confidence rests on the robustness identification result (Theorem 1) and the efficiency claim (Corollary 3). Theorem 1 appears correct: the triple-robustness algebra in Appendix B.2 checks out under Assumptions 1 and 2. However, the efficiency-bound derivation in Theorem 2 contains a concrete technical error: the observed-data influence function is computed by conditioning on W = (DeltaY, S, S D1, D2, X) rather than on the always-observed variables V = (DeltaY, D2, X). This is not a matter of convention; it changes the functional form of the EIF and, in general, produces a function with nonzero mean that cannot be an influence function. The paper's own Table SE.3 provides a red flag: R's variance is reported to be below the claimed semiparametric bound in two of three PDATTs, which is impossible for any regular estimator. This indicates that the SEB formula in Theorem 2 is overstated and that Corollary 3's claim that R attains the true bound is unsupported. The central identification and consistency results remain valuable, but the paper should be accepted only after the efficiency-bound section is corrected and the simulation evidence re-evaluated. The reader identified MAR as the weakest assumption, which is a substantive applied concern, but the efficiency-bound error is a more direct, internally checkable flaw in the paper's stated contribution.","tokens_in":50916,"tokens_out":26993,"duration_ms":219797,"concrete_test":"Recompute the observed-data efficient influence function using the correct formula phi_obs = S/q(D2,X) * F_full - (S/q(D2,X) - 1) * E[F_full | D2, DeltaY, X], with F_full as defined in the Theorem 2 proof, and verify whether it equals the F_tau in Theorem 2. For a simple DGP with only one treatment (D2 identical zero), D1 missing at random with q(X)=0.5 and X-independent, compute the mean of the paper's F_tau (it should be zero for a valid influence function) and compare its variance to that of the correct phi_obs. If the means differ from zero or the variances differ, then Theorem 2 and Corollary 3 are incorrect. Additionally, replicate the simulation in Table SE.3 and check whether R's variance is at least as large as the corrected semiparametric bound.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The semiparametric efficiency claim (Theorem 2 and Corollary 3) is not supported. In the proof of Theorem 2, the observed-data influence function is derived using F_obs = S/q (F_full - E[F_full|W]) + E[F_full|W] with W = (DeltaY, S, S*D1, D2, X). The correct Tsiatis formula conditions the augmentation on the always-observed variables V = (D2, DeltaY, X), giving phi_obs = S/q F_full - (S/q - 1) E[F_full | V]. When S=1, the paper's formula sets E[F_full|W] = F_full, so F_obs = F_full, dropping the required augmentation term. Consequently, equation (C.3) only retains the term involving m_d - m_d' - tau and omits the (DeltaY - m_d' - tau) contributions to E[F_full | V]. The resulting F_tau in Theorem 2 is not the efficient influence function; it typically has nonzero mean and the wrong variance. Indeed, Table SE.3 reports R variance below the claimed SEB for two PDATTs (e.g., 49.4 vs 51.1 for tau_(11)(00) in the 'None' row), which is impossible if SEB were a true lower bound. The robust identification result in Theorem 1 and the estimation/inference theory in Theorem 3 do not rely on this bound, but Corollary 3's efficiency claim and the paper's stated efficiency contribution are invalid as written.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper considers a two-period panel with a binary treatment whose first-period value D1 may be missing, and studies path-dependent average treatment effects (PDATTs) such as τ_(11)(00), τ_(10)(00), and τ_(01)(00). It shows that a standard DID estimand that ignores D1 identifies a non-convex weighted average of PDATTs and that complete-case DID is subject to selection bias. It then proposes an augmented inverse-probability-weighted estimand with three working models: an outcome regression, a propensity score, and a missing-treatment probability. Theorem 1 claims that this estimand identifies the PDATT whenever any two of the three models are correctly specified. The paper derives asymptotic normality for the two-step plug-in estimator (Theorem 3), claims semiparametric efficiency when all three models are correct (Theorem 2 and Corollary 3), and supports the results with Monte Carlo experiments and two empirical applications.","tokens_in":51237,"tokens_out":13351,"duration_ms":120702,"significance":"If the identification and inference results are correct, the paper makes a useful contribution to DID estimation with partially observed treatment histories: it relaxes the common requirement that the missingness model be correctly specified, and it nests several existing estimators as special cases. The Monte Carlo design is careful and extensive, covering misspecification of each model individually, multiple misspecifications, varying missingness rates, and varying degrees of misspecification, and the reported finite-sample behavior is consistent with the identification claims. The authors are also transparent about the strong missing-at-random assumption and provide an alternative IPW estimand under a weaker missingness assumption in SA.2. However, the efficiency claim is not supported by the derivation as written, and one of the robustness claims in Section 3.1 overstates what the assumptions deliver. These issues are substantive and need correction.","major_comments":[{"comment":"The derivation of the semiparametric efficiency bound is not correct. In equations (C.2)-(C.3), the observed-data influence function is built by conditioning on W=(∆Y,S,SD1,D2,X), and E[F_full|W] is asserted to equal p_{d1|d2}(X)1[D2=d2]P(D=d)^{-1}(m_d-m_d'-τ). For the missing-data projection theorem, the augmentation must condition on the variables observed for all units, V=(∆Y,D2,X), not on SD1. When S=1, E[F_full|W]=F_full, so the formula collapses to F_obs=F_full and the required augmentation term is dropped; when S=0 the expression in (C.3) also omits the (∆Y-m_d'-τ) components that remain in E[F_full|V]. Hence the F_τ in Theorem 2 is not an efficient influence function. Table SE.3 corroborates this: in the 'None' row the R variance lies below the claimed SEB for τ_(11)(00) (49.404 < 51.099) and for τ_(01)(00) (63.900 < 66.058), which is impossible if the reported bound were a valid lower bound. Corollary 3 and the abstract's efficiency claim therefore need to be corrected or withdrawn.","section":"Appendix C, Theorem 2"},{"comment":"The third illustrative claim after Corollary 2 states that identification still holds when 'missingness is driven by unobserved factors' and 'Assumption 2 depends on unobservables.' This is not supported: every case in the proof of Theorem 1 (Appendix B.2, equations (B.4)-(B.7), (B.13), (B.14)) invokes Assumption 2.1, which requires S to be independent of (D1,∆Y) given (D2,X). If missingness depends on unobserved factors correlated with treatment-effect heterogeneity, the robust estimand is generally biased because the missing-at-random condition fails. Assumption SA.2 only relaxes missingness to depend on ∆Y and supports a different IPW estimand, not the main robust estimand. The sentence should be removed or sharply qualified.","section":"Section 3.1"}],"minor_comments":[{"comment":"The conditioning set in (C.2) is written inconsistently: W is defined as (∆Y,S,SD1,D2,X), but the conditional expectations condition only on (∆Y,SD1,D2,X). This notation should be aligned.","section":"Appendix C, eq. (C.2)"},{"comment":"The sentence 'we extend our framework to 1 < T << n' appears to confuse the time dimension with the sample size; it should state that T is an integer number of time periods with 1 < T.","section":"Section 2.5"},{"comment":"There is a typo: 'standardized before before being used' should read 'standardized before being used.'","section":"Section 6.2"},{"comment":"The proof of Lemma SA.1 refers to 'Lemma 2.2,' but no Lemma 2.2 exists in the manuscript; the cross-reference should be corrected.","section":"Supplementary Appendix SA.2"},{"comment":"The proof is said to be in 'Supplementary Appendix D,' but Appendix D appears to be part of the main text; this cross-reference should be corrected.","section":"Corollary 3"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful core of this paper is Theorem 1: the triple-robust estimand for path-dependent ATTs with partially observed treatment histories, where any two of outcome, propensity, and missingness models being correct identifies the target. I checked the algebra in the appendix and it works. The Monte Carlo evidence backs it up, and the application to COVID-era turnout with 51% missing histories is a good illustration of why this matters. The paper is honestly written and engages the right literature. Credit where due: this is a practical extension that DID practitioners will want.\n\nThe soft spot is real, though. The stress-test note about Theorem 2 is right. The paper applies Tsiatis's missing-data influence-function formula with a conditioning set W = (ΔY, S, S·D1, D2, X) that contains the full data whenever S=1, so the augmentation term collapses to zero. The correct conditioning set for the always-observed variables is V = (D2, ΔY, X), giving F_obs = (S/q)F_full - (S/q - 1)E[F_full|V]. The proof drops the second piece. This is not a cosmetic slip: equation (C.3) omits the (ΔY - m_d' - τ) contributions to the projected score, and Table SE.3 shows R variances below the claimed semiparametric bound (for example 49.4 vs 51.1 in the 'None' row for τ_(11)(00)). That is impossible if the bound is correct. Corollary 3's efficiency claim is unsupported, and the paper's advertised efficiency contribution should be withdrawn or correctly derived before publication.\n\nThis does not sink Theorem 1 or the simulation-based inference results, which do not rely on the bound. But it does mean the paper's central selling point—efficiency among missingness-adjusted estimators—is currently overstated. The MAR assumption is also genuinely strong (missingness independent of the first treatment and outcome change given D2 and X), and the authors acknowledge it; the alternative in SA.2 is for a different estimand. That is a limitation, not a fatal flaw.\n\nWho should read this? Applied researchers working with short panels where treatment histories are incomplete. It deserves a serious referee: the identification result is substantial, the proofs are detailed enough to check, and the flaw is local and fixable. I would send it out, but with a specific request that the efficiency theorem be corrected or removed, and with the replication code supplied so the numerical claims can be verified.","headline":"The triple-robust DID estimator is a real contribution, but the efficiency-bound theorem is wrong as written and the paper overclaims it.","tokens_in":51732,"tokens_out":2367,"would_cite":true,"duration_ms":23106,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D10","62P20"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that a new robust difference-in-differences estimator identifies path-dependent treatment effects from panel data with partially missing treatment histories whenever any two of three working models are correctly specified.","keywords":["partially observed treatments","missing at random","difference-in-differences","path-dependent treatment effects","robust estimation","semiparametric efficiency","panel data","missing treatment histories"],"falsifier":"Generate data in which the missingness indicator $S$ depends on $D_1$ or on the individual treatment effect even after conditioning on $D_2$ and $X$, while keeping the outcome regression and propensity score correctly specified; under the paper's Assumption 2.1 the robust estimator should be unbiased, so any bias growing with that dependence would refute the identification claim.","tokens_in":50754,"feed_emoji":"📊","tokens_out":5561,"duration_ms":51599,"temperature":0.7,"pith_summary":"This paper tackles a gap in difference-in-differences: estimating how the full history of a binary treatment, such as being treated in both periods versus only the second, affects final outcomes when the first-period treatment is missing for part of the sample. It proposes a new estimator that combines three working models — the outcome regression, the propensity score, and the probability that the treatment history is observed — and proves that the target causal parameter is identified as long as any two of the three are correctly specified. Under the same missing-at-random and parallel-trends assumptions, conventional complete-case and doubly robust estimators require the missing-data model to be correct, while the new estimator remains unbiased when that model fails. The paper also derives the semiparametric efficiency bound for the parameter and shows its estimator attains it when all three models are correct. This matters because missing treatment histories are routine in surveys, rotating panels, and administrative data, and because existing estimators can be biased whether missingness is ignored or modeled.","feed_headline":"Any two of three models identify missing-history treatment effects","feed_subtitle":"New DID estimator stays unbiased when the missingness model is wrong; existing DR and complete-case methods do not.","key_machinery":"The load-bearing object is the augmented inverse-probability-weighted estimand $\\tau^R_{dd'}$ in equation (4), constructed from three working models: $\\mu_d(X)$ for the outcome mean, $\\pi_d(X)$ for the propensity score, and $\\phi_{d_2}(X)$ for the probability of observing $D_1$ given $D_2$ and $X$. Its four Hajek-type weights $w_1, w_2, w_3, w_4$ re-weight the observed treated and comparison groups, the $D_2$-only groups, and differences in outcome-model predictions. The mechanism is cancellation: when any two models are correct, the differences between pairs of weight-adjusted expectation terms vanish in expectation, leaving only the term that identifies the PDATT. This same structure nests the OR, IPW, and DR estimands as special cases, which is why those alternatives all require a correct missing-data model whereas the robust estimator does not.","core_discovery":"The paper's central claim is that the causal parameter $\\tau_{dd'} = E[Y_2(d) - Y_2(0,0) \\mid D = d]$ — the average effect of treatment path $d$ on the final outcome for individuals who followed that path — can be identified even when $D_1$ is unobserved for some units. Identification is achieved by the robust estimand in equation (4), which re-weights observed outcomes and outcome-model differences using four Hajek-normalized weights built from the missingness probability $q_{d_2}(X)$, the propensity score $p_d(X)$, and the outcome regression $m_d(X)$. Theorem 1 proves this estimand equals $\\tau_{dd'}$ whenever any two of the three models are correctly specified: outcome plus propensity score, missingness plus propensity score, or missingness plus outcome. The proof shows that, under each pair of correct models, the spurious adjustment terms in the estimand vanish in expectation and the surviving term equals the path-dependent treatment effect. When all three models are correct, Corollary 3 shows the estimator's asymptotic variance equals the semiparametric efficiency bound for $\\tau_{dd'}$, so it is efficient within the class of missingness-adjusted estimators.","pith_inferences":["The paper's robustness claim is conditional on Assumption 2.1 being true; if missingness is driven by unobserved factors correlated with treatment-effect heterogeneity, the robust estimator inherits the same bias as any MAR-based method even with perfectly specified models.","An extension the paper leaves implicit is a practical diagnostic: comparing robust and DR estimates, a large divergence that shrinks when the missingness model is respecified would point to the missingness model as the fragile component rather than the causal parameter itself.","The paper notes that a weaker assumption allowing missingness to depend on the outcome change requires a different inverse-probability-weighted estimand; the main robust estimator is not designed for that case, so its triple-robustness should not be read as covering outcome-dependent missingness.","The Monte Carlo design, which misspecifies each model by using nonlinear transformations of the covariates, suggests that in real applications the triple-robust property is most valuable when the missingness model is the hardest to justify, which is often the case with survey nonresponse and attrition."],"forward_implications":["Applied DID studies with partially missing treatment histories can report a causally interpretable PDATT rather than a non-convex mixture of effects from ignoring $D_1$ or a selection-biased complete-case estimate.","Researchers need not correctly model the missingness mechanism: as long as the outcome regression and propensity score are correct, the robust estimator identifies the target parameter even when missingness depends on covariates in unknown ways.","The OR, IPW, and DR estimators are formally nested in the robust estimand, so the paper's inference machinery provides valid standard errors for those alternatives as well.","With all three models correct, the estimator attains the semiparametric efficiency bound, so the efficiency loss from guarding against misspecification of the missing-data model disappears.","The framework extends to multiple periods and arbitrary missing patterns in the treatment history, so the result applies beyond the three-period, $D_1$-missing case."],"supporting_citations":[{"why":"supplies the doubly robust DID framework and improved-estimation ideas that the robust estimand generalizes to persistent treatment effects and missing histories.","marker":"(Sant'Anna and Zhao, 2020)"},{"why":"defines the PDATT-type causal parameters and staggered-adoption setting that the paper nests as a special case.","marker":"(Callaway and Sant'Anna, 2021)"},{"why":"provides the missing-data projection result used to derive the semiparametric efficiency bound under missing at random.","marker":"Theorem 7.2 in Tsiatis (2006)"},{"why":"supplies the covariate transformations used in the Monte Carlo experiments to create misspecified working models.","marker":"(Kang and Schafer, 2007)"},{"why":"motivates the inference-robust first-stage estimators that minimize the first-order effect of nuisance-parameter estimation.","marker":"(Vermeulen and Vansteelandt, 2015)"},{"why":"established a similar multiply robust ATE property with missing treatment information in a cross-sectional setting, which this paper extends to dynamic DID.","marker":"(Zhang et al., 2016)"},{"why":"justifies the normalized weighting scheme used in the estimand.","marker":"(Hajek, 1971)"}],"fun_headline_variants":["Two of three models suffice for missing-history effects","Partial histories? Any two models still identify the effect","Robust DID: miss one model, keep consistency","Missing treatment paths? Two correct models do the job","Path effects with gaps: two models are enough"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The missing-at-random condition $S \\perp (D_1, \\Delta Y) \\mid D_2, X$ must hold, meaning that whether the first treatment is observed is independent of the first treatment and of the outcome change once the second treatment and covariates are controlled; if missingness is driven by unobserved factors tied to treatment-effect heterogeneity, the robust estimator is biased even with perfectly specified models.","fun_headline_variants_meta":{"raw":{"variants":["Two of three models suffice for missing-history effects","Partial histories? Any two models still identify the effect","Robust DID: miss one model, keep consistency","Missing treatment paths? Two correct models do the job","Path effects with gaps: two models are enough"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000246,"raw_usage":{"total_tokens":1524,"prompt_tokens":916,"completion_tokens":608,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":532,"completion_tokens_details":{"reasoning_tokens":534}},"tokens_in":532,"tokens_out":608,"duration_ms":6048,"temperature":1.0,"reasoning_tokens":534,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:24:11.471151+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate data in which the missingness indicator $S$ depends on $D_1$ or on the individual treatment effect even after conditioning on $D_2$ and $X$, while keeping the outcome regression and propensity score correctly specified; under the paper's Assumption 2.1 the robust estimator should be unbiased, so any bias growing with that dependence would refute the identification claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the doubly robust DID framework and improved-estimation ideas that the robust estimand generalizes to persistent treatment effects and missing histories."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"defines the PDATT-type causal parameters and staggered-adoption setting that the paper nests as a special case."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the covariate transformations used in the Monte Carlo experiments to create misspecified working models."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"motivates the inference-robust first-stage estimators that minimize the first-order effect of nuisance-parameter estimation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"established a similar multiply robust ATE property with missing treatment information in a cross-sectional setting, which this paper extends to dynamic DID."}],"review_version":1}