Pith. sign in

REVIEW 2 major objections 2 minor 44 references

When Do Treatment Changes Identify Causal Effects?

T0 review · 2 major / 2 minor · reviewed 2026-06-28 · grok-4.3

Pith's one-line read Treatment changes identify causal effects by removing time-constant confounders under two non-nested structural models.

desk verdict The paper maps the assumptions behind treatment-change identification and shows an equivalence to level-based methods only under random walk on the treatment process. read the letter →

arxiv 2606.02234 v2 pith:IBYZKNZF submitted 2026-06-01 econ.EM stat.ME

classification econ.EMstat.ME
keywords causalinferencetreatmentchangesdifference-in-differencesselectiononobservablesrandomwalkdoublerobustnesspaneldatafixedeffects
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper characterizes the conditions under which differencing treatments over time identifies causal effects after conditioning on observed covariates. It does so by differencing out time-constant confounders that enter additively in the treatment equation, in two distinct structural models. These assumptions do not nest with those of selection-on-observables strategies that control for past outcomes and treatments or with difference-in-differences that differences outcomes instead. Under a random-walk restriction on the treatment process, however, using treatment changes becomes equivalent to using treatment levels conditional on lagged treatment. In partially linear models the non-nesting yields a double-robustness property for two-way fixed-effects regression that differences both outcome and treatment.

What carries the argument

Differencing out time-constant confounders additive in the treatment equation, under either a random-walk restriction or conditions that rule out dynamic effects.

What would settle it

Empirical or simulated data in which the treatment process deviates from a random walk, dynamic treatment effects are present, and estimates from treatment changes diverge from those obtained from levels conditional on lagged treatment or from standard DiD.

Watch

Extended reading notes

Core claim

The paper shows that treatment-change identification is valid conditional on covariates in two structural models with non-nested assumptions, because changes difference out time-constant confounders additive in the treatment equation. Under a random-walk restriction on the treatment process, this strategy is equivalent to using treatment levels given lagged treatment, which permits overidentification tests. Under an alternative model that rules out dynamic treatment effects, treatment changes can serve as an instrument for a constant treatment effect. In partially linear models these results imply that two-way fixed-effects regression remains consistent if either the treatment-change assumpt

Load-bearing premise

The treatment process must obey a random walk, or the alternative conditions that rule out dynamic treatment effects must hold, for the identification, equivalence, and double-robustness results to apply.

Editorial extensions

If this is right

  • Exploiting treatment changes is equivalent to using treatment levels given lagged treatment when the random-walk restriction holds.
  • Overidentification tests become feasible by comparing change-based and level-based estimates.
  • Two-way fixed-effects regression that differences both outcome and treatment is consistent if either the change-based or parallel-trends assumption is satisfied.
  • Treatment changes can identify a constant effect as an instrument when dynamic effects are ruled out.
  • The assumptions for change-based identification are generally not nested with those of selection-on-observables or conventional DiD.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Applied researchers can routinely compare change-based and level-based estimates to check consistency when panel data are available.
  • In settings where treatment follows a non-random-walk process, change-based methods may recover parameters distinct from those recovered by level-based methods.
  • The double-robustness property reduces the risk that two-way fixed-effects estimates are invalidated by violation of either the change or parallel-trends assumption alone.
  • The results suggest designing overidentification tests that exploit both changes and levels in the same sample.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper characterizes two structural models with non-nested assumptions under which treatment changes identify causal effects by differencing out additive time-constant confounders conditional on covariates. Under a random-walk restriction on the treatment process, it establishes equivalence between change-based identification and level-based methods conditional on lagged treatment; under an alternative model ruling out dynamic treatment effects, changes serve as instruments for a constant treatment effect. These results imply double robustness for two-way fixed effects regression in partially linear models (consistent if either the change-based or parallel-trends assumption holds) and motivate overidentification tests. The paper includes simulations and an empirical application to cigarette demand.

Significance. If the derivations hold, the paper offers a useful clarification of the distinct identifying assumptions behind change-based versus level-based strategies in panel data, with the non-nesting results and double-robustness implication for TWFE providing practical guidance for applied researchers. The simulations and empirical application are strengths that illustrate the theoretical points and demonstrate applicability.

major comments (2)
  1. [Section characterizing the random-walk model] The equivalence between treatment-change and lagged-treatment-level strategies (and the associated overidentification tests) is derived under the random-walk restriction on the treatment process; this assumption is load-bearing for the nesting claim, yet the manuscript does not provide a formal statement of the random-walk process (e.g., as an equation for the treatment dynamics) or discuss its implications for mean reversion common in economic data.
  2. [Section on partially linear models and TWFE] The double-robustness result for TWFE regression that differences both outcome and treatment is stated for partially linear models; the manuscript should clarify whether this holds for heterogeneous treatment effects or only average effects, as the non-nesting with parallel trends is central to the robustness claim.
minor comments (2)
  1. The abstract and introduction could more explicitly flag that all equivalence and robustness results are conditional on the stated restrictions (random walk or no dynamic effects) to avoid overgeneralization by readers.
  2. In the empirical application, report the exact specification of the TWFE estimator and the overidentification test statistic to allow direct replication of the cigarette-demand results.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive comments, which will improve the clarity of the manuscript. We address each major comment below and plan to revise accordingly.

read point-by-point responses
  1. Referee: [Section characterizing the random-walk model] The equivalence between treatment-change and lagged-treatment-level strategies (and the associated overidentification tests) is derived under the random-walk restriction on the treatment process; this assumption is load-bearing for the nesting claim, yet the manuscript does not provide a formal statement of the random-walk process (e.g., as an equation for the treatment dynamics) or discuss its implications for mean reversion common in economic data.

    Authors: We agree that an explicit formalization is needed. In the revision we will add the random-walk equation D_it = D_i,t-1 + ε_it (with E[ε_it | covariates] = 0) and discuss that this rules out mean reversion in treatment levels, which may be restrictive in some economic applications but is the condition that delivers the nesting with lagged-treatment strategies. revision: yes

  2. Referee: [Section on partially linear models and TWFE] The double-robustness result for TWFE regression that differences both outcome and treatment is stated for partially linear models; the manuscript should clarify whether this holds for heterogeneous treatment effects or only average effects, as the non-nesting with parallel trends is central to the robustness claim.

    Authors: The result is derived under the partially linear model, which imposes a treatment effect that is constant conditional on covariates. We will clarify in the revision that double robustness applies to this conditional average treatment effect (and note that fully heterogeneous effects would require a different framework). The non-nesting with parallel trends continues to hold for the average effect under the stated conditions. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: results derived from explicit structural models and stated restrictions

full rationale

The paper characterizes identification under two explicitly stated structural models with non-nested assumptions, derives equivalence only under an additional random-walk restriction on the treatment process, and shows double-robustness implications for TWFE under partially linear models. All steps rest on the paper's own model equations and assumptions rather than fitted parameters, self-citations, or renamings; no quantity is defined in terms of another by construction, and the derivation remains self-contained against the stated conditions.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The central claims rest on domain assumptions about the treatment process and absence of dynamic effects; no free parameters or invented entities are introduced in the abstract.

assumptions (2)
  • domain assumption random-walk restriction on the treatment process
    Invoked to establish equivalence between treatment-change and treatment-level identification given lagged treatment.
  • domain assumption rules out dynamic treatment effects (among other conditions)
    Required for the instrument-based identification result when the random-walk assumption is dropped.

how reviews work

0 comments
Cite this review

Pith. "Pith review of When Do Treatment Changes Identify Causal Effects?." pith.science (2026). https://pith.science/paper/IBYZKNZF

@misc{pith2026260602234,
  author       = {Pith},
  title        = {Pith review of: When Do Treatment Changes Identify Causal Effects?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IBYZKNZF}},
  note         = {Machine review of arXiv:2606.02234}
}
read the original abstract

This paper clarifies the identifying assumptions underlying causal inference based on treatment changes rather than levels, and their relationship to conventional identification strategies. We characterize two structural models, with non-nested assumptions, under which treatment-change identification is valid conditional on observed covariates by differencing out time-constant confounders that are additive in the treatment equation. The assumptions underlying treatment changes are generally not nested with those of methods relying on treatment levels, such as selection-on-observables strategies that control for past outcomes, treatments, and covariates, or difference-in-differences approaches that difference outcomes rather than treatments over time. We show, however, that under a random-walk restriction on the treatment process, exploiting treatment changes for identification is equivalent to using treatment levels given lagged treatment. This and other equivalence results motivate overidentification tests based on methods considering treatment levels and changes. Under an alternative model that does not assume a random walk but instead rules out dynamic treatment effects (among other conditions), treatment changes can still be used as an instrument to identify a treatment effect that is constant given covariates. However, without random walk, different identification strategies are generally not nested. In partially linear models, the non-nesting results carry a double robustness implication for two-way fixed effects regression that differences both the outcome and the treatment over time, which under certain conditions remains consistent if either the treatment-change assumption or the parallel-trends assumption holds. We characterize the causal models consistent with each method, run simulations for illustration, and present an empirical application to cigarette demand.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

44 extracted references · 6 canonical work pages

  1. [1]

    Abadie, A. (2005). Semiparametric difference-in-differences estimators.Review of Economic Studies,72, 1-19

  2. [2]

    Angrist, J., & Imbens, G. W. (1995). Two-stage least squares estimation of average causal effects in models with variable treatment intensity.Journal of American Statistical Association,90, 431-442

  3. [3]

    D., & Pischke, J.-S

    Angrist, J. D., & Pischke, J.-S. (2009).Mostly harmless econometrics: An empiricist’s companion. Princeton university press

  4. [4]

    Arkhangelsky, D., & Imbens, G. W. (2021).Double-robust identification for causal panel data models(Working Paper No. 28364). National Bureau of Economic Research. doi: 10.3386/w28364

  5. [5]

    Arkhangelsky, D., & Imbens, G. W. (2022). Doubly robust identification for causal panel data models.The Econometrics Journal,25(3), 649–674

  6. [6]

    Ashenfelter, O. (1978). Estimating the effect of training programms on earnings.The Review of Economics and Statistics,6, 47-57

  7. [7]

    Athey, S., Tibshirani, J., & Wager, S. (2019). Generalized random forests.Annals of Statistics,47(2), 1148–1178. doi: 10.1214/18-AOS1709

  8. [8]

    H., Dorn, D., & Hanson, G

    Autor, D. H., Dorn, D., & Hanson, G. H. (2013). The china syndrome: Local labor market effects of import competition in the united states.American economic review,103, 2121-2168

Show all 44 references
  1. [9]

    H., & Levin, D

    Baltagi, B. H., & Levin, D. (1986). Estimating dynamic demand for cigarettes using panel data: The effects of bootlegging, taxation and advertising reconsidered.Review of Economics and Statistics,68(1), 148–155

  2. [10]

    Bartik, T. J. (1991). Who benefits from state and local economic development policies?

  3. [11]

    Borusyak, K., Hull, P., & Jaravel, X. (2021). Quasi-Experimental Shift-Share Research Designs.The Review of Economic Studies,89, 181-213

  4. [12]

    (2024).A practical guide to shift-share instruments

    Borusyak, K., Hull, P., & Jaravel, X. (2024).A practical guide to shift-share instruments

  5. [13]

    Borusyak, K., Jaravel, X., & Spiess, J. (2024). Revisiting Event-Study Designs: Robust and Efficient Estimation.The Review of Economic Studies, rdae007

  6. [14]

    Callaway, B., & Sant’Anna, P. H. (2021). Difference-in-differences with multiple time periods. Journal of Econometrics,225, 200-230

  7. [15]

    Card, D., & Krueger, A. B. (1994). Minimum wages and employment: A case study of the fast-food industry in new jersey and pennsylvania.The American Economic Review, 84, 772-793. Chabé-Ferret, S. (2017). Should we combine difference in differences with conditioning on 38 pre-tr...

  8. [16]

    J., & Warner, K

    Chaloupka, F. J., & Warner, K. E. (2000). The economics of smoking. In A. J. Culyer & J. P. Newhouse (Eds.),Handbook of health economics(Vol. 1, pp. 1539–1627). Elsevier

  9. [17]

    (1958).Planning of experiments

    Cox, D. (1958).Planning of experiments. New York: Wiley

  10. [18]

    Croissant, Y., & Millo, G. (2008). Panel data econometrics in R: The plm package.Journal of Statistical Software,27(2), 1–43. Retrieved from https://doi.org/10.18637/jss.v027.i02 doi: 10.18637/jss.v027.i02 de Chaisemartin, C., Ciccia, D., D’Haultfœuille, X., & Knau, F. (2024)....

  11. [19]

    Firpo, S. (2007). Efficient semiparametric estimation of quantile treatment effects.Econo- metrica,75, 259-276

  12. [20]

    Fricke, H. (2017). Identification based on difference-in-differences approaches with multiple treatments.Oxford Bulletin of Economics and Statistics,79, 426-433

  13. [21]

    Goldsmith-Pinkham, P., Sorkin, I., & Swift, H. (2020). Bartik instruments: What, when, why, and how.American Economic Review,110

  14. [22]

    Goodman-Bacon, A. (2021). Difference-in-differences with variation in treatment timing. Journal of Econometrics,225, 254-277

  15. [23]

    F., Huber, M., & Zhang, L

    Haddad, M. F., Huber, M., & Zhang, L. Z. (2024). Difference-in-differences with time- varying continuous treatments using double/debiased machine learning.arXiv preprint 2410.21105

  16. [24]

    Hausman, J. A. (1978). Specification tests in econometrics.Econometrica,46(6), 1251–1271

  17. [25]

    Huber, M., & Oeß, E.-M. (2024). A joint test of unconfoundedness and common trends. arXiv preprint 2404.16961

  18. [26]

    Imbens, G. W. (2004, Feb.). Nonparametric estimation of average treatment effects under exogeneity: a review.The Review of Economics and Statistics,86, 4-29

  19. [27]

    W., & Angrist, J

    Imbens, G. W., & Angrist, J. (1994). Identification and estimation of local average treatment effects.Econometrica,62, 467-475

  20. [28]

    C., & Pfleiderer, H

    Knaus, M. C., & Pfleiderer, H. (2026). Causal graphs for conditional parallel trends.arXiv preprint 2604.12818

  21. [29]

    Lechner, M. (2011). The estimation of causal effects by difference-in-difference methods. Foundations and Trends in Econometrics,4, 165-224

  22. [30]

    Mogstad, M., & Torgovitsky, A. (2024). Instrumental variables with unobserved heterogeneity in treatment effects. InHandbook of labor economics(Vol. 5, pp. 1–114). Elsevier. 39

  23. [31]

    Neyman, J. (1923). On the application of probability theory to agricultural experiments. essay on principles.Statistical Science,Reprint, 5, 463-480

  24. [32]

    (1988).Probabilistic reasoning in intelligent systems: networks of plausible inference

    Pearl, J. (1988).Probabilistic reasoning in intelligent systems: networks of plausible inference. San Mateo: Morgan Kaufmann

  25. [33]

    (2000).Causality: Models, reasoning, and inference

    Pearl, J. (2000).Causality: Models, reasoning, and inference. Cambridge: Cambridge University Press

  26. [34]

    Robins, J. M. (1986). A new approach to causal inference in mortality studies with sustained exposure periods - application to control of the healthy worker survivor effect. Mathematical Modelling,7, 1393-1512

  27. [35]

    M., Hernan, M

    Robins, J. M., Hernan, M. A., & Brumback, B. (2000). Marginal structural models and causal inference in epidemiology.Epidemiology,11, 550-560

  28. [36]

    M., Rotnitzky, A., & Zhao, L

    Robins, J. M., Rotnitzky, A., & Zhao, L. (1994). Estimation of regression coefficients when some regressors are not always observed.Journal of the American Statistical Association,90, 846-866

  29. [37]

    Rubin, D. (1980). Comment on ’randomization analysis of experimental data: The fisher randomization test’ by d. basu.Journal of American Statistical Association,75, 591-593

  30. [38]

    Rubin, D.B. (1974). Estimatingcausaleffectsoftreatmentsinrandomizedandnonrandomized studies.Journal of Educational Psychology,66, 688-701

  31. [39]

    Sun, L., & Abraham, S. (2021). Estimating dynamic treatment effects in event studies with heterogeneous treatment effects.Journal of Econometrics,225, 175-199

  32. [40]

    Tibshirani, J., Athey, S., & Wager, S. (2020). grf: Generalized random forests.R package

  33. [41]

    Wager, S., & Athey, S. (2018). Estimation and inference of heterogeneous treatment effects using random forests.Journal of the American Statistical Association,113, 1228-1242

  34. [42]

    M., van der Laan, M

    Weber, A. M., van der Laan, M. J., & Petersen, M. L. (2015). Assumption trade-offs when choosing identification strategies for pre-post treatment effect estimation: An illustration of a community-based intervention in madagascar.Journal of Causal Inference,3, 109–130

  35. [43]

    Wright, P. G. (1928).The tariff on animal and vegetable oils. The Macmillan Company

  36. [44]

    Xu, Y. (2023). Causal inference with time-series cross-sectional data: a reflection.SSRN 3979613. 40

Pith tools

Reviewed June 28, 2026 · model on record in the stance chip above.