REVIEW 4 major objections 5 minor 36 references
State-policy heterogeneity analyses should bound state-specific effects instead of estimating conditional averages, because bounds can determine a policy's sign for individual states even when point identification fails.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
A difference-in-differences bounding method is proposed that reports ranges for state-specific policy effects, and it recovers effect signs more reliably than conditional average effect estimates in simulations.
T0 review reviewed 2026-08-03 challenge →
load-bearing objection State-specific effect bounds are a genuinely useful shift of target, but the data-informed tau undercalibrates in the paper's own simulations and the application reads more certainty than the method supports. the 4 major comments →
State policy heterogeneity analyses: considerations and proposals
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The central claim is that when the data cannot point-identify a state-level policy effect, the inferential goal should shift from estimating a conditional average to bounding the individual (state-specific) effect. Under a bounded-counterfactual-error assumption—the error in predicting what each state would have done without (or with) the policy is no larger than a chosen τ_i—each state's ITE lies in an interval centered on a unit-level difference-in-differences estimate. The paper develops this interval, shows how pre-treatment residuals can inform τ_i, and proves that untreated states generally require larger sensitivity parameters than treated states (Z1 > Z0). Simulations show that the r
What carries the argument
The key machinery is the unit-level difference-in-differences estimator ˆψ^{DiD}_{i,m} combined with Assumption 2 (bounded counterfactual error), which converts an unverifiable point-identifying assumption into an interval ψ_{i,m} ∈ [ˆψ^{DiD}_{i,m} − τ_i, ˆψ^{DiD}_{i,m} + τ_i]. The width τ_i is the sensitivity parameter; the paper proposes setting τ_i = Z∥˜ε_i∥ using pre-treatment residuals, with Lemma 2 and Proposition 6 showing how the residual norm can be scaled for treated versus untreated units. Treatment coarsening enters through a two-stage randomization framework that interprets coarsened DiD estimates as stochastic interventions over policy versions, and the paper offers three ways
Load-bearing premise
The load-bearing premise is that the pre-treatment residuals, scaled by a constant, genuinely cover the size of the error in predicting each state's no-policy (or with-policy) outcome; the paper notes this is not guaranteed, and its own simulations show that for control states the data-informed bounds cover the true effect only about 57–59% of the time at the chosen Z.
What would settle it
A falsifying experiment would simulate states where the post-treatment counterfactual error is, by construction, larger or more skewed than the pre-treatment residual norm used for τ; if the data-informed bounds cover the true ITE at rates well below the intended level even at the recommended Z, the calibration assumption fails. In the application, re-estimating the Illinois and New Mexico bounds with a longer pre-period and a different comparison pool would provide a direct check: if either state's interval flips sign or contains zero under a moderate Z, the headline application conclusion wo
If this is right
- Researchers can report 'the effect in state i lies between L and U' and interpret it as the causal effect for that state, rather than an average over states with similar characteristics.
- In the paper's simulations, the ITE bounds determine the sign of the true state effect more reliably than CATE confidence intervals, especially for treated states.
- Pre-treatment residuals can be used to calibrate sensitivity parameters, and varying the multiplier Z gives a tipping-point analysis that is directly interpretable on the scale of observed pre-period deviations.
- When treatment is coarsened, untreated-state bounds require larger sensitivity parameters; the paper formalizes why and gives three concrete strategies (known M(1), union of bounds, conservative τ).
- In the application, the bounds identify Illinois as having a negative and New Mexico as having a positive Medicaid-expansion effect on high-volume buprenorphine prescribing, consistent across eight specifications.
Where Pith is reading between the lines
- A natural extension the paper leaves implicit is applying the same residual-calibration bounding to staggered adoption and dynamic effects; the logic should carry over, but the counterfactual error bound would need to account for treatment timing.
- The paper's own simulation shows data-informed bounds covering control-state ITEs only about 57–59% of the time at Z=2, suggesting practitioners should treat Z as a sensitivity knob tilted toward larger values for untreated states, not as a calibrated confidence level.
- Because CATE confidence intervals target a projected average rather than individual states, published state-policy heterogeneity findings based on interaction terms may overstate confidence in which states benefit; reanalysis with these bounds could change policy rankings.
- The bounding approach could be combined with placebo-covariate balance checks or Bayesian hierarchical models as an additional robustness overlay, though the paper mentions Bayesian models only in passing.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses a recurring problem in state-level policy heterogeneity analyses: standard estimands such as CATEs and their linear projections are often not aligned with policy questions about what happened or would happen in a specific state. It formalizes the distinction between ITEs, CATEs, and CDEs, and shows how treatment coarsening further complicates interpretation. The main proposal is to abandon point identification and instead bound state-specific ITEs under a bounded counterfactual error assumption (Assumption 2), with sensitivity parameters informed by pre-treatment residuals. Formal results include worst-case bounds under bounded errors (Proposition 5), a probabilistic MSE-based lemma (Lemma 2), and a normality-based scaling result for control-state errors (Proposition 6). A simulation compares ITE bounds with CATE confidence intervals, and an application to Medicaid expansion and high-volume buprenorphine prescribing reports robust-looking negative bounds for Illinois and positive bounds for New Mexico across 6 of 8 specifications.
Significance. The conceptual contribution is genuinely useful. Distinguishing ITE, CATE, and CDE, and making the case that CATEs are often associational rather than causal for state-specific questions, is an important clarification for applied work. The bounding framework is transparent and interpretable: when the true tau* is known, the interval contains the ITE by construction, and the oracle version gives correct coverage. The paper is also unusually open about limitations: Section 4.2 explicitly concedes that no pre-data rule guarantees tau_i >= tau_i^star, and Table 4 reports the consequences. The availability of code for simulations and application is a concrete strength. However, the paper's headline claims—that the method can "more reliably determine the sign of the ITEs than CATE estimates" and that the Illinois/New Mexico findings are strongly suggested—depend on calibration choices that the paper does not deliver. The framework is defensible as a sensitivity analysis, but not as a calibrated procedure with the stated coverage/validity implications.
major comments (4)
- [Section 4.2 and Table 4] The central claim rests on the data-informed sensitivity parameter tau_i^m = Z ||epsilon_i||. The paper itself states in Section 4.2 that "none of these approaches guarantee tau_i^m >= tau_i^star", so the bounds are not guaranteed to contain the ITE. Table 4 quantifies the problem: at Z=2, coverage for treated units is 0.833–0.842 and for control units only 0.570–0.587. These are far below any nominal confidence level. Thus the abstract's claim that the bounds "can more reliably determine the sign of the ITEs than CATE estimates" and the application's "strongly suggest" language overstate what the simulations establish. The paper should either calibrate Z (e.g., by reporting coverage over a grid of Z and choosing a value that meets a target) or explicitly reframe the intervals as pure sensitivity intervals with no coverage guarantee. This is not a presentation issue; it is load-bearing f
- [Proposition 6, Remark 11, and Section 4.4] The formal support for control-state bounds requires assumptions that are not met in the simulation or application. Proposition 6 assumes normality of the errors and known mean/scale parameters alpha and gamma, plus inequality (11); none of these are verifiable from pre-period data. More importantly, Remark 11 says that for untreated units one should generally choose Z1 > Z0, but the simulation (Section 4.4) and the application (Section 5.1) set Z1 = Z0 = 2. The simulation's poor control-state coverage (0.54–0.59) is, in part, a direct consequence of not following the paper's own guidance. This matters because the control-state bounds are the weakest link. The authors should either implement a calibration that respects Remark 11 and report coverage, or explain why equal Z1 and Z0 is a deliberate choice and temper the conclusions accordingly.
- [Section 4.4 and Table 4] The simulation's headline comparison is between ITE bounds and CATE confidence intervals, but the CATE estimator targets a linear projection of a coarsened CATE, not the ITE. It is unsurprising that intervals for one estimand have poor coverage and sign recovery for a different estimand. The claim "bounding state-specific effects can more reliably determine the sign of the ITEs than CATE estimates" is therefore a comparison of incomparable objects. A fairer assessment would compare the ITE bounds against an estimator that also targets the ITE, or at least report the CATE comparison with an explicit caveat that the two methods answer different questions. The paper does note this caveat, but the abstract and Section 6 still present the comparison as evidence for the bounding method.
- [Section 5.1 and Table 3] The application's robustness claim is summarized only as counts of "6 of 8" specifications for Illinois and New Mexico. Table 3 does not report which specification produced which sign, the values of Z used, or how the alternative specifications differ in assumptions. Since the eight specifications are not independent—they share the same data, the same pre-treatment residual construction, and the same choice Z1=Z0=2—the count "6 of 8" overstates the evidential weight. The authors should provide a full specification-by-specification table and show how the classification changes as Z is varied, rather than only reporting a binary count.
minor comments (5)
- [Definition 5] The coarsened ITE is denoted tilde(psi), whereas Definition 1 uses psi for the ITE; the notation should be consistent to avoid confusion.
- [Section 3.4] "Cohon's d" should be "Cohen's d".
- [Section 1] "implemeted" is a typo for "implemented".
- [Section 4.2] The paragraph introducing Z^m would benefit from stating explicitly that Z=0 corresponds to exact parallel trends and that larger Z does not make the interval a confidence interval; the current wording is close but could be sharper given the paper's emphasis on transparency.
- [Appendix A.2.2] Equation (11), used in the text and in the skeptical summary, does not have a visible equation number in the manuscript. Please number all key equations referenced in the text.
Circularity Check
No circular reasoning: the ITE bounds are an explicit sensitivity analysis, not a fitted prediction.
full rationale
The derivation of the ITE bounds is an explicit, assumption-driven sensitivity analysis rather than a fitted prediction. Assumption 2 states |Ŷ_i,T(m)−Y_i,T(m)| ≤ τ_i^m, and equation (2) then places ψ_i,m in [ψ̂_i,m^DiD − τ̂_i^m, ψ̂_i,m^DiD + τ̂_i^m]. The paper immediately acknowledges that choosing τ = τ⋆ makes Assumption 2 hold by definition and that the open question is how to choose τ. This is a transparent tautology, not a hidden reduction: the bounds are not fitted to the target ITE, and no guarantee is claimed for data-informed τ. The simulation uses a fixed Z=2, reports control-state coverage of 0.57–0.59 in Table 4, and explicitly disclaims formal statistical guarantees, so the method is not presented as a calibrated predictor. The application's Illinois/New Mexico statements are robustness counts over eight sensitivity specifications and remain conditional on an unvalidated τ choice; that is a validity limitation, not circularity. Load-bearing support comes from external references (Manski & Pepper, Rambachan & Roth, VanderWeele & Hernán), and the authors' self-citations are contextual—prior estimates of Medicaid expansion effects and high-volume prescriber concentration—rather than used to force the paper's conclusions. No step in the claimed derivation chain reduces to its own inputs.
Axiom & Free-Parameter Ledger
free parameters (4)
- Sensitivity parameter Z^m =
Z=2 in simulation; Z in {1, 1.5, 2} in application
- Norm choice for pre-treatment residuals =
MAE in simulation; L-infinity in application
- Rurality threshold in application =
States with more than 25% of buprenorphine prescribers in rural communities
- Mean/variance scaling for untreated-state error (alpha, gamma) =
User-specified; examples alpha=1, gamma=1 or alpha=gamma=0
axioms (8)
- domain assumption SUTVA, consistency, and no anticipatory effects
- domain assumption A single well-defined control condition M=0
- domain assumption Parallel trends (Assumptions 1, 5.a, 5.b)
- domain assumption Bounded counterfactual error (Assumption 2)
- domain assumption Pre-treatment residuals are informative about post-treatment counterfactual error
- ad hoc to paper Normality and mean/variance scaling of errors (Proposition 6)
- domain assumption Exclusion restriction: A affects Y only through M
- domain assumption Positivity of treatment assignment
invented entities (1)
-
M_i(1): counterfactual policy version assignment for untreated states
no independent evidence
Cite this review
Pith. "Pith review of State policy heterogeneity analyses: considerations and proposals." pith.science (2026). https://pith.science/paper/R3WD5XOZ
@misc{pith2026260208643,
author = {Pith},
title = {Pith review of: State policy heterogeneity analyses: considerations and proposals},
year = {2026},
howpublished = {\url{https://pith.science/paper/R3WD5XOZ}},
note = {Machine review of arXiv:2602.08643}
}
read the original abstract
State-level policy studies often conduct heterogeneity analyses that quantify how treatment effects vary across state characteristics. These analyses may be used to inform state-specific policy decisions, or to infer how the effect of a policy changes in combination with other state characteristics. However, in state-level settings with varied contexts and policy landscapes, multiple versions of similar policies, and differential policy implementation, the causal quantities targeted by these analyses may not align with the inferential goals. This paper clarifies these issues by distinguishing several causal estimands relevant to heterogeneity analyses in state-policy settings, including state-specific treatment effects (ITE), conditional average treatment effects (CATE), and controlled direct effects (CDE). We argue that the CATE is often the easiest to identify and estimate, but may not be the most policy relevant target of inference. Moreover, the widespread practice of coarsening distinct policies or implementations into a single indicator further complicates the interpretation of these analyses. Motivated by these limitations, we propose bounding ITEs as an alternative inferential goal, yielding ranges for each state's policy effect under explicit assumptions that quantify deviations from the ideal identifying conditions. These bounds target a well-defined and policy-relevant quantity, the effect for specific states. We develop this approach within a difference-in-differences framework and discuss how sensitivity parameters may be informed using pre-treatment data. Through simulations we demonstrate that bounding state-specific effects can more reliably determine the sign of the ITEs than CATE estimates. We then illustrate this method to examine the effect of the Affordable Care Act Medicaid expansion on high-volume buprenorphine prescribing.
Figures
Reference graph
Works this paper leans on
-
[1]
Adaptive interventions in child and adolescent mental health.Journal of Clinical Child & Adolescent Psychology, 45(4):383–395, 2016
Daniel Almirall and Andrea Chronis-Tuscano. Adaptive interventions in child and adolescent mental health.Journal of Clinical Child & Adolescent Psychology, 45(4):383–395, 2016
2016
-
[2]
Princeton university press, 2009
Joshua D Angrist and J¨ orn-Steffen Pischke.Mostly harmless econometrics: An empiricist’s companion. Princeton university press, 2009
2009
-
[3]
Joseph Antonelli, Max Rubinstein, Denis Agniel, Rosanna Smart, Elizabeth Stuart, Matthew Cefalu, Terry Schell, Joshua Eagan, Elizabeth Stone, Max Griswold, et al. Autoregressive models for panel data causal inference with application to state-level opioid policies.arXiv preprint arXiv:2408.09012, 2024
arXiv 2024
-
[4]
Models as approximations i.Statistical Science, 34(4):523–544, 2019
Andreas Buja, Lawrence Brown, Richard Berk, Edward George, Emil Pitkin, Mikhail Traskin, Kai Zhang, and Linda Zhao. Models as approximations i.Statistical Science, 34(4):523–544, 2019. 27
2019
-
[5]
Difference-in-differences with multiple time periods.Jour- nal of econometrics, 225(2):200–230, 2021
Brantly Callaway and Pedro HC Sant’Anna. Difference-in-differences with multiple time periods.Jour- nal of econometrics, 225(2):200–230, 2021
2021
-
[6]
Double/debiased machine learning for treatment and structural parameters, 2018
Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins. Double/debiased machine learning for treatment and structural parameters, 2018
2018
-
[7]
All medicaid expansions are not created equal: the geography and targeting of the affordable care act
Craig Garthwaite, John A Graves, Tal Gross, Zeynal Karaca, Victoria R Marone, and Matthew J Notowidigdo. All medicaid expansions are not created equal: the geography and targeting of the affordable care act. Technical report, National Bureau of Economic Research, 2019
2019
-
[8]
Cambridge university press, 2007
Andrew Gelman and Jennifer Hill.Data analysis using regression and multilevel/hierarchical models. Cambridge university press, 2007
2007
-
[9]
Phillip Heiler. Heterogeneous treatment effect bounds under sample selection with an application to the effects of social media on political polarization.Journal of Econometrics, 244(1):105856, 2024
2024
-
[10]
Effect or treatment heterogeneity? policy evaluation with aggregated and disaggregated treatments
Phillip Heiler and Michael Knaus. Effect or treatment heterogeneity? policy evaluation with aggregated and disaggregated treatments. 2022
2022
-
[11]
A causal framework for evaluating drivers of policy effect heterogeneity using difference-in-differences.Health Services and Outcomes Research Methodology, pages 1–22, 2025
Gary Hettinger, Youjin Lee, and Nandita Mitra. A causal framework for evaluating drivers of policy effect heterogeneity using difference-in-differences.Health Services and Outcomes Research Methodology, pages 1–22, 2025
2025
-
[12]
Jennifer L. Hill. Bayesian nonparametric modeling for causal inference.Journal of Computational and Graphical Statistics, 20(1):217–240, 2011
2011
-
[13]
Statistics and causal inference.Journal of the American statistical Association, 81(396):945–960, 1986
Paul W Holland. Statistics and causal inference.Journal of the American statistical Association, 81(396):945–960, 1986
1986
-
[14]
Nonparametric estimation of average treatment effects under exogeneity: A review
Guido W Imbens. Nonparametric estimation of average treatment effects under exogeneity: A review. Review of Economics and statistics, 86(1):4–29, 2004
2004
-
[15]
Edward H Kennedy. Semiparametric doubly robust targeted double machine learning: a review.arXiv preprint arXiv:2203.06469, 2022
Pith/arXiv arXiv 2022
-
[16]
Wild bootstrap inference for wildly different cluster sizes
James G MacKinnon and Matthew D Webb. Wild bootstrap inference for wildly different cluster sizes. Journal of Applied Econometrics, 32(2):233–254, 2017
2017
-
[17]
The wild bootstrap for few (treated) clusters.The Econometrics Journal, 21(2):114–135, 2018
James G MacKinnon and Matthew D Webb. The wild bootstrap for few (treated) clusters.The Econometrics Journal, 21(2):114–135, 2018. 28
2018
-
[18]
Randomization inference for difference-in-differences with few treated clusters.Journal of Econometrics, 218(2):435–450, 2020
James G MacKinnon and Matthew D Webb. Randomization inference for difference-in-differences with few treated clusters.Journal of Econometrics, 218(2):435–450, 2020
2020
-
[19]
How do right-to-carry laws affect crime rates? coping with ambiguity using bounded-variation assumptions.Review of Economics and Statistics, 100(2):232–244, 2018
Charles F Manski and John V Pepper. How do right-to-carry laws affect crime rates? coping with ambiguity using bounded-variation assumptions.Review of Economics and Statistics, 100(2):232–244, 2018
2018
-
[20]
Emma E McGinty, Nicholas J Seewald, Sachini Bandara, Magdalena Cerd´ a, Gail L Daumit, Matthew D Eisenberg, Beth Ann Griffin, Tak Igusa, John W Jackson, Alene Kennedy-Hendricks, et al. Scaling interventions to manage chronic disease: innovative methods at the intersection of health policy research and implementation science.Prevention Science, 25(Suppl 1)...
2024
-
[21]
Bayesian hierarchical models.Jama, 320(22):2365–2366, 2018
Anna E McGlothlin and Kert Viele. Bayesian hierarchical models.Jama, 320(22):2365–2366, 2018
2018
-
[22]
Direct and indirect effects
Judea Pearl. Direct and indirect effects. InProbabilistic and causal inference: the works of Judea Pearl, pages 373–392. 2022
2022
-
[23]
A more credible approach to parallel trends.Review of Economic Studies, 90(5):2555–2591, 2023
Ashesh Rambachan and Jonathan Roth. A more credible approach to parallel trends.Review of Economic Studies, 90(5):2555–2591, 2023
2023
-
[24]
Heterogeneous interventional effects with multiple mediators: Semiparametric and nonparametric approaches.Journal of Causal Inference, 11(1):20220070, 2023
Max Rubinstein, Zach Branson, and Edward H Kennedy. Heterogeneous interventional effects with multiple mediators: Semiparametric and nonparametric approaches.Journal of Causal Inference, 11(1):20220070, 2023
2023
-
[25]
Max Rubinstein, Amelia Haviland, and David Choi. Balancing weights for region-level analysis: The effect of medicaid expansion on the uninsurance rate among states that did not expand medicaid.The Annals of Applied Statistics, 17(2):1469–1490, 2023
2023
-
[26]
dabblers
Brendan Saloner, Barbara Andraka Christou, Adam J Gordon, and Bradley D Stein. Article commen- tary: It will end in tiers: A strategy to include “dabblers” in the buprenorphine workforce after the x-waiver.Substance Abuse, 42(2):153–157, 2021
2021
-
[27]
Growing importance of high-volume buprenorphine prescribers in oud treatment: 2009–2018
Megan S Schuler, Andrew W Dick, Adam J Gordon, Brendan Saloner, Rose Kerber, and Bradley D Stein. Growing importance of high-volume buprenorphine prescribers in oud treatment: 2009–2018. Drug and alcohol dependence, 259:111290, 2024
2009
-
[28]
Schuler, Beth Ann Griffin, Magdalena Cerd´ a, Emma E
Megan S. Schuler, Beth Ann Griffin, Magdalena Cerd´ a, Emma E. McGinty, and Elizabeth A. Stuart. Methodological challenges and proposed solutions for evaluating opioid policy effectiveness.Health Services and Outcomes Research Methodology, 21:21–41, 2021. 29
2021
-
[29]
Schuler, Sara E
Megan S. Schuler, Sara E. Heins, Rosanna Smart, Beth Ann Griffin, David Powell, Elizabeth A. Stuart, Bryce Pardo, Sierra Smucker, Stephen W. Patrick, Rosalie Liccardo Pacula, and Bradley D. Stein. The state of the science in opioid policy research.Drug and Alcohol Dependence, 214:108137, 2020
2020
-
[30]
High-volume buprenorphine prescribers: examining state policy contexts.Drug and Alcohol Dependence Reports, page 100406, 2026
Megan S Schuler, Flora Sheng, Brendan Saloner, Adam J Gordon, and Bradley D Stein. High-volume buprenorphine prescribers: examining state policy contexts.Drug and Alcohol Dependence Reports, page 100406, 2026
2026
-
[31]
In- vestigating the complexity of naloxone distribution: Which policies matter for pharmacies and potential recipients.Journal of health economics, 97:102917, 2024
Rosanna Smart, David Powell, Rosalie Liccardo Pacula, Evan Peet, Rahi Abouk, and Corey S Davis. In- vestigating the complexity of naloxone distribution: Which policies matter for pharmacies and potential recipients.Journal of health economics, 97:102917, 2024
2024
-
[32]
Bradley D Stein, Rachel K Landis, Flora Sheng, Brendan Saloner, Adam J Gordon, Mark Sorbero, and Andrew W Dick. Buprenorphine treatment episodes during the first year of covid: a retrospective examination of treatment initiation and retention.Journal of general internal medicine, 38(3):733–737, 2023
2023
-
[33]
Causal inference under multiple versions of treatment
Tyler J VanderWeele and Miguel A Hernan. Causal inference under multiple versions of treatment. Journal of causal inference, 1(1):1–20, 2013
2013
-
[34]
conditional average treatment effects
Brian G Vegetabile. On the distinction between “conditional average treatment effects” (cate) and “individual treatment effects” (ite) under ignorability assumptions.arXiv preprint arXiv:2108.04939, 2021. 30 A Theoretic results A.1 Identification results: treatment coarsening We consider causal identification in the presence of treatment coarsening from a...
Pith/arXiv arXiv 2021
-
[35]
selection- on-observables
To motivate Proposition 1, we make the following assumptions: Assumption 3.a(Consistency). Y= kX m=0 I(M=m)Y(m). Consistency states that the observed values ofYis equal to the potential outcome under the respective treatment. Assumption 3.b(Ignorability). Y(m)⊥M|X, m= 0, . . . , k. Ignorability states thatMis effectively randomized with respect toY(m) con...
-
[36]
fundamental problem of causal inference:
We consider the following model of state-level potential outcomes with respect to both a multi-valued interventionM, some effect modifying covariate of interest covariateX, and a second effect modifierU, where, for simplicity, we have that (X i, Ui) iid ∼ N 0, σ2 x σxu σxu σ2 u . The potential outcomesY i(m, x) are then given by Yi(0, x)...
2013
This paper was first reviewed by deepseek-v4-flash on August 3, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.