Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Under parallel growths, the counterfactual untreated vector of categorical counts for the treated group is identified in closed form, along with its total and shares.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 10:04 UTC pith:I2I6VAPA

load-bearing objection The core 2x2 identification is sound and worth knowing, but the paper overclaims: the abstract promises asymptotics and a pre-trends test that are not in the manuscript, and Theorem 2's partial-identification bounds for the compositional effect are actually wrong. the 3 major comments →

arxiv 2510.11659 v4 pith:I2I6VAPA submitted 2025-10-13 econ.EM

Compositional difference-in-differences

classification econ.EM
keywords difference-in-differencescategorical outcomescompositional data analysisparallel growthscausal inferencerandom utility modelvotingRGGI
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper claims that when an outcome is a vector of counts across unordered categories, the difference-in-differences logic can be repaired by moving from parallel trends in levels to parallel growths in logs: absent treatment, each category's size would grow or shrink by the same proportion in treated and control groups. Under that assumption, the counterfactual vector for the treated group is identified in closed form by multiplying its pre-treatment counts by the control group's category-by-category growth factors, equivalently exp(log q_10 + log q_01 - log q_00). From that vector one reads off both the counterfactual total and the counterfactual shares, so the same framework quantifies effects on total participation and on reallocation across categories. The paper also gives the assumption economic content—parallel evolution of expected utilities and relative preferences in a random-utility model—and geometric content—parallel trajectories in the standard simplex geometry. Applications to early voting and a regional carbon market show the kinds of claims the method supports.

Core claim

The paper's central claim is Theorem 1: under common support and parallel growths, the counterfactual untreated quantity for each category in the treated, post-treatment cell is q_N11(c_k) = exp(log q_N10(c_k) + log q_N01(c_k) - log q_N00(c_k)); the counterfactual total is the sum of these quantities, and the counterfactual shares are the normalized quantities. This makes the counterfactual distribution point-identified in closed form and always a valid probability vector. The paper defines two target parameters: the growth treatment effect on the treated (GTT), a proportional effect on category and total counts, and the compositional treatment effect on the treated (CTT), a normalized vecto

What carries the argument

The load-bearing object is the parallel growths assumption: log q_1,1 - log q_1,0 = log q_0,1 - log q_0,0 componentwise. It is a multiplicative parallel-trends condition that is scale-invariant, keeps counterfactual counts positive, and, through the log-odds transformation, makes the implied share evolution a parallel shift in the simplex under its standard compositional geometry. The assumed equality on log-counts also transfers to equality of changes in expected utilities in a random-utility model, so relative preferences evolve in parallel absent treatment. This single assumption carries all the identification: once it is imposed, the closed-form formula follows by algebra.

Load-bearing premise

The load-bearing premise is that, absent treatment, every outcome category grows or shrinks by the same proportional rate in the treated and control groups; if that shared-log-growth condition fails, the closed-form counterfactual and all treatment-effect estimates built on it are not identified.

What would settle it

Use two or more pre-treatment periods and compute the cross-group log-count gap for each category: log q_1,t(c_k) - log q_0,t(c_k). If that gap drifts or trends across pre-treatment periods in the early-voting or RGGI data, or in a placebo 'treatment' period, parallel growths is contradicted and Theorem 1's counterfactual will be systematically off. A formal pre-trend test of Assumption 2 on either application, or on simulated data generated with heterogeneous growth rates, would settle the claim.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If parallel growths holds in a 2x2 design, counterfactual totals and shares are identified without estimating a full discrete-choice model, and all counterfactual shares lie inside the probability simplex.
  • The same construction extends to staggered adoption either with never-treated units or with not-yet-treated units as controls, giving cohort-by-time counterfactual quantities and totals.
  • When pre-treatment growth differentials are bounded, the closed-form point estimates become sharp bounds on GTT and CTT, so imperfectly parallel trends need not destroy the analysis.
  • A synthetic version reweights control groups to match the treated group's pre-treatment log trajectory, yielding treatment effects on both margins in panel settings.
  • In the reported applications, early voting raises turnout by about 4.4% and the Democratic vote share by 0.92 percentage points, while RGGI's first compliance period reduces total generation by up to 13.6% and moves generation from coal and oil toward gas and nuclear rather than renewables.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A natural extension not pursued in the paper is to apply the same log-ratio construction to count outcomes with zeros or measurement error, adding a trimming or smoothing step to keep the exponential formula defined.
  • Because the share counterfactual equals a softmax of a linear expression, CoDiD estimates can be compared directly with multinomial-logit difference-in-differences coefficients; disagreement between the two would reveal where functional form rather than sampling noise matters.
  • A testable placebo extension is to apply the formula to pairs of pre-treatment periods, treating one as a fake post-period; systematic bias there would show how sensitive the counterfactual is to departures from parallel growths.
  • The breadth of the partial-identification bounds depends on the researcher's choice of weights across pre-treatment periods, so a data-driven weight-selection rule would be a natural complement to the paper's convex-hull relaxation.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes CoDiD, a difference-in-differences method for categorical outcomes. Under a 'parallel growths' assumption (Assumption 2: equal pre/post log-quantity changes across groups), Theorem 1 gives a closed-form counterfactual quantity vector, total, and shares. The paper interprets this assumption in a random-utility model and in Aitchison geometry, and extends the framework to partial identification under relaxed assumptions (Theorem 2), staggered adoption (Theorem 3), and a synthetic control analog. Two empirical applications illustrate the method. The central point-identification algebra is correct, but the advertised partial-identification bounds, asymptotic distribution, and pre-trends test are either absent or incorrect in the current manuscript.

Significance. If limited to Theorem 1 and Proposition 2, the paper makes a useful and clean contribution: log-linear parallel trends for counts yield counterfactual shares that respect the simplex and connect naturally to multinomial logit. The closed-form nature, with no fitted parameters in the identification step, is a genuine strength. However, the additional advertised results—sharp partial identification, joint asymptotic distribution, and a pre-trends test—are not delivered correctly or at all. As it stands, the manuscript's significance is substantially below the abstract's claims, though the central identification formula is worth publishing after a major revision.

major comments (3)
  1. [Section 2.1, Theorem 2] The CTT partial-identification set is not sharp and can exclude the true counterfactual. Under Assumption 3, the identified set for q^N_{1,1} is the box Π_k [b_min(c_k), b_max(c_k)], so the identified set of counterfactual shares is {q/Σq : b_min ≤ q ≤ b_max}. Theorem 2 instead writes s ∈ ℓ^{-1}({r ∈ R^{p-1} : log b_min(c_k) ≤ r_k ≤ log b_max(c_k)}). But r_k = log(s_k/s_p) = log(q_k/q_p), and q_p appears in the denominator. The box constraint on q does not imply log b_min(c_k) ≤ r_k ≤ log b_max(c_k). For p=2 with b_min=(1,2), b_max=(2,4), the true counterfactual share s_1 lies in [0.2, 0.5], whereas the theorem's characterization gives r_1 ∈ [log 1, log 2] = [0, log 2], i.e. s_1 ∈ [1/3, 2/3], which excludes true values such as s_1 = 0.25. Since the abstract explicitly promises 'sharp partial identification bounds', this is a load-bearing error. The set should be characterized directly as
  2. [Abstract / Sections 1–5, Appendix C] The abstract promises 'the joint asymptotic distribution of all treatment-effect estimators under multinomial sampling' and 'a pre-trends test'. Neither appears in the body or appendices. The only inference procedure is a parametric bootstrap (Appendix C), with no theorem on the joint limiting distribution and no formal pre-trends test anywhere. The empirical sections rely on visual inspection of pre-treatment log-trend plots (Figures 6–8, 10). This is a substantial gap between what is advertised and what is delivered. The author should either add the missing theoretical results or revise the abstract to reflect the actual content.
  3. [Section 2.3, synthetic CoDiD] The GTT formula in the synthetic CoDiD extension appears inconsistent with the parallel-growths identifying assumption. Under parallel growths, the counterfactual treated post quantity is q_T,pre × (q_C,post / q_C,pre), so the growth treatment effect should be GTT = q_T,post × q_C,pre / (q_T,pre × q_C,post) − 1. The displayed formula, 'GTT = q_treated,post / (q_treated,pre + q_control,post − q_control,pre) − 1', is an additive DiD on the exp-smoothed quantities, not a proportional-growth effect. If this is a typographical error from equation formatting, it should be corrected; as written, it undermines the synthetic control extension.
minor comments (4)
  1. [Appendix A.2, proof of Proposition 1] Typo in the last display: 'log q^N_{g1}(c_k)−log q^N_{g1}(c_k)' should read 'log q^N_{g1}(c_k)−log q^N_{g0}(c_k)'.
  2. [Section 1.3, Proposition 3 and proof A.4] The proposition states π^N_{1,1} ⊖ π^N_{1,0} = π^N_{0,1} ⊖ π^N_{0,0}, but the proof begins by analyzing π^N_{1,1} ⊖ π^N_{0,1} = π^N_{1,0} ⊖ π^N_{0,0}. These are algebraically equivalent under the log-odds characterization, but the notation should be aligned to avoid confusion.
  3. [Section 1.1.1, equation (1)] The display 'π^N_{1,1} ⊖ π^N_{1,1}' appears to be a typo; the compositional difference for the control group should be π^N_{0,1} ⊖ π^N_{0,0}.
  4. [General] The paper would benefit from a careful proofreading pass for notation and equation formatting; several displays are garbled or have inconsistent subscripts (e.g., Section 2.3, Theorem 2's b_min/b_max definitions).

Circularity Check

1 steps flagged

No significant circularity: Theorem 1 follows by direct algebra from Assumption 2; the only minor tautological step is the RUM 'justification' in Proposition 1, which restates Assumption 2 in utility notation.

specific steps
  1. self definitional [Section 1.2, Proposition 1; Appendix A.2]
    "V N_gt(c_k)=μ^N_gt(c_k)−log(Σ_k e^{μ^N_gt(c_k)}) + log(S^N_gt). ... Proposition 1 (Implication for expected utilities). Under assumptions 1 and the RUM model assumptions, assumption 2 is equivalent to parallel trends of expected utilities: E[U^N_11(c_k)]−E[U^N_10(c_k)] = E[U^N_01(c_k)]−E[U^N_00(c_k)] ∀ c_k∈Ȳ."

    Given q^N_gt(c_k)=π^N_gt(c_k) S^N_gt and π^N_gt(c_k)=e^{μ^N_gt(c_k)}/Σ_j e^{μ^N_gt(c_j)}, the paper's decomposition makes E[U^N_gt(c_k)] = log q^N_gt(c_k)+γ (up to the Gumbel mean). Thus Eq. (6) is exactly Assumption 2 after adding/subtracting the constant γ. The RUM section therefore restates parallel growths as 'parallel trends in expected utilities' rather than providing an independent micro-foundation. This is an interpretive equivalence by construction, not an input to Theorem 1; the main identification result remains direct algebra from Assumption 2.

full rationale

The central identification claim is not circular. Theorem 1 is a one-line algebraic rearrangement of Assumption 2 (parallel growths) followed by normalization; no parameter is fitted to the treated post-treatment outcome and then relabeled as a prediction. The share-space formula (10), the log-odds implication (Proposition 2), the compositional-difference interpretation (Proposition 3), and the Aitchison-geometry reading (Proposition 4) are all algebraic restatements of the same log-translation logic, not independent sources of the result. There is no load-bearing self-citation chain: the paper cites external prior work (Aitchison, Callaway-Sant'Anna, Arkhangelsky et al., Ban-Kédagni, etc.) and does not invoke a uniqueness theorem by the present author to force its choice. The only mildly circular element is Proposition 1's RUM 'justification,' where the expected-utility decomposition is set up so that parallel growths and parallel expected-utility trends are the same equation; this is a non-load-bearing interpretation. Separately, Theorem 2's claimed sharp CTT bounds appear to treat log-quantity bounds as log-odds bounds, which is a correctness/validity concern rather than a circularity concern, so it is not scored here. Score 1 reflects the minor tautological RUM restatement; the paper's main derivation is self-contained.

Axiom & Free-Parameter Ledger

0 free parameters · 6 axioms · 0 invented entities

The central claim is an algebraic consequence of Assumption 2; there are no fitted free parameters. The main assumptions are the parallel growths restriction itself, the implicit positivity required for logs, the random-utility interpretation, and the sampling model used for bootstrap inference.

axioms (6)
  • domain assumption Assumption 1: common support and strictly positive category quantities for all untreated potential outcomes
    Ensures the componentwise logarithm is defined. The proof of Theorem 1 asserts strict positivity 'by Assumption 1', but Assumption 1 as stated only requires common support, not nonzero counts.
  • ad hoc to paper Assumption 2: parallel growths, log qN_1,1 - log qN_1,0 = log qN_0,1 - log qN_0,0 componentwise
    Core identifying assumption of the paper; every counterfactual and treatment-effect estimate follows from it. It is imposed, not formally tested.
  • domain assumption Random utility model with i.i.d. type-1 extreme value errors
    Used in Propositions 1 and 2 to interpret parallel growths as parallel trends in expected utilities. Not needed for the algebraic identification formula in Theorem 1.
  • domain assumption Multinomial sampling model for inference
    Appendix C treats each group-period count vector as Multinomial(S_gt, pi_gt). The validity of the parametric bootstrap confidence intervals is assumed, not proven.
  • ad hoc to paper Assumption 3: post-treatment log-difference lies in the convex hull of weighted pre-treatment log-differences
    Underlies Theorem 2's partial identification bounds. The theorem states sharpness but no sharpness proof is provided.
  • domain assumption Assumptions 4-6: irreversibility and parallel growths to never-treated or not-yet-treated cohorts
    Needed for the staggered-adoption extension in Theorem 3; they are the log-growth analogue of Callaway and Sant'Anna parallel trends.

pith-pipeline@v1.3.0-alltime-deepseek · 28427 in / 13037 out tokens · 110980 ms · 2026-08-04T10:04:33.076290+00:00 · methodology

0 comments
read the original abstract

Many causal questions concern vectors of quantities across mutually exclusive categories, votes by party, employment by status, generation by energy source, where both the shares and the total matter. Standard practice runs a separate linear difference-in-differences (DiD) on each share, which violates the simplex constraint, cannot describe reallocation across categories, and is silent on the total. This paper develops Compositional Difference-in-Differences (CoDiD). For the share margin, parallel trends in log-odds identify the counterfactual composition in closed form, always inside the simplex, and admit two readings: parallel evolution of relative utilities in a random-utility model, and parallel trajectories in the Aitchison geometry of the simplex. Strengthening the assumption to parallel growth in log-counts jointly identifies effects on shares and on the total, and reveals a composition adjustment factor, the ratio of inclusive-value growth across groups, that corrects the aggregation bias of log-total DiD. I derive the joint asymptotic distribution of all treatment-effect estimators under multinomial sampling, provide a pre-trends test, and obtain sharp partial identification bounds when pre-treatment growth differentials are only bounded. Applying CoDiD to early voting in the 2008 U.S.\ presidential election, I find a $4.4\%$ increase in turnout and a $0.92$ percentage-point increase in the Democratic vote share.

Figures

Figures reproduced from arXiv: 2510.11659 by Onil Boussim.

Figure 1
Figure 1. Figure 1: Illustration of how parallel trends can lead to invalid counterfactuals in the 2-dimensional probability simplex. The control group moves from (0.7, 0.2, 0.1) to (0.3, 0.3, 0.4), while applying the same linear shift to the treated group at (0.2, 0.3, 0.5) produces an invalid counterfactual (−0.2, 0.4, 0.8), shown above as ly￾ing outside the probability space. Ideally, one would like to consider a nonlinear… view at source ↗
Figure 2
Figure 2. Figure 2: Visual representation of the DiD identification strategy. The counterfac￾tual outcome for the treated group in the post-treatment period (q N 1,1 ) is constructed by extrapolating the time trend observed in the control group (blue arrows) to the treated group’s pre-treatment level. The causal effect is the difference between the observed treated outcome (q I 1,1 ) and this counterfactual (q N 1,1 ), shown … view at source ↗
Figure 3
Figure 3. Figure 3: Illustration of parallel growths in the simplex. The red curve represents the trajectory of the control group from pre-treatment (π N 0,0 ) to post-treatment (π N 0,1 ), while the blue curve shows the counterfactual trajectory of the treated group from pre-treatment (π N 1,0 ) to post-treatment (π N 1,1 ). Dashed lines indicate the linear translation (parallel growths). evolution of the distributions is co… view at source ↗
Figure 4
Figure 4. Figure 4: Geometric Illustration of parallel trends [PITH_FULL_IMAGE:figures/full_fig_p027_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Geometric illustration of Rank invariance identification. If multiple transformations could map the control group across periods, the counterfactual is no longer unique. In that case, we have identified a set of feasible outcomes, as in the case of Changes-in-Changes (CiC) with a discrete ordered outcome. CoDiD shares this same logic, but for probability mass functions, where the space is the simplex. Here… view at source ↗
Figure 6
Figure 6. Figure 6: log-ratios evolution 1992-2008 (treated vs. control) [PITH_FULL_IMAGE:figures/full_fig_p031_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: shares evolution 1992-2008 (treated vs. control) [PITH_FULL_IMAGE:figures/full_fig_p031_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: quantities evolution 1992-2008 (treated vs. control) My method appears to be the most appropriate in this context, as it relies on the visual assessment of parallel trends between the treatment and control groups. The observed similarity in pre-treatment trajectories provides reassuring evidence that the identifying assumptions are reasonable, thereby strengthening the credibility of the causal analysis. E… view at source ↗
Figure 9
Figure 9. Figure 9: Map of the United States showing the nine RGGI states in purple and six excluded control states (due to market leakage or carbon pricing policies) in dark gray, and the control group consists of the gray states. organized into three-year compliance periods, with the first spanning January 1, 2009, to December 31, 2011. I focus my analysis on this initial period because the parallel growth assumption become… view at source ↗
Figure 10
Figure 10. Figure 10: log-quantity evolution 2000-2011 (treated vs control group) the impact of RGGI over the first three years of implementation (2009–2011), corresponding to the program’s initial compliance period. As shown in the dynamic estimates below, the policy induced a statistically significant and growing reduction in total electricity generation (see figure11). These results show an 8.2% decline in total electricity… view at source ↗
Figure 11
Figure 11. Figure 11: GTT on total electricity production year, increasing to 8.9% in the second year, and reaching 13.6% by the end of the initial compliance period in 2011. The growing effect suggests that RGGI not only reduced gen￾eration in the short term but also triggered accelerating adjustments over time, potentially through coal plant retirements, fuel switching, or improvements in demand-side efficiency. While total … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Compositional Synthetic Controls

    econ.EM 2026-07 conditional novelty 5.0

    For outcomes that are shares summing to one, the paper estimates counterfactuals as weighted geometric means of donor compositions in log-odds space, with weights fit before treatment.

Reference graph

Works this paper leans on

41 extracted references · 4 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Aitchison

    J. Aitchison. The statistical analysis of compositional data. Journal of the Royal Statistical Society: Series B (Methodological), 44 0 (2): 0 139--160, 1982

  2. [2]

    Aitchison

    J. Aitchison. Relative variation diagrams for describing patterns of compositional variability. Mathematical Geology, 22 0 (4): 0 487--511, 1990

  3. [3]

    Aitchison

    J. Aitchison. On criteria for measures of compositional difference. Mathematical Geology, 24 0 (4): 0 365--379, 1992

  4. [4]

    Aitchison

    J. Aitchison. Simplicial inference. In M. A. G. Viana and D. S. P. Richards, editors, Algebraic Methods in Statistics and Probability, volume 287 of Contemporary Mathematics, pages 1--22. American Mathematical Society, Providence, Rhode Island, 2002

  5. [5]

    Arkhangelsky, S

    D. Arkhangelsky, S. Athey, D. A. Hirshberg, G. W. Imbens, and S. Wager. Synthetic difference-in-differences. American Economic Review, 111 0 (12): 0 4088--4118, 2021

  6. [6]

    K. F. Arnold, L. Berrie, P. W. Tennant, and M. S. Gilthorpe. A causal inference perspective on the analysis of compositional data. International journal of epidemiology, 49 0 (4): 0 1307--1313, 2020

  7. [7]

    Athey and G

    S. Athey and G. W. Imbens. Identification and inference in nonlinear difference-in-differences models. Econometrica, 74 0 (2): 0 431--497, 2006

  8. [8]

    Baker, B

    A. Baker, B. Callaway, S. Cunningham, A. Goodman-Bacon, and P. H. Sant'Anna. Difference-in-differences designs: A practitioner's guide. arXiv preprint arXiv:2503.13323, 2025

  9. [9]

    Ban and D

    K. Ban and D. K \'e dagni. Robust difference-in-differences models. arXiv preprint arXiv:2211.06710, 2022

  10. [10]

    Barceló-Vidal, J

    C. Barceló-Vidal, J. A. Martín-Fernández, and V. Pawlowsky-Glahn. Mathematical foundations of compositional data analysis. In G. Ross, editor, Proceedings of the Sixth Annual Conference of the International Association for Mathematical Geology, Cancun, Mexico, 2001. CD-ROM

  11. [11]

    S. T. Berry, C. Cox, and P. Haile. Selective turnout, voting policy, and partisan bias: Evidence from multi-level data. Technical report, National Bureau of Economic Research, 2025

  12. [12]

    Billheimer, P

    D. Billheimer, P. Guttorp, and W. F. Fagan. Statistical interpretation of species composition. Journal of the American Statistical Association, 96 0 (456): 0 1205--1214, 2001

  13. [13]

    Bonhomme and U

    S. Bonhomme and U. Sauder. Recovering distributions in difference-in-differences models: A comparison of selective and comprehensive schooling. The Review of Economics and Statistics, 93 0 (2): 0 479--494, 2011. ISSN 00346535, 15309142

  14. [14]

    Cafri, W

    G. Cafri, W. Wang, P. H. Chan, and P. C. Austin. A review and empirical comparison of causal inference methods for clustered observational data with application to the evaluation of the effectiveness of medical devices. Statistical Methods in Medical Research, 28 0 (10-11): 0 3142--3162, 2019

  15. [15]

    Callaway and T

    B. Callaway and T. Li. Quantile treatment effects in difference in differences models with panel data. Quantitative Economics, 10 0 (4): 0 1579--1618, 2019

  16. [16]

    Callaway and P

    B. Callaway and P. H. Sant’Anna. Difference-in-differences with multiple time periods. Journal of Econometrics, 2020. ISSN 0304-4076

  17. [17]

    Callaway and P

    B. Callaway and P. H. Sant’Anna. Difference-in-differences with multiple time periods. Journal of econometrics, 225 0 (2): 0 200--230, 2021

  18. [18]

    Callaway, T

    B. Callaway, T. Li, and T. Oka. Quantile treatment effects in difference in differences models under dependence restrictions and with only two time periods . Journal of Econometrics, 206 0 (2): 0 395--413, 2018

  19. [19]

    De Chaisemartin and X

    C. De Chaisemartin and X. d’Haultfoeuille. Two-way fixed effects and differences-in-differences with heterogeneous treatment effects: A survey. The econometrics journal, 26 0 (3): 0 C1--C30, 2023

  20. [20]

    J. J. Egozcue, V. Pawlowsky-Glahn, G. Mateu-Figueras, and C. Barcelo-Vidal. Isometric logratio transformations for compositional data analysis. Mathematical geology, 35 0 (3): 0 279--300, 2003

  21. [21]

    Ghanem, D

    D. Ghanem, D. K \'e dagni, and I. Mourifi \'e . Evaluating the impact of regulatory policies on social welfare in difference-in-difference settings. arXiv preprint arXiv:2306.04494, 2023

  22. [22]

    J. A. Graves, C. Fry, J. M. McWilliams, and L. A. Hatfield. Difference-in-differences for categorical outcomes. Health Services Research, 57 0 (3): 0 681--692, 2022

  23. [23]

    M. J. Hamilton. Chapter 1 lie groups and lie algebras: Basic concepts. In Mathematical Gauge Theory: With Applications to the Standard Model of Particle Physics, pages 3--82. Springer, 2017

  24. [24]

    Havnes and M

    T. Havnes and M. Mogstad. Is universal child care leveling the playing field? Journal of Public Economics, 127: 0 100--114, 2015. ISSN 0047-2727. The Nordic Model

  25. [25]

    J. L. Horowitz. Bootstrap methods in econometrics. Annual Review of Economics, 11 0 (1): 0 193--224, 2019

  26. [26]

    M. Lechner. The estimation of causal effects by difference-in-difference methods. Foundations and Trends® in Econometrics, 4 0 (3): 0 165--224, 2011. ISSN 1551-3076

  27. [27]

    C. F. Manski and J. V. Pepper. How do right-to-carry laws affect crime rates? coping with ambiguity using bounded-variation assumptions. Review of Economics and Statistics, 100 0 (2): 0 232--244, 2018

  28. [28]

    McFadden

    D. McFadden. Conditional logit analysis of qualitative choice behavior. 1972

  29. [29]

    McFadden

    D. McFadden. The measurement of urban travel demand. Journal of public economics, 3 0 (4): 0 303--328, 1974

  30. [30]

    McFadden

    D. McFadden. Modelling the choice of residential location. 1977

  31. [31]

    B. D. Meyer, W. K. Viscusi, and D. Durbin. Workers' compensation and injury duration: evidence from a natural experiment, 1990

  32. [32]

    difference-in-differences

    P. A. Puhani. The treatment effect, the cross difference, and the interaction term in nonlinear “difference-in-differences” models. Economics Letters, 115 0 (1): 0 85--87, 2012

  33. [33]

    Rambachan and J

    A. Rambachan and J. Roth. An honest approach to parallel trends. Working Paper, 2020

  34. [34]

    Roth and P

    J. Roth and P. Sant'Anna. When is parallel trends sensitive to functional form? Unpublished Manuscript, 2021

  35. [35]

    J. Roth, P. H. Sant’Anna, A. Bilinski, and J. Poe. What’s trending in difference-in-differences? a synthesis of the recent econometrics literature. Journal of Econometrics, 235 0 (2): 0 2218--2244, 2023

  36. [36]

    E. J. Tchetgen Tchetgen, C. Park, and D. B. Richardson. Universal difference-in-differences for causal inference in epidemiology. Epidemiology, 35 0 (1), 2024

  37. [37]

    Torous, F

    W. Torous, F. Gunsilius, and P. Rigollet. An optimal transport approach to causal inference. arXiv preprint arXiv:2108.05858, 2021

  38. [38]

    K. E. Train. Discrete choice methods with simulation. Cambridge university press, 2009

  39. [39]

    J. M. Wooldridge. Simple approaches to nonlinear difference-in-differences with panel data. The Econometrics Journal, 26 0 (3): 0 C31--C66, 2023

  40. [40]

    J. Yan. The impact of climate policy on fossil fuel consumption: Evidence from the regional greenhouse gas initiative (rggi). Energy Economics, 100: 0 105333, 2021

  41. [41]

    Y. Zhou, D. Kurisu, T. Otsu, and H.-G. M \"u ller. Geodesic difference-in-differences. arXiv preprint arXiv:2501.17436, 2025