Pith. sign in

REVIEW 3 major objections 5 minor 17 references

A Distributional Perspective on Pearl's Causal Hierarchy: From Marginal to Joint and Individualized Potential Outcomes

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read This paper recasts Pearl's causal hierarchy in potential-outcomes terms, sorting causal estimands by whether they require marginal, joint, or individual counterfactual information.

desk verdict Useful systematic classification of causal estimands by potential-outcome information, but the ATT placement in layer 2 is a real error that contradicts the SCM-based hierarchy the paper itself cites. read the letter →

arxiv 2601.20405 v2 pith:K5KGEY4B submitted 2026-01-28 stat.OT

classification stat.OT MSC 62D20
keywords causalhierarchypotentialoutcomescounterfactualsinterventionlayerjointdistributionofindividualizedtreatmenteffectsestimandsidentificationassumptions
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Pearl's causal hierarchy sorts causal questions into association, intervention, and counterfactuals. This paper argues that the boundary between the intervention and counterfactual layers can be drawn precisely in potential-outcomes language: an estimand belongs to the intervention layer exactly when it depends only on the marginal distributions of potential outcomes, and to the counterfactual layer when it depends on their joint distribution, on nested outcomes across different interventions, or on individual-level counterfactual outcomes. On this view, randomization identifies the intervention layer, while counterfactual estimands need extra assumptions about dependence between potential outcomes. The paper applies this criterion across a broad range of estimands, including benefit and harm rates, persuasion, principal effects, fairness metrics, and mediation, and uses it to explain why individual-level decisions cannot rest on conditional average treatment effects alone. The practical payoff is a map from a scientific question to the estimand, the probabilistic object it requires, and the assumptions needed for identification.

What carries the argument

The load-bearing object is the pair of probabilistic distinctions the paper draws: the marginal distributions of potential outcomes versus their joint distribution, and population-level joint quantities versus individual realized counterfactual outcomes. The classification rule is that an estimand belongs to the intervention layer exactly when it can be written as a functional of marginal distributions under single interventions, and to the counterfactual layer when it requires the joint distribution, nested cross-world outcomes, or individual-level outcomes. Because each unit realizes only one potential outcome, the joint distribution is never directly observed, so every identification strategy for a counterfactual-layer estimand must add an assumption that pins down the dependence between potential outcomes; the paper catalogs these assumptions and shows how they achieve point identification or partial identification.

What would settle it

Construct any causal estimand the paper assigns to Layer 3 and show, for some data-generating process, that it is point-identified from the marginal potential-outcome distributions alone under randomization — or symmetrically, an estimand assigned to Layer 2 whose identified set depends on the joint distribution. A concrete test: in a binary-treatment, binary-outcome randomized trial, compute the sharp identified set for the persuasion rate $P(Y(1)=1\mid Y(0)=0)$ from the marginals $P(Y(1))$ and $P(Y(0))$ alone; if this set is ever a single point without any dependence assumption, the claimed boundary collapses.

Watch

Extended reading notes

Core claim

The central claim is that Pearl's three-layer causal hierarchy can be operationalized at the level of estimands by a distributional criterion. The intervention layer consists of estimands that are functionals of the marginal potential-outcome distributions $P(Y(a)\mid X)$; the counterfactual layer consists of estimands that are functionals of the joint distribution $P(Y(0),Y(1)\mid X)$, involve nested potential outcomes such as $Y(1,M(0))$, or target individual realized counterfactual outcomes. The paper splits the counterfactual layer into cross-world queries and individual-level queries, and argues that this ordering tracks identifiability demands: randomization identifies marginals; monotonicity, association parameters, copula restrictions, rank preservation, or data fusion are needed for the joint; and individual outcomes require a deterministic view of counterfactuals plus an SCM-based abduction-action-prediction step or rank preservation, with conformal inference providing prediction intervals. This correspondence with Pearl's hierarchy is illustrated in Figure 2 and applied systematically in Examples 1–11.

Load-bearing premise

The classification assumes that Pearl's layer distinctions coincide exactly with whether an estimand depends on marginal, joint, or individual-level potential-outcome information; the paper illustrates this alignment with examples and Figure 2 but does not prove exhaustiveness from the formal SCM definitions.

Editorial extensions

If this is right

  • Researchers can use the marginal/joint/individual criterion to check whether a chosen estimand matches the scientific question, rather than settling for a convenient proxy.
  • Because randomization identifies only marginal potential-outcome distributions, any counterfactual-layer conclusion drawn from a randomized trial is valid only under an additional, often implicit assumption about the dependence between potential outcomes.
  • Individual-level treatment recommendations cannot be justified by conditional average treatment effects alone; harm rates, benefit rates, or prediction intervals for the individual treatment effect are the relevant third-layer quantities.
  • Mediation estimands such as natural direct and indirect effects are inherently cross-world, so mediation analyses require assumptions beyond those identifying average treatment effects.
  • The framework clarifies when proposed identifiability assumptions are sufficient or overly restrictive for a given estimand, as summarized in the paper's Table 2.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A corollary the paper does not fully spell out is an 'information hierarchy' for causal functionals: each move up a layer must add exactly one new piece of dependence information, so the framework could be used to audit causal claims for hidden assumptions.
  • The same marginal-versus-joint lens could be extended to dynamic treatment regimes, where nested potential outcomes under sequences of interventions may form additional sublayers beyond the two the paper distinguishes.
  • The asserted equivalence between the potential-outcome layers and Pearl's SCM hierarchy is testable: if a formal SCM layer-3 query could not be translated into potential-outcome notation, the two hierarchies would not be isomorphic.
  • A practical extension would be an automated layer checker: given a symbolic expression for a causal estimand, detect whether it contains cross-world or individual-level terms and assign the corresponding layer.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper proposes a potential-outcomes interpretation of Pearl's causal hierarchy. The central rule (Section 3.1) classifies an estimand as second-layer if it is a functional of marginal potential-outcome distributions, and third-layer if it requires the joint distribution of potential outcomes, nested/cross-world quantities, or individual-level counterfactuals. The authors apply this rule to a broad set of estimands (ATE, ATT, QTE, DRF, PN, TBR, persuasion rate, ITE distribution, principal causal effects, mediation, and fairness metrics) and review identification strategies such as monotonicity, association parameters, copulas, rank preservation, and conformal inference. The paper is written as a perspective/survey rather than as a new technical result.

Significance. If the proposed classification were valid, it would provide a practical and useful map from scientific questions to the probabilistic objects required for identification, organizing a large and fragmented literature. The paper's strengths are its broad coverage, its many explicit examples, and its systematic review of identification strategies for joint and individual-level estimands. However, the central equivalence between the proposed classification and Pearl's SCM-based hierarchy is not established, and there is a concrete counterexample (ATT) where the paper's rule contradicts the SCM hierarchy it cites. Because this equivalence is the paper's main conceptual contribution, the current version cannot be accepted as a faithful recasting of Pearl's hierarchy.

major comments (3)
  1. [Section 3.1, Example 1] The paper classifies ATT = E[Y(1)-Y(0)|A=1] as a second-layer estimand because it is written in terms of marginals conditioned on observed treatment. Under the SCM hierarchy cited in Section 2.2, however, E[Y(0)|A=1] is a counterfactual conditional: it conditions on the endogenous treatment value A, and the interventional distributions P(Y(0)) and P(Y(1)) do not determine it in general. For example, in the SCM with U~Bernoulli(0.5), A=U, and Y=A+U+epsilon, E[Y(0)] = 0.5+E[epsilon] but E[Y(0)|A=1] = 1+E[epsilon]; the two interventional marginals are the same across models that differ in this conditional, so layer-2 information is insufficient for ATT. Pearl's hierarchy therefore places ATT at layer 3, not layer 2. This is not a terminological quibble: ATT is the first worked example, and the parenthetical definition of 'marginal' is what drives the misclassification. The proposed criterion is not equivalent to Pearl's hierarchy as stated.
  2. [Section 3.1, Figure 2] The second-layer definition permits conditioning on observed treatment and covariates, but Figure 2 displays only P(Y(1)|X) and P(Y(0)|X), not P(Y(a)|A=1). This inconsistency matters because Pearl's hierarchy distinguishes pre-treatment covariates X, which may be conditioned on at layer 2, from post-treatment variables such as A, conditioning on which is a counterfactual (abduction) operation. Collapsing X and A into a single notion of 'marginal' produces the ATT error and leaves the classification rule ambiguous for other conditional estimands. The paper should state formally which variables may be conditioned on at each layer; if conditioning on A is moved to layer 3, then ATT and similar conditional-on-treatment estimands must be reclassified, and the 'marginal vs joint' dichotomy needs to be replaced by a more precise criterion.
  3. [Section 3.1 (general framework)] The claimed correspondence with Pearl's hierarchy is asserted rather than derived. Section 3.1 and Figure 2 present the classification as the definition of the layers, but no proof or formal argument establishes that membership in Pearl's layers is equivalently characterized by whether an estimand depends on marginal, joint, or individual-level potential outcomes. The ATT counterexample shows that this is not merely a missing proof: the asserted equivalence is false as stated. The authors should either revise the classification rule and re-derive the layer assignments, or explicitly present their taxonomy as a distinct potential-outcomes hierarchy that is related to, but not identical with, Pearl's.
minor comments (5)
  1. [Table 3] The comparison table uses symbols such as '/' and ',' in the 'Rank Preservation' and 'Conformal Inference' columns without explaining their meaning; the comparison would be clearer with explicit statements of whether each method has weaker conditions, point identification, and generalizability.
  2. [Section 3.3.2] The sentence 'a Gaussian copula assumes that (Y(1),Y(0)) follows a joint Gaussian distribution' is imprecise; a Gaussian copula with arbitrary marginal distributions does not imply joint normality of the potential outcomes.
  3. [Example 10] Counterfactual parity is stated as an equality constraint, P(S(1)=1)=P(S(0)=1), rather than as an estimand; the intended estimand or target of inference should be clarified.
  4. [Table 1] The notation P(Y(a)|A=a',Y(a')) in the layer-3 row of Table 1 is not defined in the text; it should be defined or cross-referenced to the SCM notation in Section 2.2.
  5. [Throughout] There are several encoding artifacts in the references and text (e.g., 'Hern´ an', '¤ects', missing spaces in citations); these should be cleaned before publication.

Circularity Check

1 steps flagged · score 6.0 of 10

ATT is assigned to layer 2 only because the paper's definition of 'marginal' includes conditioning on treatment, so the claimed correspondence to Pearl's hierarchy is self-definitional for this estimand.

  1. self definitional [Introduction, 'Potential Outcomes Perspective' box; Example 1 in Section 3.2.]
    "Second layer: Intervention. A causal estimand that depends only on the marginal distributions of potential outcomes (conditional on observed treatment and covariates) belongs to the second layer of causation. ... Example 1 (Binary Treatment, Second Layer). ... ATT =E[Y(1)−Y(0)|A= 1]."

    By the paper's own Table 1, Pearl's third layer includes queries of the form P(Y(a)|A=a', Y(a')); the second term of ATT, E[Y(0)|A=1], is such a counterfactual conditional because Y(0) is unobserved for units with A=1 and conditioning on A=1 requires the joint dependence between A and Y(0), not merely the interventional marginals P(Y|do(A=0)) and P(Y|do(A=1)). Figure 2 lists P(Y(a)|X), not P(Y(a)|A=1), as the second-layer objects. The Introduction's definition, however, inserts '(conditional on observed treatment and covariates)' into 'marginal distributions of potential outcomes', thereby converting P(Y(0)|A=1) into a 'marginal' quantity.

full rationale

No fitted-input or prediction-style circularity is present: the paper fits no parameters and derives no numerical predictions; its self-citations (e.g., Wu et al. 2024a, 2025a; Wang et al. 2025b) are used to support identification results for joint distributions and counterfactual outcomes, and those results do not presuppose the paper's layer taxonomy. The central circularity is definitional rather than statistical: the paper asserts that layer membership is governed by whether an estimand depends on marginal, joint, or individual-level potential-outcome information, but it defines 'marginal' to include distributions conditional on the observed treatment. That single broadening lets Example 1 classify ATT as a second-layer estimand even though Pearl's hierarchy, as quoted in the paper's own Table 1, treats P(Y(0)|A=1) as a counterfactual layer-3 query. The claimed correspondence is therefore partly imposed by definition, not derived. Score 6 reflects partial circularity: the ATT classification reduces to the definition, while other classifications (ATE, QTE, DTE, probability of causation, mediation effects) follow the stated marginal/joint distinction and remain meaningful.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The paper's framework rests on standard causal-inference assumptions (SUTVA, ignorability, overlap), a deterministic view of counterfactuals for the individual-level sublayer, and its own definitional classification rule. The classification rule is the main nonstandard premise: it asserts an equivalence between Pearl's layers and the marginal/joint/individual information ordering of potential outcomes, supported by examples rather than a formal proof.

assumptions (6)
  • domain assumption Stable unit treatment value assumption (SUTVA): no interference and no multiple versions of treatment.
    Invoked in Section 2.1 to justify defining the potential outcome Y(a) and the consistency relation Y = Y(A).
  • domain assumption Well-defined potential outcomes Y(a) for all a in the support of A.
    Section 2.1 defines potential outcomes prior to measurement; existence of all counterfactual versions is presupposed.
  • domain assumption Ignorability and overlap conditions for second-layer identification.
    Section 3.2 states A perpendicular to Y(a) and 0 < P(A=1) < 1 as the standard identification conditions; these are assumed when the paper says randomization identifies marginals.
  • ad hoc to paper Assumption 1 (deterministic viewpoint): each individual's potential outcome Y(a) is a fixed quantity.
    Section 3.4.1 introduces this explicitly to make individual-level counterfactuals well-defined, acknowledging the stochastic alternative (Dawid 2000).
  • ad hoc to paper The classification rule itself: layer membership is determined by whether an estimand depends on marginal, joint/nested, or individual-level potential outcome information.
    Section 3.1 states this as the paper's perspective; it is a definitional mapping asserted to correspond to Pearl's hierarchy rather than derived.
  • domain assumption Structural causal model with mutually independent exogenous variables for the three-step procedure.
    Section 2.2 defines SCM with independent U; Section 3.4.2 relies on this for abduction-action-prediction.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Distributional Perspective on Pearl's Causal Hierarchy: From Marginal to Joint and Individualized Potential Outcomes." pith.science (2026). https://pith.science/paper/K5KGEY4B

@misc{pith2026260120405,
  author       = {Pith},
  title        = {Pith review of: A Distributional Perspective on Pearl's Causal Hierarchy: From Marginal to Joint and Individualized Potential Outcomes},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/K5KGEY4B}},
  note         = {Machine review of arXiv:2601.20405}
}
read the original abstract

Pearl's causal hierarchy is a foundational lens for formulating causal questions and is most often discussed within the framework of structural causal models. We recast the hierarchy in potential outcomes language and make its information ordering operational at the level of causal estimands. Specifically, we classify an estimand according to whether it requires marginal potential outcome distributions, their joint distribution or nested cross-world quantities, or individual-level counterfactual outcomes. We apply this criterion systematically to a broad range of estimands, including cases whose classification depends on the formulation of the scientific question. We then clarify what additional assumptions are needed when an estimand depends not only on the marginal distributions of potential outcomes, but also on their unobserved joint distribution. In particular, randomization identifies the marginals, whereas monotonicity, copula restrictions, rank preservation, and partial identification restrict or characterize the remaining uncertainty about the joint distribution. The resulting framework provides a practical map from a scientific question to an estimand, the probabilistic object it requires, and the assumptions needed for identification.

Figures

Figures reproduced from arXiv: 2601.20405 by the authors.

Figure 1
Figure 1. Observed data for binary treatment, where [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. The proposed potential outcomes perspective on Pearl’s causal hierarchy. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

17 extracted references · 6 canonical work pages

  1. [4]

    Marginal causal effect estimation with continuous instrumental variables.arXiv preprint arXiv:2510.14368,

    Mei Dong, Lin Liu, Dingke Tang, Geoffrey Liu, Wei Xu, and Linbo Wang. Marginal causal effect estimation with continuous instrumental variables.arXiv preprint arXiv:2510.14368,

  2. [7]

    Learning the effect of persuasion via difference-in-differences.arXiv preprint arXiv:2410.14871,

    Sung Jae Jun and Sokbae Lee. Learning the effect of persuasion via difference-in-differences.arXiv preprint arXiv:2410.14871,

  3. [13]

    Complete identification methods for the causal hierarchy.Journal of Machine Learning Research, 9:1941–1979,

    21 Ilya Shpitser and Judea Pearl. Complete identification methods for the causal hierarchy.Journal of Machine Learning Research, 9:1941–1979,

  4. [14]

    Semiparametric principal stratification analysis beyond monotonicity.arXiv preprint arXiv:2501.17514,

    Jiaqi Tong, Brennan Kahan, Michael O Harhay, and Fan Li. Semiparametric principal stratification analysis beyond monotonicity.arXiv preprint arXiv:2501.17514,

  5. [15]

    Counterfactually fair reinforcement learning via sequential data preprocessing.arXiv preprint arXiv:2501.06366, 2025a

    Jitao Wang, Chengchun Shi, John D Piette, Joshua R Loftus, Donglin Zeng, and Zhenke Wu. Counterfactually fair reinforcement learning via sequential data preprocessing.arXiv preprint arXiv:2501.06366, 2025a. Linbo Wang and Eric Tchetgen Tchetgen. Bounded, efficient and multiply robust estimation of average treatment effects using instrumental variables.Jou...

  6. [17]

    Quantifying individual risk for binary outcome

    Peng Wu, Peng Ding, Zhi Geng, and Yue Liu. Quantifying individual risk for binary outcome. arXiv preprint arXiv:2402.10537, 2024a. Peng Wu, Ziyu Shen, Feng Xie, Zhongyao Wang, Chunchen Liu, and Yan Zeng. Policy learning for balancing short-term and long-term rewards. InInternational Conference on Machine Learning, 2024b. 22 Peng Wu, Haoxuan Li, Chunyuan Z...

  7. [1990]

    Regression adjustment for estimating distributional treatment effects in randomized controlled trials.arXiv preprint arXiv:2407.14074,

    20 Tatsushi Oka, Shota Yasui, Yuta Hayakawa, and Undral Byambadalai. Regression adjustment for estimating distributional treatment effects in randomized controlled trials.arXiv preprint arXiv:2407.14074,

  8. [2005]

    Semiparametric Estimation of Long-Term Treatment Effects

    Jiafeng Chen and David M Ritzwoller. Semiparametric estimation of long-term treatment effects. arXiv preprint arXiv:2107.14405,

Show all 17 references
  1. [2006]

    Identification and estimation of joint potential outcome distribu- tions from a single study.arXiv preprint arXiv:2509.20506,

    Zach Shahn and David Madigan. Identification and estimation of joint potential outcome distribu- tions from a single study.arXiv preprint arXiv:2509.20506,

  2. [2011]

    Cross-world assumption and refining prediction intervals for individual treatment effects.arXiv:2507.12581,

    Juraj Bodik, Yaxuan Huang, and Bin Yu. Cross-world assumption and refining prediction intervals for individual treatment effects.arXiv:2507.12581,

  3. [2013]

    Distributional treatment effect with latent rank invariance.arXiv:2403.18503,

    Myungkou Shin. Distributional treatment effect with latent rank invariance.arXiv:2403.18503,

  4. [2017]

    Identifying the distribution of treatment effects under support restrictions.arXiv preprint arXiv:1410.5885,

    Ju Hyun Kim. Identifying the distribution of treatment effects under support restrictions.arXiv preprint arXiv:1410.5885,

  5. [2018]

    Causal analysis of ordinal treatments and binary outcomes under truncation by death.Journal of the Royal Statistical Society Series B: Statistical Methodology, 79(3):719–735, 2017a

    Linbo Wang, Thomas S Richardson, and Xiao-Hua Zhou. Causal analysis of ordinal treatments and binary outcomes under truncation by death.Journal of the Royal Statistical Society Series B: Statistical Methodology, 79(3):719–735, 2017a. Linbo Wang, Xiao-Hua Zhou, and Thomas S Ric...

  6. [2022]

    A first course in causal inference.arXiv:2305.18793,

    16 Peng Ding. A first course in causal inference.arXiv:2305.18793,

  7. [2023]

    Identifying the effect of persuasion.Journal of Political Economy, 131(8):2032–2058,

    Sung Jae Jun and Sokbae Lee. Identifying the effect of persuasion.Journal of Political Economy, 131(8):2032–2058,

  8. [2024]

    Assessing heterogeneity of treatment effects.arXiv:2306.15048,

    Tetsuya Kaji and Jianfei Cao. Assessing heterogeneity of treatment effects.arXiv:2306.15048,

  9. [2025]

    From probability to counterfactuals: the increasing complexity of satisfiability in Pearl’s causal hierarchy.arXiv preprint arXiv:2405.07373,

    Julian D¨ orfler, Benito van der Zander, Markus Bl¨ aser, and Maciej Liskiewicz. From probability to counterfactuals: the increasing complexity of satisfiability in Pearl’s causal hierarchy.arXiv preprint arXiv:2405.07373,

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.