REVIEW 2 major objections 5 minor 12 references
Moments of Causal Effects
T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This paper establishes that, under exogeneity and monotonicity, every moment of the individual causal effect $Y_1-Y_0$ is identified from observational data as an integral functional of the conditional CDFs of $Y$ given $X$.
desk verdict A genuinely useful toolkit for higher moments of causal effects, held back by a proof-completeness gap and an overstated convergence claim; the main identification is right, but it needs a revision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the integral decomposition of Lemma 1: $(Y_1-Y_0)^m = \int_{\Omega_Y^m} I(Y_0<y_1\le Y_1,\dots,Y_0<y_m\le Y_1)\,dy_1\cdots dy_m + (-1)^m \int_{\Omega_Y^m} I(Y_1<y_1\le Y_0,\dots,Y_1<y_m\le Y_0)\,dy_1\cdots dy_m$. Taking expectations turns these into joint potential-outcome probabilities of the form appearing in probabilities of causation for continuous outcomes. The actual identification step is imported from Theorem 5.2 of Kawakami et al. [2024a], which supplies the conditional-CDF formula for such joint probabilities under exogeneity and monotonicity; this paper's theorems substitute that formula into the decomposed moment integrals. For the bounds, the Fréchet inequalities bound the same joint probabilities by sums and minima of observed conditional CDFs, yielding Theorems 2 and 4.
What would settle it
Simulate a known structural causal model satisfying exogeneity and monotonicity (for instance $Y=a(X)+b(X)U$ with $U$ standard normal and $X$ independent Bernoulli), compute the true $\mu^{(m)}=E[(Y_1-Y_0)^m]$ by Monte Carlo over $U$, and compare it with $\sigma^{(m)}$ evaluated from the induced observational conditional CDFs; any systematic mismatch for some monotone $b(X)$ would refute Theorem 1. A second check is to inspect whether the cited source theorem requires an additional monotonicity condition and, if so, to construct a model satisfying the paper's assumptions but violating that condition and test the formula numerically.
Extended reading notes
Core claim
The paper's core discovery is a closed-form identification formula for the $m$-th moment $\mu^{(m)}=E[(Y_1-Y_0)^m]$ of the individual causal effect. Theorem 1 states that under exogeneity and monotonicity, $\mu^{(m)}=\sigma^{(m)}$, where $\sigma^{(m)}$ is an $m$-fold integral built from the two conditional CDFs $P(Y<\cdot|X=0)$ and $P(Y<\cdot|X=1)$ through min/max and positive-part operations. The argument decomposes $(Y_1-Y_0)^m$ into integrals of indicators over the interval between the potential outcomes, so the moment is a sum of probabilities like $P(Y_0<y_1\le Y_1,\dots,Y_0<y_m\le Y_1)$; those joint probabilities are then identified from the conditional CDFs. Theorem 3 extends the same mechanism to product moments $\rho_{i,j;k,h}=E[(Y_i-Y_j)(Y_k-Y_h)]$, and Theorems 2 and 4 bound all these quantities by Fréchet inequalities when monotonicity is dropped. The result is that the whole shape of the distribution of causal effects, not just its mean, is recoverable from observational data in closed form.
Load-bearing premise
The load-bearing premise is that the separate theorem the paper cites for identifying joint potential-outcome probabilities holds under exactly the exogeneity and monotonicity assumptions stated here, with no additional hidden condition needed.
Editorial extensions
If this is right
- The variance, skewness, and kurtosis of individual causal effects become estimable from ordinary observational data, not just the average effect.
- For multi-valued treatments, the covariance and correlation between two causal effects (for example consecutive dose increments) are identifiable, revealing whether patients who benefit from one change tend to benefit from the next.
- When monotonicity is implausible, every moment still has computable bounds from exogeneity alone, with the upper bound sharp for even moments.
- Plug-in estimators using empirical conditional CDFs and Monte Carlo integration are consistent, so the formulas translate directly into practice.
- The $m=1$ case recovers the average causal effect and requires no monotonicity, situating the result as a strict generalization of standard ACE identification.
Reading between the lines
- Beyond the paper: the same identification route could in principle be inverted to recover the full cumulative distribution function of $Y_1-Y_0$, since under boundedness the moments determine the distribution; this would give a nonparametric answer to 'how are causal effects distributed?' without any parametric assumption.
- A testable extension: in a randomized crossover trial where both potential outcomes are measured on each subject, the plug-in estimator of the variance of causal effects can be compared directly with the sample variance of observed individual differences; agreement would empirically validate the monotonicity-based identification.
- An editor's guess: the bounds in the paper are quite wide in the cholesterol application, so the practical usefulness of this line will hinge on narrowing them, perhaps by adding covariate information as has been done for probabilities of causation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies estimation of moments and product moments of individual causal effects Y1−Y0 under a structural causal model with binary or multi-valued treatment. Under exogeneity and a monotonicity assumption on the outcome function, Theorem 1 expresses the m-th moment E[(Y1−Y0)^m] as a closed-form integral functional of the conditional CDFs of Y given X=0 and X=1 (Eq. 6); analogous results are given for central moments, product moments, covariance, and correlation, with Fréchet-inequality bounds under exogeneity alone. The paper includes finite-sample simulations and an application to a cholesterol-reduction dataset.
Significance. If the identification theorems hold, the main result is valuable: it shows that all moments of the individual causal effect, not just the ACE, are nonparametrically identified under monotonicity plus exogeneity, with a parameter-free formula involving only conditional CDFs. The bounding results under exogeneity alone are also useful. The paper's main formula is specific and falsifiable, and the simulations include ground-truth comparisons. However, the proof currently depends on an external theorem with unstated conditions, and the consistency proof contains a rate error; both are fixable, but they are load-bearing for the paper's claims.
major comments (2)
- [Appendix A, proof of Theorem 1] The identification of P(Y0<y1≤Y1,...,Y0<ym≤Y1) is imported from Theorem 5.2 of Kawakami et al. (2024a) without stating that theorem's conditions. Appendix D then states that the cited paper's identification uses an additional 'conditional monotonicity over Yx' assumption and that this is equivalent to the present Assumption 8 only under 'Assumption 3.6' of that paper; neither condition is stated in the main text. As written, Theorems 1, 3, 5, and 7 are not established by the cited proof unless Theorem 5.2 holds under exactly Assumptions 1 and 2. The gap is eliminable: under Assumption 2 the pair (Y0,Y1) is comonotonic, so P(Y0<min_p y_p, Y1≥max_p y_p)=max{min_p F0(y_p)-max_p F1(y_p),0}, which yields Eq. (6) directly; the authors should include such a proof or explicitly verify the conditions of the cited theorem.
- [Appendix E, Eq. (115) and surrounding text] The consistency proof claims that the empirical-process term in |σ̂(m)−σ(m)| is O_p(1/√N^m) and that central-moment estimators are O_p(1/√N^{2m}). These rates are dimensionally wrong: each empirical CDF converges at rate O_p(N^{-1/2}) pointwise, and the integrand is 1-Lipschitz in the empirical CDFs, so the integral of the empirical-process discrepancy over the bounded domain is O_p(N^{-1/2}), not O_p(N^{-m/2}). Consistency itself may still hold, but the proof as written does not support the stated rates. The authors should correct the rates and avoid citing the delta method for the nondifferentiable max/min functional.
minor comments (5)
- [Section 3.2, after Eq. (6)] The display for the second moment contains a typo: 'max{P(Y<y1|X=1), P(Y<y1|X=1)}' should read 'max{P(Y<y1|X=1), P(Y<y2|X=1)}'.
- [Appendix A, proof of Lemma 2] The proof states the result for '(i,j) ∈ {(1,0),(0,0)}'; the second pair should be '(0,1)'.
- [Appendix A, proof of Lemma 1] In the second displayed equation of the proof, 'I(Y1 > Y0)^m' should be 'I(Y0 > Y1)'.
- [Section 6, real-world application] The interpretive paragraphs draw strong conclusions about positive skewness, high kurtosis, and outliers from estimates with 95% CIs that include the opposite sign (e.g., skewness 21.027 with CI [−6.747, 34.504] for N=10 per group); the text should acknowledge that the data support little more than a wide range of plausible values.
- [Appendix E] Assumption 9 says P(Y<y|X=x) is 'differential' in y; this should be 'differentiable'. Also, the citation of the delta method for the max/min functional is not appropriate because that map is not differentiable at points where the max argument changes sign.
Circularity Check
No significant circularity: the identification theorems cite an independent prior identification result, and the moment formulas follow by integration; the only caveat is a proof-completeness note in Appendix D, not circularity.
full rationale
The paper's main identification theorems (Theorems 1 and 3) reduce the joint potential-outcome probabilities P(Y0<y1≤Y1,...) to Theorem 5.2 of Kawakami et al. (2024a), a self-citation. The proof states: 'P(Y0 < y1 ≤ Y1, Y0 < y2 ≤ Y1, . . . , Y0 < ym ≤ Y1) and P(Y1 < y1 ≤ Y0, . . . , Y1 < ym ≤ Y0) are identifiable by max{minp=1,...,m{P(Y < yp|X = 0)} − maxp=1,...,m{P(Y < yp|X = 1)}, 0} and max{minp=1,...,m{P(Y < yp|X = 1)}− maxp=1,...,m{P(Y < yp|X = 0)}, 0} respectively by Theorem 5.2 in [Kawakami et al., 2024a] under Assumptions 1 and 2.' This is load-bearing, but the cited theorem is a parameter-free identification result about probabilities of causation for continuous variables, not about moments of causal effects, so the target result is not an input to the citation. The step from those identified joint probabilities to Eq. (6) is a genuine derivation via Lemma 1's integral decomposition and expectation; the product-moment version is the same argument for pair events. The bounds in Theorems 2 and 4 are derived directly from Fréchet inequalities, and the estimators are plug-in empirical versions with a consistency argument; no fitted parameter is renamed as a prediction. Appendix D notes that the cited paper's 'conditional monotonicity over Yx' is equivalent to the present Assumption 8 only under 'Assumption 3.6' of the cited paper, which is not restated; that is a proof-completeness caveat for the conditional results, not a circularity. Overall, the derivation chain is not equivalent to its inputs by construction.
Assumptions & free parameters
assumptions (4)
- domain assumption Assumption 1 (Exogeneity): Yx ⊥⊥ X for all x.
- domain assumption Assumption 2 (Monotonicity over fY): fY(x, UY) is monotone increasing or decreasing in UY for all x.
- domain assumption Assumptions 3-6 (Finiteness and existence of integrals).
- standard math Theorem 5.2 of Kawakami et al. 2024a identifies P(Yj<y1≤Yi, Yh<y2≤Yk) from conditional CDFs under Assumptions 1 and 2.
Cite this review
Pith. "Pith review of Moments of Causal Effects." pith.science (2026). https://pith.science/paper/4BVVXORQ
@misc{pith2026250504971,
author = {Pith},
title = {Pith review of: Moments of Causal Effects},
year = {2026},
howpublished = {\url{https://pith.science/paper/4BVVXORQ}},
note = {Machine review of arXiv:2505.04971}
}
read the original abstract
The moments of random variables are fundamental statistical measures for characterizing the shape of a probability distribution, encompassing metrics such as mean, variance, skewness, and kurtosis. Additionally, the product moments, including covariance and correlation, reveal the relationships between multiple random variables. On the other hand, the primary focus of causal inference is the evaluation of causal effects, which are defined as the difference between two potential outcomes. While traditional causal effect assessment focuses on the average causal effect, this work provides definitions, identification theorems, and bounds for moments and product moments of causal effects to analyze their distribution and relationships. We conduct experiments to illustrate the estimation of the moments of causal effects from finite samples and demonstrate their practical application using a real-world medical dataset.
Reference graph
Works this paper leans on
-
[1]
} . (104) The estimators ˆσ (i, j ; k, h ) and ˆ ¯σ (i, j ; k, h ) are ˆσ (i, j ; k, h ) = (b − a)2 N1N2 N1∑ k1=1 N2∑ k2=1 max { min{ˆP(Y < y 1 k1 |X = j), ˆP(Y < y 2 k2 |X = h)} − max{ˆP(Y < y 1 k1 |X = i), ˆP(Y < y 2 k2 |X = k)}, 0 } − (b − a)2 N1N2 N1∑ k1=1 N2∑ k2=1 max { min{ˆP(Y < y 1 k1 |X = i), ˆP(Y < y 2 k2 |X = h)} − max{ˆP(Y < y 1 k1 |X = j), ˆP...
work page 1935
-
[2]
, y m; i, j ) ≤ P(Yj < y 1 ≤ Yi, Y j < y2 ≤ Yi,
Under SCM M and Assumptions 1 and 3, given m ≥ 1, we have l(y1, . . . , y m; i, j ) ≤ P(Yj < y 1 ≤ Yi, Y j < y2 ≤ Yi, . . . , Y j < y m ≤ Yi) ≤ u(y1, . . . , y m; i, j ) where l(y1, . . . , y m; i, j ) = max { ∑ p=1,...,m P(Y < y p|X = j) − ∑ p=1,...,m P(Y < y p|X = i) − m + 1, 0 } , (33) u(y1, . . . , y m; i, j ) = min { min p=1,...,m {P(Y < y p|X = j)},...
work page 1935
-
[4]
Under SCM M and Assumptions 1 and 4, for any i, j, k, h ∈ {1, . . . , R }, we have l(y1, y 3; i, j, k, h ) ≤ P(Yj < y1 ≤ Yi, Y h < y 2 ≤ Yk) ≤ u(y1, y 3; i, j, k, h ), where l(y1, y 2; i, j, k, h ) = max { P(Y < y 1|X = j) − P(Y < y 1|X = i) + P(Y < y 2|X = h) − P(Y < y 2|X = k) − 1, 0 } , (48) u(y1, y 2; i, j, k, h ) = min { min{P(Y < y 1|X = j), P(Y < y...
work page 1935
-
[5]
, y m; i, j ) ≤ P(Yj − E[Yj ] < y 1 ≤ Yi − E[Yi], Y j − E[Yj ] < y 2 ≤ Yi − E[Yi],
Under SCM M and Assumptions 1 and 5, given m ≥ 1, we have ¯l(y1, . . . , y m; i, j ) ≤ P(Yj − E[Yj ] < y 1 ≤ Yi − E[Yi], Y j − E[Yj ] < y 2 ≤ Yi − E[Yi], . . . , Y j − E[Yj ] < y m ≤ Yi − E[Yi]) ≤ ¯u(y1, . . . , y m; i, j ), where ¯l(y1, . . . , y m; i, j ) = max { ∑ p=1,...,m P(Y − E[Y |X = j] < y p|X = j) − ∑ p=1,...,m P(Y − E[Y |X = i] < y p|X = i) − (...
work page 1935
-
[6]
Under SCM M and Assumptions 1 and 6, for any i, j, k, h ∈ { 1, . . . , R }, we have l(y1, y 3; i, j, k, h ) ≤ P(Yj − E[Yj ] < y 1 ≤ Yi − E[Yi], Y h − E[Yh] < y 2 ≤ Yk − E[Yk]) ≤ u(y1, y 3; i, j, k, h ), where l(y1, y 2; i, j, k, h ) = max { P(Y − E[Y |X = j] < y 1|X = j) − P(Y − E[Y |X = i] < y 1|X = i) + P(Y − E[Y |X = h] < y 2|X = h) − P(Y − E[Y |X = k]...
work page 1935
-
[1928]
APPENDIX TO “MOMENTS OF CAUSAL EFFECTS" We provide the proofs of the lemmas and theorems in the body of the paper in Appendix A, the discussion of the central moments of causal effects in Appendix B, the discussion of th e central product moments of causal effects in Appendix C, the discussion of the conditional moments of causal effects in Appendix D, th...
work page 1935
-
[1959]
Marginal and dependence uncertainty: bounds, optimal transport, and sharpness
Daniel Bartl, Michael Kupper, Thibaut Lux, Antonis Papa- pantoleon, and Stephan Eckstein. Marginal and depen- dence uncertainty: bounds, optimal transport, and sharp- ness. arXiv preprint arXiv:1709.00641 ,
-
[2007]
Jerzy Neyman. Sur les applications de la theorie des prob- abilites aux experiences agricoles: Essai des principes (in polish). english translation by dm dabrowska and tp speed (1990). Statistical Science, 5:465–480,
work page 1990
Show all 12 references
-
[2017]
Fortin, and Thomas Lemieux
John DiNardo, Nicole M. Fortin, and Thomas Lemieux. Labor market institutions and the distribution of wages, 1973-1992: A semiparametric approach. Econometrica, 64(5):1001–1044,
1973
-
[2018]
Kurtosis as peakedness, 1905 -
Peter H Westfall. Kurtosis as peakedness, 1905 -
1905
-
[2021]
Probabil- ities of causation for continuous and vector variables
Y uta Kawakami, Manabu Kuroki, and Jin Tian. Probabil- ities of causation for continuous and vector variables. Proceedings of the 40th Conference on Uncertainty in Artificial Intelligence (UAI-2024) , 2024a. Y uta Kawakami, Manabu Kuroki, and Jin Tian. Identifica- tion and estim...
2024
-
[2024]
URL https://arxiv.org/abs/1806.02935. Sören R. Künzel, Jasjeet S. Sekhon, Peter J. Bickel, and Bin Y u. Metalearners for estimating heterogeneous treatment effects using machine learning. Proceedings of the Na- tional Academy of Sciences , 116(10):4156–4165,
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.