REVIEW 4 minor 12 references
Identification and Bounding of Central Moments of Causal Effects Using Marginal Moments Information
T0 review · 0 major / 4 minor · reviewed 2026-07-11 · grok-4.5
Pith's one-line read Central moments of individual treatment effects can be identified or sharply bounded from only the marginal moments of each potential outcome.
desk verdict Clean, usable theorems that turn published moments into statements about ICE heterogeneity; IED is the only real soft spot and the authors treat it honestly. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The recursive identity (Theorem 1) that expands the m-th moment of Y1 under independence of the individual effect and the baseline, together with the L^p-norm triangle inequalities that produce the sharp even-order bounds when independence is dropped.
What would settle it
If, in a setting where both full joint data and only marginal moments are available, the moment recovered from the recursive formula (or lying inside the claimed sharp bounds) systematically differs from the moment computed directly from the joint sample, the identification or bounding claim is false.
Extended reading notes
Core claim
The m-th central moment of the individual causal effect is identified, under the Independent Effect Deviation condition, by a simple recursive formula that uses only the marginal central moments of the two potential outcomes; without that condition the same moments admit sharp, closed-form bounds expressed solely in terms of those marginal moments.
Load-bearing premise
Point identification requires that a person's treatment effect is statistically independent of their baseline outcome—an untestable modelling assumption that must be justified by subject-matter knowledge.
Editorial extensions
If this is right
- Any RCT that reports only group means and variances immediately yields sharp bounds on the variance of individual treatment effects.
- When higher moments are also published, skewness and kurtosis of the ICE become either point-identified under IED or bounded without it.
- Historical medical and social-science trials that never released micro-data can be re-analyzed for treatment-effect heterogeneity.
- An upper bound on ICE variance translates, via Cantelli’s inequality, into a bound on the probability that an individual experiences an effect of the opposite sign from the average.
- Stratum-specific versions of the same formulae give within-subgroup heterogeneity measures once conditional moments are available.
Reading between the lines
- The same recursive structure should extend, with only notational change, to continuous treatments or multi-valued treatments once the appropriate contrast is defined.
- Because the bounds remain valid under unmeasured confounding once the marginal moments themselves have been bounded, the method can be chained with existing partial-identification techniques for observational data.
- Reporting conventions that currently list only means and standard deviations could usefully add third and fourth moments, instantly unlocking skewness and kurtosis statements about individual effects.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper derives identification and sharp bounding results for the central moments µ(m) of the individual causal effect Y1-Y0 using only the marginal central moments σ(k)1 and σ(k)0 of the two potential outcomes. Under the Independent Effect Deviation (IED) assumption, Theorem 1 and Corollaries 1–2 give recursive point identification of µ(m) (hence of variance, skewness and kurtosis). Without IED, Theorems 3, 6, 10, 12, 14 and 16 supply closed-form sharp bounds that depend on a single pair (σ(k)1,σ(k)0); multi-moment intersections are also stated (Theorems 7, 11, 15, Corollary 3) and shown by counter-examples not to be sharp in general. Two empirical re-analyses of published RCTs that report only summary moments illustrate the methods.
Significance. The contribution is practically useful: many RCTs and observational studies release only means, SDs, skewness or kurtosis for privacy or historical reasons, rendering full-distribution methods inapplicable. The sharp single-k bounds (attained by explicit two-point constructions) and the IED identification formulae therefore enlarge the set of studies in which treatment-effect heterogeneity can be quantified. The proofs rely on standard Lp inequalities and are complete; the sharpness claims are supported by matching constructions and by counter-examples for the non-sharp multi-moment cases. These features make the results immediately usable and falsifiable.
minor comments (4)
- [Table 1 / §5] Table 1 caption and the surrounding text in §5 could more explicitly flag that the multi-k intersections (Theorems 7, 11, 15) are generally not sharp; the counter-examples appear only in the appendix.
- [§7] In the case-study tables the plug-in estimates are reported without any uncertainty quantification; a short remark that simultaneous confidence intervals for the marginal moments can be propagated through the closed-form expressions (as sketched in Appendix D) would be helpful for applied readers.
- [§5.1] Notation for the correlation λ (Eq. 9) is introduced only in §5.1; a forward reference earlier would improve readability.
- A few minor typographical inconsistencies appear (e.g., “Moseley et al. [2002]” vs. “[Moseley et al., 2002]”; occasional missing spaces around mathematical operators).
Circularity Check
No significant circularity: pure algebraic consequences of moment definitions, IED independence, and standard Lp inequalities with explicit attaining constructions.
full rationale
The derivation chain is self-contained mathematics. Theorem 1 follows directly from the binomial expansion of E[(eY0 + eDelta)^m] under the IED independence (Y1-Y0) ⊥ Y0, yielding the recursive formula µ(m) = σ(m)1 - σ(m)0 - sum binom terms; Corollaries 1-2 are immediate special cases. The sharp single-k bounds (Theorems 3, 6, 10, 14, 16) are obtained from Minkowski/triangle inequalities in L^p together with two-point (or sign-flip) constructions that attain the endpoints (Appendix C proofs). Intersections of those bounds are correctly flagged as non-sharp via explicit counter-examples (A.2). No parameters are fitted to data and then re-labeled as predictions; the two case studies simply plug reported summary moments into the closed-form expressions. Self-citations to Kawakami & Tian (2025) and Post & Van Den Heuvel supply background context and comparison baselines but are not load-bearing for the new moment-only identification or bounds. The IED assumption is stated as an untestable domain-knowledge condition, not derived circularly. The paper therefore contains no self-definitional loops, fitted-input predictions, uniqueness smuggling, or renaming of known results.
Assumptions & free parameters
assumptions (3)
- domain assumption Independent Effect Deviation (IED): (Y1 − Y0) ⊥⊥ Y0
- standard math Minkowski and Cauchy–Schwarz inequalities in L^p spaces
- domain assumption Existence of the relevant central moments of the potential outcomes
Cite this review
Pith. "Pith review of Identification and Bounding of Central Moments of Causal Effects Using Marginal Moments Information." pith.science (2026). https://pith.science/paper/KMO42R2L
@misc{pith2026260704957,
author = {Pith},
title = {Pith review of: Identification and Bounding of Central Moments of Causal Effects Using Marginal Moments Information},
year = {2026},
howpublished = {\url{https://pith.science/paper/KMO42R2L}},
note = {Machine review of arXiv:2607.04957}
}
read the original abstract
Evaluating the causal effect of a treatment on an outcome is a central objective in causal inference. While the average causal effect summarizes the mean impact of treatment, the central moments of the individual causal effect (ICE) characterize the shape of the ICE distribution, thereby revealing the extent and structure of treatment effect heterogeneity across individuals. This paper investigates the identification and bounding of the central moments of the ICE using only the marginal central moments of each potential outcome (PO). Compared with existing approaches that require knowledge of the full marginal distributions of the POs, marginal moment information is often substantially easier to obtain in empirical applications. Finally, we illustrate the practical relevance of our results through two empirical case studies.
Figures
Reference graph
Works this paper leans on
-
[1]
+ 2x− 1 2 + √ 3 6 −x (2− √
-
[2]
(42) Therefore2 √ 3−2≤ µ(2) ≤4
=−1 + 6x, µ(2) =E[( eY1 −eY0)2] =E[ eY 2 1 ]−2E[ eY1eY0] +E[eY 2 0 ] = 2−2E[ eY1eY0] = 2−2(−1 + 6x) = 4−12x. (42) Therefore2 √ 3−2≤ µ(2) ≤4. On the other hand, (19) gives a bound0≤ µ(2) ≤4 √
-
[3]
Counterexample for sharpness of the bound(20)
Hence (19) is not sharp. Counterexample for sharpness of the bound(20). Lower bound. Let σ(2) 1 = σ(2) 0 = 1 , σ(3) 1 = √ 2, σ(3) 0 =− √ 2 and σ(4) 1 = σ(4) 0 = 3 . Let (eY1,eY0) be a distribution compatible with it. By Lemma 1, such an (eY1,eY0) exists, andeY1 andeY0 have the distribution in (38) and (39), respectively. Eq. (20) gives a bound0≤ µ(2) ≤4. ...
-
[4]
The case( σ(3) 1 , σ(3) 0 )and( σ(4) 1 , σ(4) 0 )are available
In particular, neither bound in (24) is sharp. The case( σ(3) 1 , σ(3) 0 )and( σ(4) 1 , σ(4) 0 )are available. Let σ(3) 1 = 0, σ(4) 1 = 1, σ(3) 0 = 0, and σ(4) 0 = 0. (These are realized by, e.g.,eY0 = 0and eY1 withP(eY1 = 1) =P( eY1 =−1) = 1/2.) Then σ(4) 0 = 0 implieseY0 = 0 almost surely and hencee∆ = eY1 almost surely. Thereforeµ(3) =E[ e∆3] =E[ eY 3 ...
-
[5]
The case( σ(2) 1 , σ(2) 0 ),( σ(3) 1 , σ(3) 0 )and( σ(4) 1 , σ(4) 0 )are available
In particular, neither bound in (24) is sharp. The case( σ(2) 1 , σ(2) 0 ),( σ(3) 1 , σ(3) 0 )and( σ(4) 1 , σ(4) 0 )are available. Let σ(2) 1 = σ(4) 1 = 1 , σ(3) 1 = 0 , and σ(2) 0 = σ(3) 0 = σ(4) 0 = 0. (These are realized by, e.g.,eY0 = 0and eY1 withP(eY1 = 1) =P( eY1 =−1) = 1/2.) Then σ(4) 0 = 0 implieseY0 = 0 almost surely and hencee∆ = eY1 almost sur...
-
[6]
Counterexample for sharpness of the bound(31)
In particular, neither bound in (24) is sharp. Counterexample for sharpness of the bound(31). Lower bound. Let σ(3) 1 = √ 2, σ(3) 0 =− √ 2 and σ(4) 1 = σ(4) 0 = 3, and let (eY1,eY0) be a distribution compatible with it. By Lemma 1,eY1 andeY0 have the distribution in (38) and (39), respectively. Eq. (31) gives0≤ µ(4) ≤48. Letx :=P eY1 = √ 2+ √ 6 2 , eY0 =−...
-
[7]
Therefore, we have eY 2 1 =c almost surely for some c∈R , and we obtain c= 1 since σ(2) 1 = 1
= 1 200 .(47) Since σ(4) 1 = (σ(2) 1 )2, the equality in the Cauchy-Schwarz inequality (σ(2) 1 )2 =E[ eY 2 1 ]2 ≤E[( eY 2 1 )2]E[12] = σ(4) 1 holds. Therefore, we have eY 2 1 =c almost surely for some c∈R , and we obtain c= 1 since σ(2) 1 = 1 . Thus eY 3 1 = eY1 and E[eY 2 1 eY 2 0 ] =E[ eY 2 0 ] = 1 10. Now we have µ(4) =E[( eY1 −eY0)4] = 13 5 −4E[eY1eY0...
-
[8]
Moreover, since √ 10<4 we have (1− 1√ 10)4 <(3/4) 4 < 3
Show all 12 references
-
[9]
Also, (32) coincides with (30), so (32) is not sharp either
Thus neither bound in (30) is sharp. Also, (32) coincides with (30), so (32) is not sharp either. B BOUNDING µ(m) USING ONLY MARGINAL SKEWNESS AND KURTOSIS If we have access to variance, skewness and kurtosis, then we can compute 3rd and 4th central moments. In some cases, res...
-
[10]
If σ(3) 0 ̸= 0, let U be a random variable E[U] = 0, E[U2] =η −2/3, and E[U3] = σ(3) 0 /η (Lemma 5)
and let S be a random variable with P(S= 1) =P(S=−1) = 1/2. If σ(3) 0 ̸= 0, let U be a random variable E[U] = 0, E[U2] =η −2/3, and E[U3] = σ(3) 0 /η (Lemma 5). Otherwise set U := 0. LetWbe a random variable with P(W=η −7/16) =P(W=−η −7/16) =P W= η−7/16 2 =P W=− η−7/16 2 = 1 4...
-
[11]
In addition, we have µ(4) =E[ e∆4] = (1−2η)(c 1 −c 0)4 +ηE[H 4].(112) If σ(3) 0 ̸= 0, thenηE[U 2] =η 1/3, while if σ(3) 0 = 0, thenηE[U 2] = 0
Moreover, E[eY 3 0 ] =ηE[U 3] = σ(3) 0 and E[eY 3 1 ] = ηE[U 3] +ηE[(W+H) 3] = σ(3) 1 . In addition, we have µ(4) =E[ e∆4] = (1−2η)(c 1 −c 0)4 +ηE[H 4].(112) If σ(3) 0 ̸= 0, thenηE[U 2] =η 1/3, while if σ(3) 0 = 0, thenηE[U 2] = 0. Moreover, we have ηE[W 2] = 5 8 η1/8, ηE[H 2]...
-
[12]
Replacing ZA by −ZA yields µ(m) <−M
Taking A large yields µ(m) > M . Replacing ZA by −ZA yields µ(m) <−M . By Lemma 2, these can be realized by SCMs. D NUMERICAL EXPERIMENTS In this appendix, we conduct numerical experiments to illustrate the properties of our results in a finite-sample setting. Table 5: Estimat...
2025
Reviewed July 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.