Pith. sign in

REVIEW 5 minor 100 references

Latent treatment effects under unobserved confounding are the eigenvalues of one compressed proxy operator.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

After shared-subspace compression, the difference of treatment-arm proxy quotient operators is similar to the diagonal of latent treatment effects, whose eigenvalues and lifted eigenvectors recover the full mixture.

T0 review reviewed 2026-07-14 challenge →

load-bearing objection Clean spectral re-foundation of SPO: the latent effects are eigenvalues of one compressed difference operator, with full proofs and better finite-sample behavior than the scalar recursion.

arxiv 2607.10926 v1 pith:VVXTSAWP submitted 2026-07-12 cs.LG stat.ML

The Spectral Structure of Latent Treatment Effects

classification cs.LG stat.ML
keywords heterogeneous treatment effectsproximal causal inferencelatent confoundingobservable operatorsspectral methodssynthetic potential outcomesmixture of treatment effects
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

When treatment effects vary across unobserved groups and only proxies are observed, earlier work recovered the mixture of those effects by building a long chain of scalar moments and extracting roots from a Hankel pencil. This paper shows that those moments are simply bilinear projections of a single, finite-dimensional operator that can be formed directly from observable second- and third-order moments. After the shared signal subspace of the proxies is extracted, the difference of the two arm-specific quotient operators is similar to the diagonal matrix of latent treatment effects; its eigenvalues are exactly those effects, and its left eigenvectors recover the feature matrix and the mixture weights. The construction works with more proxy coordinates than latent classes, replaces recursive high-order inversion by a single spectral decomposition, and comes with high-probability first-order error bounds. A sympathetic reader cares because the same causal target is recovered more stably and under weaker dimensional restrictions than the scalar recursion.

Core claim

Under the same population factorizations used by Synthetic Potential Outcomes, there exists an exact compressed observable operator: after projection onto the shared k-dimensional proxy signal subspace, the difference of the two treatment-arm quotient operators is similar to the diagonal of latent treatment effects. Its eigenvalues are precisely the latent effects; its lifted left eigenvectors, after anchor normalization, recover the target-proxy feature matrix and thence the latent mixture proportions. Every scalar SPO moment is a bilinear functional of a power of this operator.

What carries the argument

The compressed difference operator ΔQ̃ = Q̃₁ − Q̃₀. After the shared row space of the stacked cross-moment matrix is extracted by truncated SVD, each arm yields a k × k quotient Q̃_t; their difference is similar to diag(τ(u)), so spectral analysis on this single matrix recovers the entire mixture of treatment effects.

Load-bearing premise

The unobserved confounder must be a finite discrete set of known size in which every class appears under both treatments and the class-specific treatment effects are all different from one another.

What would settle it

Generate synthetic data from the exact proxy mixture model with known distinct π values, run the compressed spectral estimator, and check whether the recovered eigenvalues match the planted effects to within the predicted n^{-1/2} rate; systematic failure under full rank and positivity would refute the claim.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Overcomplete proxy systems (more coordinates than latent classes) can be used directly without forced square truncation.
  • High-order scalar inversion and Hankel root-finding are replaced by a single finite-dimensional eigendecomposition.
  • First-order high-probability bounds become available simultaneously for the effects, the feature rows, and the simplex-projected mixture weights.
  • Latent causal homogeneity is equivalent to the compressed difference operator being a scalar matrix, giving an exact operator test for effect constancy.
  • The same operator generates the entire synthetic-moment hierarchy, so every polynomial functional of the treatment-effect law is recovered from one object.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same subspace-compression idea may extend to continuous latent confounders via compact integral operators or RKHS analogues, as the discussion already hints.
  • Because the operator is environment-invariant when the potential-outcome diagonals are shared, multi-environment data could be pooled for a single spectral estimate without re-deriving moments.
  • The complex-bifurcation diagnostic already present in the paper could serve as a practical rank-selection or model-check statistic: large imaginary parts flag that the sample is outside the separated-eigenvalue regime.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 5 minor

Summary. The paper studies identification of a mixture of treatment effects (MTE) under a discrete latent confounder with proxy variables. Under the same conditional-independence and full-rank factorization assumptions used by Mazaheri et al. (2025) for Synthetic Potential Outcomes, it constructs ambient and compressed quotient operators from observable second- and third-order moment matrices. After projection onto the shared k-dimensional proxy signal subspace, the difference of the two arm-specific compressed operators is similar to the diagonal of latent treatment effects; its eigenvalues are exactly the latent effects, and its lifted left eigenvectors (after anchor normalization by X1=1) recover the target-proxy feature matrix B and the mixture weights p. The scalar SPO moment sequence is recovered as bilinear functionals of powers of this operator. The construction accepts overcomplete proxies, and the paper supplies geometric rank/positivity characterizations, high-probability n^{-1/2} perturbation bounds for eigenvalues, lifted rows, and simplex-projected weights, and synthetic experiments showing improved stability relative to the recursive scalar baseline.

Significance. If the results hold, the paper supplies a clean operator-theoretic foundation for latent MTE identification in the discrete-proxy setting. The central similarity (Theorem 5.2) is an exact algebraic consequence of the stated factorizations and shared-subspace compression; the operator-moment equivalence (Theorem 5.4) unifies the prior recursive construction; and the finite-sample bounds (Theorems 7.1–7.2) are first-order and use standard matrix-perturbation tools. The ability to handle overcomplete proxies without square inversion, together with the geometric diagnostics for rank and positivity, is a genuine practical and conceptual advance over the scalar Hankel-pencil route. Strengths include fully written appendix proofs, an explicit algorithm, and synthetic experiments that match the predicted rate and demonstrate clear gains over the baseline.

minor comments (5)
  1. The complex-bifurcation diagnostic (Remark 3) is useful but could be stated more operationally: e.g., a concrete threshold on the imaginary part relative to the estimated eigengap, or a short note on how often residual imaginary parts appear in the reported experiments.
  2. Figure 4 and Table 1 report median absolute eigenvalue error; adding interquartile ranges or a brief note on trial-to-trial variability would make the stability claim easier to assess at a glance.
  3. Assumption 3 (spectral separation) is used for simple-eigenvalue pairing and eigenvector recovery; a short remark on the multiple-eigenvalue case (already alluded to in Remark 1 for zero effects) would clarify what is still identified when some τ(u) coincide.
  4. Notation for the compressed operators (˜Qt vs. Δ˜Q) is consistent but dense; a one-line glossary or a small table of the main population objects would help readers navigating Sections 5–7.
  5. The related-work discussion of spectral OOMs and tensor methods is appropriate; a sentence distinguishing the present asymmetric causal quotient from simultaneous diagonalization of symmetric tensors would further locate the contribution.

Circularity Check

0 steps flagged

No significant circularity: spectral identification is an algebraic consequence of stated proxy factorizations, not a fit or self-citation loop.

full rationale

The paper’s load-bearing chain is: Assumption 1 (proxy conditional independencies) plus law of total expectation yield the factorizations M_ZX|t = A_t D_U|t B^T and M_ZXY|t = A_t D_U|t D_Y|t B^T; full column rank (Assumption 2) and positivity give ambient quotients Q_t = (B^T)^† D_Y|t B^T (Theorem 5.1); shared row-space compression then produces ˜Q_t = R^{-1} D_Y|t R and Δ˜Q = R^{-1} D_τ R (Theorem 5.2), so eigenvalues are exactly the latent effects and left eigenvectors recover B after the X_1=1 anchor (Proposition 5.3). Every step is proved from these hypotheses via pseudo-inverse identities and subspace geometry (Appendix B, Lemmas B.1–B.2); no parameter is fitted to force the spectrum to match a target, and the operator-moment equivalence (Theorem 5.4) shows prior SPO scalar moments are bilinear functionals of powers of the same operator rather than an independent premise. Citation of Mazaheri–Squires–Uhler ’25 is historical/comparative (same author set) and is not used as a uniqueness theorem that forbids alternatives or that supplies the similarity identity. Finite-sample bounds are standard perturbation consequences of the population operator, not circular predictions. The derivation is therefore self-contained against its stated assumptions; score 0.

Axiom & Free-Parameter Ledger

1 free parameters · 4 axioms · 1 invented entities

Identification rests on standard causal proxy independencies, a finite discrete latent class model of known size k, full-column-rank proxy feature matrices, strict positivity, and spectral separation of the treatment effects. No free parameters are fitted to force the spectral claim; k may be chosen by a separate rank procedure but is treated as known for the main theorems. The compressed operator itself is a derived object, not an invented physical entity.

free parameters (1)
  • latent dimension k
    Treated as known for population identification and finite-sample bounds; when unknown it is selected by a separate rank procedure on the stacked moment matrix whose error is not absorbed into the main rates.
axioms (4)
  • domain assumption Proxy conditional independencies Z ⊥ (X,Y)|(T,U) and X ⊥ (Y,T)|U together with latent ignorability Y(t) ⊥ T|U (Assumption 1)
    These are the standard proximal-causal factorization assumptions that produce the low-rank moment factorizations M_ZX|t = A_t D_U|t B^T and M_ZXY|t = A_t D_U|t D_Y|t B^T.
  • domain assumption Proxy richness: A_0, A_1, B have full column rank k and d_x, d_z ≥ k (Assumption 2)
    Guarantees that the stacked moment matrix has exact rank k and that the compressed operators are well-defined and similar to the latent diagonals.
  • domain assumption Strict latent positivity: P(T=t|U=u)>0 for all u,t, and spectral separation of the τ(u) (Assumption 3 and positivity hypotheses)
    Positivity equates the treated and control row spaces with the row space of B; separation makes eigenvalues simple so left eigenvectors recover rows of R up to scale.
  • standard math Bounded almost-sure norms of the moment matrices and positive treatment probability π (Assumption 4)
    Standard concentration hypotheses that convert matrix Bernstein and Wedin bounds into the stated n^{-1/2} rates.
invented entities (1)
  • compressed difference operator ΔQ̃ independent evidence
    purpose: Finite-dimensional observable operator whose spectrum equals the latent treatment effects and whose powers generate all SPO moments
    Constructed directly from observable second- and third-order moments after SVD projection; not a new physical object but the central algebraic device of the paper.

reviewed 2026-07-14 · how reviews work

0 comments
Cite this review

Pith. "Pith review of The Spectral Structure of Latent Treatment Effects." pith.science (2026). https://pith.science/paper/VVXTSAWP

@misc{pith2026260710926,
  author       = {Pith},
  title        = {Pith review of: The Spectral Structure of Latent Treatment Effects},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VVXTSAWP}},
  note         = {Machine review of arXiv:2607.10926}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Identifying heterogeneous treatment effects under unobserved confounding is central in observational causal inference. In proxy models with a discrete latent confounder, prior Synthetic Potential Outcomes (SPO) [Mazaheri-Squires-Uhler '25] recover the mixture of treatment effects through recursively constructed scalar moments. We show that this sequence is one projection of a more fundamental object. Under the same population factorization assumptions, there is an exact compressed observable operator: after projecting onto the shared proxy signal subspace, the difference of two treatment-arm quotient operators is similar to the diagonal matrix of latent treatment effects. Its eigenvalues are the latent effects; its lifted left eigenvectors, after anchor normalization, recover the target-proxy feature matrix and then the latent mixture proportions. Every scalar SPO moment is a bilinear functional of a power of this operator. The resulting estimator handles overcomplete proxy systems, replaces high-order scalar inversion with finite-dimensional spectral analysis, and admits high-probability first-order perturbation bounds for treatment effects, feature rows, and simplex-projected mixture weights.

Figures

Figures reproduced from arXiv: 2607.10926 by Bijan Mazaheri, Hamza Virk, Yihren Wu.

Figure 1
Figure 1. Figure 1: Causal triptych for the latent proxy mixture model. (a) depicts the proxy causal structure: [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Distributional view of latent treatment effects. (a) The observed arm-specific outcome [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Visual representation of the proxy matrix factorization [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Left: Median absolute eigenvalue error of the compressed spectral estimator across latent dimensions k ∈ {2, . . . , 6} and sample sizes N ∈ {103 , . . . , 2.5×104}. Center: Corresponding errors for the base SPO moment-chain baseline. Errors are larger and less stable, especially as k increases. Right: Empirical convergence rate of the spectral estimator at k = 3 against the O(n −1/2 ) reference. working s… view at source ↗
Figure 5
Figure 5. Figure 5: Naive ATE estimates across 300 trials for the three-class mixture with true effects [PITH_FULL_IMAGE:figures/full_fig_p017_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Base SPO eigenvalue estimates across 300 trials. Although the scalar moment construction [PITH_FULL_IMAGE:figures/full_fig_p018_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Compressed spectral estimator eigenvalue estimates across 300 trials. The empirical [PITH_FULL_IMAGE:figures/full_fig_p018_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Row-space compression underlying Lemma B.2. Left: The rows of M lie in the signal subspace R(M⊤) = R(Vk), so projecting onto Vk preserves the population row space. Right: Compression is exact because M = MVkV⊤ k ; equivalently, M and MVk have the same nonzero singular values. Proof. Because the columns of V form an orthonormal basis for R(M⊤), the matrix P ≜ VV⊤ is the orthogonal projector onto R(M⊤). Henc… view at source ↗
Figure 9
Figure 9. Figure 9: Counterfactual pairing through the common similarity basis. [PITH_FULL_IMAGE:figures/full_fig_p037_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Three-way heatmap comparison on the (k, N) grid of Section 8.1, with all panels on a shared logarithmic color scale. Left: Spectral SPO. Center: base SPO truncated to the first k proxy coordinates. Right: base SPO applied to the full d = k + 3 moment matrices without truncation. The no-truncation variant is uniformly worse than even the truncated baseline. 49 [PITH_FULL_IMAGE:figures/full_fig_p049_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Convergence at k = 3 and dz = dx = k + 3. The spectral estimator follows the predicted O(n −1/2 ) rate; base SPO truncated to the first k coordinates is noisy and roughly flat at much higher error; base SPO without truncation is higher still and non-monotone. 4 3 2 1 0 1 2 3 4 Estimated Treatment Effect Base SPO (no truncation, full d = k + 3) [PITH_FULL_IMAGE:figures/full_fig_p050_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Base SPO without truncation, eigenvalue estimates across 300 trials with [PITH_FULL_IMAGE:figures/full_fig_p050_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Base SPO eigenvalue estimates across 300 trials. [PITH_FULL_IMAGE:figures/full_fig_p051_13.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

100 extracted references · 3 canonical work pages

  1. [1]

    IEEE Transactions on Electronic Computers , number=

    A theory of adaptive pattern classifiers , author=. IEEE Transactions on Electronic Computers , number=. 1967 , publisher=

  2. [2]

    2026 , eprint=

    Comparing Two Proxy Methods for Causal Identification , author=. 2026 , eprint=

  3. [3]

    2025 , eprint=

    Spectral Representation for Causal Estimation with Hidden Confounders , author=. 2025 , eprint=

  4. [4]

    Biometrika , volume =

    Identifying Causal Effects with Proxy Variables of an Unmeasured Confounder , author =. Biometrika , volume =. 2018 , publisher =

  5. [5]

    BMC Medical Research Methodology , volume =

    Evaluating Sensitivity to Classification Uncertainty in Latent Subgroup Effect Analyses , author =. BMC Medical Research Methodology , volume =. 2022 , publisher =

  6. [6]

    2026 , eprint =

    Estimating Aleatoric Uncertainty in the Causal Treatment Effect , author =. 2026 , eprint =. doi:10.48550/arXiv.2602.08461 , url =

  7. [7]

    2023 , eprint =

    Causal Discovery under Latent Class Confounding , author =. 2023 , eprint =. doi:10.48550/arXiv.2311.07454 , url =

  8. [8]

    Proceedings of The 28th International Conference on Artificial Intelligence and Statistics , series =

    Density Ratio-based Proxy Causal Learning Without Density Ratios , author =. Proceedings of The 28th International Conference on Artificial Intelligence and Statistics , series =. 2025 , publisher =

  9. [9]

    Journal of Educational and Behavioral Statistics , volume =

    Estimating Heterogeneous Treatment Effects Within Latent Class Multilevel Models: A Bayesian Approach , author =. Journal of Educational and Behavioral Statistics , volume =. 2023 , publisher =. doi:10.3102/10769986221115446 , url =

  10. [10]

    Current Epidemiology Reports , volume =

    A Selective Review of Negative Control Methods in Epidemiology , author =. Current Epidemiology Reports , volume =. 2020 , publisher =

  11. [11]

    Journal of the American Statistical Association , volume =

    Kendrick Qijun Li, Xu Shi, Wang Miao and Eric Tchetgen Tchetgen , title =. Journal of the American Statistical Association , volume =. 2024 , publisher =. doi:10.1080/01621459.2023.2220935 , URL =

  12. [12]

    The Annals of Statistics , volume =

    Identifiability of Parameters in Latent Structure Models with Many Observed Variables , author =. The Annals of Statistics , volume =. 2009 , publisher =

  13. [13]

    Journal of the Royal Statistical Society: Series B (Methodological) , volume=

    The interpretation of interaction in contingency tables , author=. Journal of the Royal Statistical Society: Series B (Methodological) , volume=. 1951 , publisher=

  14. [14]

    Proceedings of the 51st Annual ACM SIGACT Symposium on Theory of Computing , pages =

    Beyond the Low-Degree Algorithm: Mixtures of Subcubes and Their Applications , author =. Proceedings of the 51st Annual ACM SIGACT Symposium on Theory of Computing , pages =. 2019 , publisher =

  15. [15]

    SIAM Journal on Computing , volume=

    Learning mixtures of product distributions over discrete domains , author=. SIAM Journal on Computing , volume=. 2008 , publisher=

  16. [16]

    FEBS letters , volume=

    Chromatin accessibility and guide sequence secondary structure affect CRISPR-Cas9 gene editing efficiency , author=. FEBS letters , volume=. 2017 , publisher=

  17. [17]

    Journal of Machine Learning Research , volume =

    Tensor Decompositions for Learning Latent Variable Models , author =. Journal of Machine Learning Research , volume =. 2014 , url =

  18. [18]

    Proceedings of the national academy of sciences , volume=

    Metalearners for estimating heterogeneous treatment effects using machine learning , author=. Proceedings of the national academy of sciences , volume=. 2019 , publisher=

  19. [19]

    Journal of the American Statistical Association , volume =

    Estimation and Inference of Heterogeneous Treatment Effects Using Random Forests , author =. Journal of the American Statistical Association , volume =. 2018 , publisher =

  20. [20]

    Advances in Multilevel Modeling for Educational Research: Addressing Practical Issues Found in Real-World Applications , pages =

    Mixture Modeling Methods for Causal Inference with Multilevel Data , author =. Advances in Multilevel Modeling for Educational Research: Addressing Practical Issues Found in Real-World Applications , pages =. 2015 , publisher =

  21. [21]

    Quantitative Psychology Research: The 79th Annual Meeting of the Psychometric Society, Madison, Wisconsin, 2014 , pages=

    Multilevel propensity score methods for estimating causal effects: A latent class modeling strategy , author=. Quantitative Psychology Research: The 79th Annual Meeting of the Psychometric Society, Madison, Wisconsin, 2014 , pages=. 2015 , organization=

  22. [22]

    Journal of Educational and Behavioral Statistics , volume=

    Hybridizing machine learning methods and finite mixture models for estimating heterogeneous treatment effects in latent classes , author=. Journal of Educational and Behavioral Statistics , volume=. 2021 , publisher=

  23. [23]

    Probabilistic and Causal Inference: The Works of Judea Pearl , pages =

    Detecting Latent Heterogeneity , author =. Probabilistic and Causal Inference: The Works of Judea Pearl , pages =. 2022 , publisher =

  24. [24]

    Biometrika , volume =

    Quasi-Oracle Estimation of Heterogeneous Treatment Effects , author =. Biometrika , volume =. 2021 , publisher =. doi:10.1093/biomet/asaa076 , url =

  25. [25]

    2015 , publisher=

    Causal inference in statistics, social, and biomedical sciences , author=. 2015 , publisher=

  26. [26]

    Foundations and Trends

    An introduction to matrix concentration inequalities , author=. Foundations and Trends. 2015 , publisher=

  27. [27]

    Journal of the American Statistical Association , volume =

    The Blessings of Multiple Causes , author =. Journal of the American Statistical Association , volume =. 2019 , publisher =

  28. [28]

    Journal of the American Statistical Association , volume =

    Comment on ``Blessings of Multiple Causes'' , author =. Journal of the American Statistical Association , volume =. 2019 , publisher =

  29. [29]

    IEEE Transactions on Information Theory , volume=

    Hadamard extensions and the identification of mixtures of product distributions , author=. IEEE Transactions on Information Theory , volume=. 2022 , publisher=

  30. [30]

    Proceedings of the forty-seventh annual ACM symposium on Theory of computing , pages=

    Learning mixtures of gaussians in high dimensions , author=. Proceedings of the forty-seventh annual ACM symposium on Theory of computing , pages=

  31. [31]

    Proceedings of the 4th conference on Innovations in Theoretical Computer Science , pages=

    Learning mixtures of spherical gaussians: moment methods and spectral decompositions , author=. Proceedings of the 4th conference on Innovations in Theoretical Computer Science , pages=

  32. [32]

    Proceedings of Thirty Seventh Conference on Learning Theory , pages =

    Identification of Mixtures of Discrete Product Distributions in Near-Optimal Sample and Time Complexity , author =. Proceedings of Thirty Seventh Conference on Learning Theory , pages =. 2024 , volume =

  33. [33]

    Proceedings of the 5th Conference on Innovations in Theoretical Computer Science , pages =

    Learning Mixtures of Arbitrary Distributions over Large Discrete Domains , author =. Proceedings of the 5th Conference on Innovations in Theoretical Computer Science , pages =. 2014 , publisher =

  34. [34]

    2010 IEEE 51st Annual Symposium on Foundations of Computer Science , pages=

    Settling the polynomial learnability of mixtures of gaussians , author=. 2010 IEEE 51st Annual Symposium on Foundations of Computer Science , pages=. 2010 , organization=

  35. [35]

    arXiv preprint arXiv:2001.06555 , year =

    Counterexamples to ``The Blessings of Multiple Causes'' by Wang and Blei , author =. arXiv preprint arXiv:2001.06555 , year =

  36. [36]

    Science , volume=

    Protein structure relationships revealed by mutational analysis , author=. Science , volume=. 1964 , publisher=

  37. [37]

    2009 , publisher=

    Probabilistic graphical models: principles and techniques , author=. 2009 , publisher=

  38. [38]

    1996 , publisher=

    Graphical models , author=. 1996 , publisher=

  39. [39]

    IEEE Transactions on Acoustics, Speech, and Signal Processing , volume=

    Matrix pencil method for estimating parameters of exponentially damped/undamped sinusoids in noise , author=. IEEE Transactions on Acoustics, Speech, and Signal Processing , volume=. 1990 , publisher=

  40. [40]

    Research in Computational Molecular Biology: 23rd Annual International Conference, RECOMB 2019, Washington, DC, USA, May 5-8, 2019, Proceedings 23 , pages=

    How many subpopulations is too many? Exponential lower bounds for inferring population histories , author=. Research in Computational Molecular Biology: 23rd Annual International Conference, RECOMB 2019, Washington, DC, USA, May 5-8, 2019, Proceedings 23 , pages=. 2019 , organization=

  41. [41]

    Proceedings of the 38th International Conference on Machine Learning , pages =

    MSA Transformer , author =. Proceedings of the 38th International Conference on Machine Learning , pages =. 2021 , editor =

  42. [42]

    2013 , publisher=

    Matrix computations , author=. 2013 , publisher=

  43. [43]

    Biometrika , volume=

    The central role of the propensity score in observational studies for causal effects , author=. Biometrika , volume=. 1983 , publisher=

  44. [44]

    Journal Polytechnique ou Bulletin du Travail fait a l'Ecole Centrale des Travaux Publics , year=

    Essai experimental et analytique: sur les lois de la dilatabilite des fluides elastique et sur celles de la force expansive de la vapeur de l'eau et de la vapeur de l'alkool, a differentes temperatures , author=. Journal Polytechnique ou Bulletin du Travail fait a l'Ecole Centrale des Travaux Publics , year=

  45. [45]

    arXiv preprint arXiv:2212.11279 , year=

    Annotated History of Modern AI and Deep Learning , author=. arXiv preprint arXiv:2212.11279 , year=

  46. [46]

    Nature , volume=

    A global reference for human genetic variation , author=. Nature , volume=. 2015 , publisher=

  47. [47]

    Contemporary Oncology/Wsp

    Review The Cancer Genome Atlas (TCGA): an immeasurable source of knowledge , author=. Contemporary Oncology/Wsp. 2015 , publisher=

  48. [48]

    bioRxiv , pages =

    Short Tandem Repeats Information in TCGA is Statistically Biased by Amplification , author =. bioRxiv , pages =. 2019 , publisher =

  49. [49]

    Plos one , volume=

    Glioblastoma signature in the DNA of blood-derived cells , author=. Plos one , volume=. 2021 , publisher=

  50. [50]

    2017 , publisher=

    Elements of causal inference: foundations and learning algorithms , author=. 2017 , publisher=

  51. [51]

    2009 , publisher=

    Causality , author=. 2009 , publisher=

  52. [52]

    The scientific world journal , volume=

    A review of data fusion techniques , author=. The scientific world journal , volume=. 2013 , publisher=

  53. [53]

    Proceedings of the National Academy of Sciences , volume=

    Causal inference and the data-fusion problem , author=. Proceedings of the National Academy of Sciences , volume=. 2016 , publisher=

  54. [54]

    2023 , eprint=

    Causal Information Splitting: Engineering Proxy Features for Robustness to Distribution Shifts , author=. 2023 , eprint=

  55. [55]

    Proceedings of Thirty Fourth Conference on Learning Theory , pages =

    Source Identification for Mixtures of Product Distributions , author =. Proceedings of Thirty Fourth Conference on Learning Theory , pages =. 2021 , volume =

  56. [56]

    Proceedings of the Second Conference on Causal Learning and Reasoning , pages =

    Causal Inference Despite Limited Global Confounding via Mixture Models , author =. Proceedings of the Second Conference on Causal Learning and Reasoning , pages =. 2023 , volume =

  57. [57]

    Mathematical Social Sciences , volume=

    Induced binary probabilities and the linear ordering polytope: A status report , author=. Mathematical Social Sciences , volume=. 1992 , publisher=

  58. [58]

    The American Mathematical Monthly , volume=

    The paradox of nontransitive dice , author=. The American Mathematical Monthly , volume=. 1994 , publisher=

  59. [59]

    arXiv preprint arXiv:2007.08101 , year=

    The sparse Hausdorff moment problem, with application to topic models , author=. arXiv preprint arXiv:2007.08101 , year=

  60. [60]

    arXiv preprint arXiv:2107.07054 , year=

    Expert Graphs: Synthesizing New Expertise via Collaboration , author=. arXiv preprint arXiv:2107.07054 , year=

  61. [61]

    Combining Binary Classifiers Leads to Nontransitive Paradoxes , author=

  62. [62]

    arXiv preprint arXiv:2006.07691 , year=

    Synthetic interventions , author=. arXiv preprint arXiv:2006.07691 , year=

  63. [63]

    Proceedings of the First Conference on Causal Learning and Reasoning , pages =

    Causal Imputation via Synthetic Interventions , author =. Proceedings of the First Conference on Causal Learning and Reasoning , pages =. 2022 , volume =

  64. [64]

    Proceedings of Thirty Sixth Conference on Learning Theory , pages =

    Causal Matrix Completion , author =. Proceedings of Thirty Sixth Conference on Learning Theory , pages =. 2023 , volume =

  65. [65]

    Journal of the American statistical Association , volume=

    Synthetic control methods for comparative case studies: Estimating the effect of California’s tobacco control program , author=. Journal of the American statistical Association , volume=. 2010 , publisher=

  66. [66]

    Journal of Economic Literature , year =

    Abadie, Alberto , title =. Journal of Economic Literature , year =

  67. [67]

    Epidemiology , volume=

    Negative controls: a tool for detecting confounding and bias in observational studies , author=. Epidemiology , volume=. 2010 , publisher=

  68. [68]

    Statistical Theory and Related Fields , volume =

    A Confounding Bridge Approach for Double Negative Control Inference on Causal Effects , author =. Statistical Theory and Related Fields , volume =. 2024 , publisher =

  69. [69]

    arXiv preprint arXiv:2009.10982 , year =

    An Introduction to Proximal Causal Learning , author =. arXiv preprint arXiv:2009.10982 , year =. doi:10.48550/arXiv.2009.10982 , url =

  70. [70]

    Proceedings of the 38th International Conference on Machine Learning , pages =

    Proximal Causal Learning with Kernels: Two-Stage Estimation and Moment Restriction , author =. Proceedings of the 38th International Conference on Machine Learning , pages =. 2021 , volume =

  71. [71]

    Advances in Neural Information Processing Systems , volume =

    Kernel Instrumental Variable Regression , author =. Advances in Neural Information Processing Systems , volume =. 2019 , url =

  72. [72]

    Political Analysis , volume=

    Estimation of heterogeneous treatment effects from randomized experiments, with application to the optimal planning of the get-out-the-vote campaign , author=. Political Analysis , volume=. 2011 , publisher=

  73. [73]

    Observational Studies , volume =

    CausalToolbox---Estimator Stability for Heterogeneous Treatment Effects , author =. Observational Studies , volume =. 2019 , publisher =

  74. [74]

    Statistics in Medicine , volume =

    Comparing Methods for Estimation of Heterogeneous Treatment Effects Using Observational Data from Health Care Databases , author =. Statistics in Medicine , volume =. 2018 , publisher =

  75. [75]

    Sociological Methodology , volume =

    Estimating Heterogeneous Treatment Effects with Observational Data , author =. Sociological Methodology , volume =. 2012 , publisher =

  76. [76]

    Statistics in medicine , volume=

    Estimating heterogeneous treatment effects for latent subgroups in observational studies , author=. Statistics in medicine , volume=. 2019 , publisher=

  77. [77]

    Bayesian Analysis , year =

    Shahn, Zach and Madigan, David , title =. Bayesian Analysis , year =. doi:10.1214/16-BA1022 , url =

  78. [78]

    Prevention Science , volume=

    For whom does it work? Subgroup differences in the effects of a school-based universal prevention program , author=. Prevention Science , volume=. 2013 , publisher=

  79. [79]

    BMC medical research methodology , volume=

    From concepts, theory, and evidence of heterogeneity of treatment effects to methodological approaches: a primer , author=. BMC medical research methodology , volume=. 2012 , publisher=

  80. [80]

    Journal of Machine Learning Research , year =

    Jean Kossaifi and Yannis Panagakis and Anima Anandkumar and Maja Pantic , title =. Journal of Machine Learning Research , year =

Showing first 80 references.

This paper was first reviewed by grok-4.5 on July 14, 2026.