REVIEW 6 minor 37 references
Trek-Based Parameter Identification for Linear Causal Models With Arbitrarily Structured Latent Variables
T0 review · 0 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A purely graphical criterion certifies when causal effects remain identifiable despite arbitrary latent structure.
desk verdict A genuinely new sufficient criterion for identifiability with arbitrary latent structure, honestly limited and largely sound; worth a serious round of review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the latent-subgraph criterion (LSC): a quadruple (Y,Z,H1,H2) satisfying size equalities relating Y to pa(v) and Z to H1,H2; trek separation of Y from Z ∪ {v} inside the latent subgraph Glat (the subgraph of all edges that are not tailed by an observed node); and the existence of a system of treks with no sided intersection from Y to pa(v) ∪ Z whose left parts, and whose right parts ending in Z, use only edges of Glat. The proof of the main theorem rests on a second new object, trek separation in subgraphs (Theorem 6.4), which generalizes the classical trek-separation criterion to block matrices whose entries come from different subgraphs; this is what guarantees that the block matrix being inverted has nonzero generic determinant. The decision algorithm is built from an integer linear program that extends maximum flow to settings where some flows are restricted to a subgraph.
What would settle it
Compute the symbolic determinant of the block matrix (A B) from Claim 5 for the graph in Figure 2 (b) with generic symbolic edge weights and the LSC tuple given in Example 3.7; if the determinant simplifies to the zero polynomial, Theorem 3.5 is false for that graph. More directly, search over small directed graphs for two distinct systems of directed paths from the same sources to the same sinks whose monomials are identical, one system having an intersection or a cycle; Lemma 6.2 asserts no such pair exists.
Extended reading notes
Core claim
The central claim is Theorem 3.5: if a 4-tuple (Y,Z,H1,H2) of observed and latent node sets satisfies the latent-subgraph criterion with respect to an observed node v, and all semi-direct effects into the auxiliary nodes Z ∪ (Y ∩ elr_{H2,H1}(Z ∪ {v})) are already known to be rationally identifiable, then all semi-direct effects into v from its semi-direct parents pa(v) are rationally identifiable. Because the criterion is purely graphical, it can be checked column by column: once every observed column is certified, the full semi-direct effect matrix Λ is identified as a rational function of the covariance matrix. The authors emphasize that this is, to their knowledge, the first identifiability criterion for semi-direct effects that does not assume latent nodes are source nodes, and they show by example that identifiability of a model and of its canonicalization are logically independent.
Load-bearing premise
The identification formula divides by a determinant, and the proof that this determinant is generically nonzero rests on a combinatorial lemma asserting that an intersection-free acyclic system of directed paths contributes a monomial to the determinant that no other path system can reproduce; if that uniqueness lemma had any counterexample, the whole column-by-column identification procedure would fail.
Editorial extensions
If this is right
- When the criterion certifies every observed column, the full semi-direct effect matrix $\Lambda$ is rationally identifiable from the observed covariance matrix $\Sigma$, so each such effect has an explicit closed-form estimator.
- Once $\Lambda$ is known, the residual matrix $\Omega = (I-\Lambda)^{\top}\Sigma(I-\Lambda)$ is also identifiable, and $\Omega$ is the covariance matrix of a simpler measurement model in which effects among latent variables can be identified by existing rules.
- The decision problem 'is G LSC-identifiable?' is handled by a sound and complete algorithm that solves integer linear programs; bounding $|H_1|+|H_2|$ makes the number of ILP calls polynomial in $|O|$ and $|L|$, provided Conjecture 4.3 holds, and without the bound the problem is NP-hard.
- Confounding-free acyclic graphs are rationally identifiable (Corollary 3.9), giving a broad structural class where the new criterion applies automatically.
Reading between the lines
- The new trek-separation-in-subgraphs condition is not tied to identifiability: a converse characterization of when such block determinants vanish could yield new polynomial constraints on covariance matrices, which in turn could be used for model equivalence and constraint-based testing, as the paper itself suggests in its discussion.
- Because the LSC is sufficient but not necessary, one can test how close it is to necessary by comparing it with dimension-based obstructions: for a random sparse graph that fails the LSC, the dimension of the image of the parametrization will often exceed the dimension of the identifiable parameter space, and a systematic comparison would quantify the gap.
- The independence of identifiability from canonicalization found in Section 5 suggests that empirical studies that reported non-identifiability of latent-variable structural equation models after canonicalization may have been too pessimistic; rechecking such examples with the LSC may turn some 'unidentified' effects into identified ones.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a graphical criterion, the latent-subgraph criterion (LSC), for rational identifiability of semi-direct causal effects in linear structural equation models with arbitrarily structured latent variables. The main result, Theorem 3.5, states that if a tuple (Y,Z,H1,H2) satisfies the LSC with respect to an observed node v and all semi-direct effects into a certain set Z ∪ (Y ∩ elr_{H2,H1}(Z ∪ {v})) are already rationally identifiable, then all semi-direct effects p⇝v with p ∈ pa(v) are rationally identifiable. The proof uses the factorization Σ = (I−Λ)^{-⊤}Ω(I−Λ)^{-1}, a new trek-separation-in-subgraphs determinant criterion (Theorem 6.4), and a constructive block linear system. The paper also gives a sound and complete algorithm for checking LSC-identifiability via an integer linear program (Algorithm 1, Theorem 4.7), presents numerical experiments, and compares the new criterion with the canonical model obtained by making all latent nodes source nodes.
Significance. If the main theorem holds, this is a substantial advance: it is, to my knowledge, the first identifiability criterion for semi-direct effects that does not require latent variables to be source nodes, and it applies to arbitrarily structured latent subgraphs. The proof is detailed and constructive, and the determinant-nonzero step is supported by a separate theorem with its own proof. The paper is honest about its limitations: Theorem 6.4 is only sufficient, as Example 6.5 shows, and Conjecture 4.3 about polynomial-time solvability of the integer program is left open. I specifically stress-tested Lemma 6.2, the most fragile combinatorial premise, and found the induction sound. The reproducible code for the simulations is a further strength. The paper is a strong fit for the journal and the central claim is, in my assessment, correct.
minor comments (6)
- [Lemma 4.5] In the statement and proof of Lemma 4.5, 'a subset YZ ⊆ Z' should read 'a subset YZ ⊆ Y'; as printed, the claim is nonsensical and inconsistent with Condition (iii) of the LSC.
- [Appendix A, Claim 3] Claim 3 states X ⊆ V \ (Z ∪ {v}), but the expressions Ω_{X,Z}, Ω_{X,v}, and Φ_{X,Z∪{v}} only make sense for X ⊆ O; please state X as a subset of the observed nodes to remove ambiguity.
- [Appendix A, Claim 5] The notation in Claim 5 overloads Λ, using it for both the full coefficient matrix and the semi-direct effect matrix; this overloading already appears in Section 2.1. Introducing distinct notation, e.g., Λ̄ for the semi-direct effect matrix, would make the dimension checks in the displayed block matrix transparent.
- [Section 6 and Section 7] There are minor typos: 'no system of of directed paths' in Example 6.5 and 'as we we show' in the discussion preceding it; please correct these.
- [Example 1.2] Equation (1.4) appears typeset incorrectly in the manuscript, with the square root sign and the fraction garbled; please repair the formula.
- [Remark 4.8] The complexity bound 'O2+kL2k' is not typeset properly; it should presumably read O(2^k |L|^{2k}).
Circularity Check
No significant circularity identified; the latent-subgraph criterion is a graph-combinatorial condition and the main identifiability proof is self-contained apart from standard external background results.
full rationale
The derivation chain is not circular. Definition 3.4 defines the latent-subgraph criterion purely graph-theoretically: conditions (i)-(iii) refer to cardinalities, trek separation, and the existence of a system of treks, not to the semidirect effects that the criterion is used to identify. Theorem 3.5 is an inductive statement: its hypothesis concerns effects into the strictly prior set Z ∪ (Y ∩ elr_{H2,H1}(Z ∪ {v})), while the conclusion concerns effects into v, which is excluded from that set. In the Appendix A proof, Claim 1 uses only these prior effects to establish rational identifiability of the matrices A, B, and c; no target coefficient λ_{p,v} appears in the induction hypothesis. Claim 5 reduces the required determinant nonvanishing of [A B] to Theorem 6.4, which is proved independently from Lemma 6.3 and Lemma 6.2. Lemma 6.2 establishes a purely combinatorial monomial-uniqueness fact for an acyclic vertex-disjoint directed path system; it does not assume the identifiability conclusion. The cited trek-separation facts from Sullivant et al. (2010) are external standard results, and self-citations such as Barber et al. (2022), Foygel et al. (2012), and Sturma et al. (2025) are contextual or used only for background remarks, not as load-bearing support for the main theorem. Explicit limitations, including the open Conjecture 4.3 and the non-if-and-only-if status of Theorem 6.4, are stated as such and do not make any step circular. No fitted parameters are renamed as predictions, and no known result is merely relabeled. I therefore find no circular step.
Assumptions & free parameters
assumptions (6)
- domain assumption Linear structural equation model with independent noise: X = Λ^T X + ε, ε independent with finite variance.
- domain assumption Invertibility of I-Λ and I-Λ_{L,L} for the parameter values considered.
- domain assumption Generic identifiability: failure allowed on a proper algebraic subset (measure zero).
- standard math Trek rule for covariance entries (Wright 1934; Spirtes et al. 2000).
- standard math Trek separation rank theorem of Sullivant et al. (2010).
- domain assumption Acyclicity or a valid recursive order for the algorithm to make progress.
Cite this review
Pith. "Pith review of Trek-Based Parameter Identification for Linear Causal Models With Arbitrarily Structured Latent Variables." pith.science (2026). https://pith.science/paper/5TGPHTPP
@misc{pith2026250718170,
author = {Pith},
title = {Pith review of: Trek-Based Parameter Identification for Linear Causal Models With Arbitrarily Structured Latent Variables},
year = {2026},
howpublished = {\url{https://pith.science/paper/5TGPHTPP}},
note = {Machine review of arXiv:2507.18170}
}
read the original abstract
We develop a criterion to certify whether causal effects are identifiable in linear structural equation models with latent variables. Linear structural equation models correspond to directed graphs whose nodes represent the random variables of interest and whose edges are weighted with linear coefficients that correspond to direct causal effects. In contrast to previous identification methods, we do not restrict ourselves to settings where the latent variables constitute independent latent factors (i.e., to source nodes in the graphical representation of the model). Our novel latent-subgraph criterion is a purely graphical condition that is sufficient for identifiability of causal effects by rational formulas in the covariance matrix. To check the latent-subgraph criterion, we provide a sound and complete algorithm that operates by solving an integer linear program. While it targets effects involving observed variables, our new criterion is also useful for identifying effects between latent variables, as it allows one to transform the given model into a simpler measurement model for which other existing tools become applicable.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
H., Chernozhukov, V., and Fernández-Val, I
Abbring, J. H., Chernozhukov, V., and Fernández-Val, I. (2025). P hilip G . W right, directed acyclic graphs, and instrumental variables. The Econometrics Journal , 28(1):1--20
work page 2025
-
[2]
Ankan, A., Wortel, I., Bollen, K., and Textor, J. (2023). Combining graphical and algebraic approaches for parameter identification in latent variable structural equation models. In Ruiz, F., Dy, J., and van de Meent, J.-W., editors, Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , volume 206 of Proceedings of M...
work page 2023
-
[3]
Baja, E. S., Schwartz, J. D., Coull, B. A., Wellenuis, G. A., Vokonas, P. S., and Suh, H. H. (2013). Structural equation modeling of the inflammatory response to traffic air pollution. Journal of Exposure Science & Environmental Epidemiology , 23(3):268–274
work page 2013
-
[4]
F., Drton, M., Sturma, N., and Weihs, L
Barber, R. F., Drton, M., Sturma, N., and Weihs, L. (2022). Half-trek criterion for identifiability of latent variable models . The Annals of Statistics , 50(6):3174 -- 3196
work page 2022
-
[5]
Bollen, K. A. (1989). Structural equations with latent variables . Wiley Series in Probability and Mathematical Statistics: Applied Probability and Statistics. John Wiley & Sons, Inc., New York. A Wiley-Interscience Publication
work page 1989
-
[6]
Bollen, K. A. and Bauldry, S. (2011). Three C s in measurement models: Causal indicators, composite indicators, and covariates. Psychological Methods , 16(3):265–284
work page 2011
-
[7]
Brito, C. and Pearl, J. (2006). Graphical condition for identification in recursive SEM . In Dechter, R. and Richardson, T. S., editors, Proceedings of the 22nd Conference on Uncertainty in Artificial Intelligence , pages 47--54. AUAI Press
work page 2006
-
[8]
H., Leiserson, C
Cormen, T. H., Leiserson, C. E., Rivest, R. L., and Stein, C. (2009). Introduction to algorithms . MIT Press, Cambridge, MA, third edition
2009
Show all 37 references
-
[9]
Cox, D., Little, J., and O'Shea, D. (2007). Ideals, varieties, and algorithms . Undergraduate Texts in Mathematics. Springer, New York, third edition. An introduction to computational algebraic geometry and commutative algebra
2007
-
[10]
Dong, X., Ng, I., Huang, B., Sun, Y., Jin, S., Legaspi, R., Spirtes, P., and Zhang, K. (2024). On the parameter identifiability of partially observed linear causal models. In Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., and Zhang, C., editors, Adv...
2024
-
[11]
Drton, M., Robeva, E., and Weihs, L. (2020). Nested covariance determinants and restricted trek separation in G aussian graphical models. Bernoulli , 26(4):2503--2540
2020
-
[12]
Foygel, R., Draisma, J., and Drton, M. (2012). Half-trek criterion for generic identifiability of linear structural equation models. The Annals of Statistics , 40(3):1682--1713
2012
-
[13]
D., Spielvogel, S., and Sullivant, S
Garcia-Puente, L. D., Spielvogel, S., and Sullivant, S. (2010). Identifying causal effects with computer algebra. In Gr\" u nwald, P. and Spirtes, P., editors, Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence . AUAI Press
2010
-
[14]
Henckel, L., Buttenschoen, M., and Maathuis, M. H. (2024). Graphical tools for selecting conditional instrumental sets. Biometrika , 111(3):771--788
2024
-
[15]
O., Shimizu, S., Kerminen, A
Hoyer, P. O., Shimizu, S., Kerminen, A. J., and Palviainen, M. (2008). Estimation of causal effects using linear non- G aussian causal models with hidden variables. International Journal of Approximate Reasoning , 49(2):362--378
2008
-
[16]
and Vreeken, J
Kaltenpoth, D. and Vreeken, J. (2023). Nonlinear causal discovery with latent confounders. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J., editors, Proceedings of the 40th International Conference on Machine Learning , volume 202 of Proceed...
2023
-
[17]
Kumor, D., Cinelli, C., and Bareinboim, E. (2020). Efficient identification in linear structural causal models with auxiliary cutsets. In III, H. D. and Singh, A., editors, Proceedings of the 37th International Conference on Machine Learning , volume 119 of Proceedings of Mach...
2020
-
[18]
Mayer, A. (2019). Causal effects based on latent variable models. Methodology , 15(Supplement 1):15--28
2019
-
[19]
Nemhauser, G. L. and Wolsey, L. A. (1988). Integer and combinatorial optimization . Wiley-Interscience Series in Discrete Mathematics and Optimization. John Wiley & Sons, Inc., New York. A Wiley-Interscience Publication
1988
-
[20]
Okamoto, M. (1973). Distinctness of the eigenvalues of a quadratic form in a multivariate sample. The Annals of Statistics , 1:763--765
1973
-
[21]
Pearl, J. (2009). Causality . Cambridge University Press, Cambridge, second edition. Models, reasoning, and inference
2009
-
[22]
A., and Nolan, G
Sachs, K., Perez, O., Pe'er, D., Lauffenburger, D. A., and Nolan, G. P. (2005). Causal protein-signaling networks derived from multiparameter single-cell data. Science , 308(5721):523--529
2005
-
[23]
J., Rosenfeld, E., Ravikumar, P
Saengkyongam, S. J., Rosenfeld, E., Ravikumar, P. K., Pfister, N., and Peters, J. (2024). Identifying representations for intervention extrapolation. In Kim, B., Yue, Y., Chaudhuri, S., Fragkiadaki, K., Khan, M., and Sun, Y., editors, International Conference on Representation...
2024
-
[24]
Salehkaleybar, S., Ghassami, A., Kiyavash, N., and Zhang, K. (2020). Learning linear non- G aussian causal models in the presence of latent variables. Journal of Machine Learning Research , 21(39):1--24
2020
-
[25]
Schrijver, A. (1986). Theory of linear and integer programming . Wiley-Interscience Series in Discrete Mathematics. John Wiley & Sons, Ltd., Chichester. A Wiley-Interscience Publication
1986
-
[26]
R., Kalchbrenner, N., Goyal, A., and Bengio, Y
Schölkopf, B., Locatello, F., Bauer, S., Ke, N. R., Kalchbrenner, N., Goyal, A., and Bengio, Y. (2021). Toward causal representation learning. Proceedings of the IEEE , 109(5):612--634
2021
-
[27]
Shafarevich, I. R. (2013). Basic Algebraic Geometry 1 . Springer Berlin Heidelberg, third edition. Varieties in projective space
2013
-
[28]
Spirtes, P., Glymour, C., and Scheines, R. (2000). Causation, prediction, and search . Adaptive Computation and Machine Learning. MIT Press, Cambridge, MA, second edition. With additional material by David Heckerman, Christopher Meek, Gregory F. Cooper and Thomas Richardson, A...
2000
-
[29]
S., and Uhler, C
Squires, C., Seigal, A., Bhate, S. S., and Uhler, C. (2023). Linear causal disentanglement via interventions. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J., editors, Proceedings of the 40th International Conference on Machine Learning , vo...
2023
-
[30]
F., Zhou, X., and Steenbergen, M
Stoetzer, L. F., Zhou, X., and Steenbergen, M. (2024). Causal inference with latent outcomes. American Journal of Political Science , page ajps.12871
2024
-
[31]
Sturma, N., Kranzlmueller, M., Portakal, I., and Drton, M. (2025). Matching criterion for identifiability in sparse factor analysis. arXiv preprint arXiv:2502.02986
2025
-
[32]
Sturma, N., Squires, C., Drton, M., and Uhler, C. (2023). Unpaired multi-domain causal representation learning. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S., editors, Advances in Neural Information Processing Systems , volume 36, pages 34465--34...
2023
-
[33]
Sullivant, S., Talaska, K., and Draisma, J. (2010). Trek separation for G aussian graphical models. The Annals of Statistics , 38(3):1665--1685
2010
-
[34]
Tramontano, D., Kivva, Y., Salehkaleybar, S., Drton, M., and Kiyavash, N. (2024). Causal effect identification in L i NGAM models with latent confounders. In Salakhutdinov, R., Kolter, Z., Heller, K., Weller, A., Oliver, N., Scarlett, J., and Berkenkamp, F., editors, Proceedin...
2024
-
[35]
I., Nguyen, N., Robeva, E., and Drton, M
Weihs, L., Robinson, B., Dufresne, E., Kenkel, J., Kubjas Reginald McGee II, K., Reginald, M. I., Nguyen, N., Robeva, E., and Drton, M. (2017). Determinantal generalizations of instrumental variables. Journal of Causal Inference , 6(1)
2017
-
[36]
Wright, S. (1934). The method of path coefficients. The Annals of Mathematical Statistics , 5(3):161–215
1934
-
[37]
Xie, F., Cai, R., Huang, B., Glymour, C., Hao, Z., and Zhang, K. (2020). Generalized independent noise condition for estimating latent variable causal graphs. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H., editors, Advances in Neural Information Processi...
2020
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.